A very helpful comparison of endpoint accuracy across GLM and other open-source models. It would be great to see the community test the official GLM-5.2 API as an additional reference point. It may score above 100%.
Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon
Providers trade off





