GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks.
Results:
- GLM-5.1: 21/70
- GLM-5.2: 48/70
- Claude Fable 5: 56/70
That's more than a twofold improvement from GLM-5.1 to GLM-5.2.
These come from an internal benchmark of 35 challenging mobile development tasks, each run twice for a total of 70 trials. We measured task completion, defined as core features working without major issues.
Zixuan Li on X: "GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an https://t.co/SkbQKvkFTf"
- Will share a demo prompt shortly if people are interested. (Much more complex than the prompt in the video, which we summarized to make it look nice.)
- The question is: where are the GLM iOS and Android apps??? The best way to show what the model is capable of is by building world class apps used by billions of people from all around the world.


