Z.ai заявила о резком росте возможностей GLM-5.2 для разработки приложений
По данным Z.ai, GLM-5.2 выполнила 48 из 70 испытаний во внутреннем тесте мобильной разработки против 21 у GLM-5.1. В тест вошли 35 сложных задач, каждая запускалась дважды.
Z.ai представила GLM-5.2 как модель для длительных многоэтапных задач с улучшенными возможностями программирования и агентных рабочих процессов. Во внутреннем тесте разработки мобильных приложений модель выполнила 48 из 70 испытаний, тогда как GLM-5.1 получила результат 21 из 70. В публикациях также указан результат 56 из 70 для модели под названием «Claude Fable 5».
Компания оценивала успешность по тому, работали ли ключевые функции без существенных проблем. Поскольку полный набор задач, детали методики и независимое воспроизведение результатов в приведённых источниках отсутствуют, эти показатели следует считать данными самой Z.ai, а не окончательным доказательством превосходства модели. В официальных материалах также заявлены контекстное окно примерно на 1 млн токенов и регулируемые уровни вычислительного усилия.
Источники
glm-5.2 Model by Z-aibuild.nvidia.com · supporting# GLM-5.2 ## Description: GLM-5.2 is the latest flagship large language model from Z.ai (zai-org), designed for long-horizon tasks with a solid 1M-token context window. It represents a substantial leap in extended-context capability over its predecessor GLM-5.1, featuring the IndexShare architecture that reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at 1M context length. GLM-5.2 delivers state-of-the-art performance across reasoning, coding, and agentic benchmarks, with multiple thinking effort levels to balance performance and latency. GLM-5.2 was developed by Z.ai (zai-org) as a part of the GLM model family. _This model is ready for commercial or non-commercial use._ ## Third-Party Community Consideration [...] ### Deployment Geography: Global ### Use Case: Developers and researchers can use GLM-5.2 for long-horizon reasoning tasks requiring extended context, complex software engineering and agentic workflows, mathematic
GLM-5.2 & GLM-5.1 & GLM-5github.com · supporting# GLM-5.2 & GLM-5.1 & GLM-5 👋 Join our Wechat or Discord community. 📖 Check out the GLM-5.2 blog and GLM-5 Technical report. 📍 Use GLM-5.2 API services on Z.ai API Platform. 🔜 Try GLM-5.2 at z.ai. ## Introduction ### GLM-5.2 GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: [...] GLM-5.2's new capabilities include: Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency Improved Architecture: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP layer for specu
README.md · zai-org/GLM-5.2 at mainhuggingface.co · supporting[Paper] [GitHub] ## Introduction We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: bench_52 bench_52 ## Benchmark
GLM-5.2 boosts app dev capabilities by 2x | Z.ai posted on the topic | LinkedInlinkedin.com · supportingAgree & Join LinkedIn By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy. View organization page for Z.ai 13,900 followers GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an internal benchmark of 35 challenging mobile development tasks, each run twice for a total of 70 trials. We measured task completion, defined as core features working without major issues. Jen W., graphic [...] Jen W., graphic Impressive jump, especially given how long-horizon app-development tasks are. I’ve been analyzing the anti-hack section of your blog post, and it seems like app-building and debugging may have different reward-hacking modes: app-building can hack by downloading/copying the target solution, while debugging can
Z.ai on X: "Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way." / Xx.com · supportingLog inSign up ## Post user avatar Z.ai @Zai\_org Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way. user avatar Zixuan Li Z.ai @ZixuanLi\_ Jun 19 GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an 00:00 3:05 AM · Jun 19, 202671KViews user avatar Himanshu @codingstark Jun 19 ngl, I tested glm 5.2 , and it's the best OSS model I've used so far. pretty impressive. 517 user avatar Revoxan @Revoxan [...] 517 user avatar Revoxan @Revoxan Jun 19 21 → 48 in one generation.Claude Fable 5 is still ahead at 56, but at this pace, the gap won’t last long. 418 user avatar SunX @SunX\
GLM-5.2: Built for Long-Horizon Tasksz.ai · supportingImage 1 2026-06-16 · Research # GLM-5.2: Built for Long-Horizon Tasks Image 2Try it at Z.ai Image 3Call it at Z.ai Image 4Z.ai Coding Plan Image 5GitHub Image 6HuggingFace We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: [...] GLM-5.2 also introduces effort level control, enabling users to explicitly balance model capability against task execution speed and computational cost. As shown in the figure, GLM-5.2 delivers substantially stronger agentic coding performance than GLM-5.1 at comparable token budgets, with its capability roughly positioned between Claude Opus 4.7 and Claude Opus 4.8 under similar token consumption. Moreover, the Max effort level allows users to allocate additional computation when higher performance is required in challenging task