公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
product

Z.ai称GLM-5.2在移动应用开发测试中较GLM-5.1显著提升

Z.ai表示,GLM-5.2在一项内部移动应用开发基准测试中完成了70次试验中的48次,高于GLM-5.1的21次。该测试覆盖35项高难度任务,每项运行两次。

76% VERIFIED

Z.ai在官方社交账号和博客中介绍了GLM-5.2,称其面向长周期任务,并强化了编码和智能体工作流能力。公司表示,在内部移动应用开发测试中,GLM-5.1完成21/70次,GLM-5.2完成48/70次;帖子还列出一个名为“Claude Fable 5”的对比结果,为56/70次。

这项结果由Z.ai自行披露,测试包含35项任务、每项重复两次,完成标准是核心功能可用且不存在重大问题。目前公开材料未提供足以独立复现的完整方法和结果明细,因此不能将该成绩视为广泛适用的第三方结论。官方资料还称,GLM-5.2支持约100万令牌上下文窗口,并提供可调节的推理强度。

来源证据

glm-5.2 Model by Z-aibuild.nvidia.com · supporting

# GLM-5.2 ## Description: GLM-5.2 is the latest flagship large language model from Z.ai (zai-org), designed for long-horizon tasks with a solid 1M-token context window. It represents a substantial leap in extended-context capability over its predecessor GLM-5.1, featuring the IndexShare architecture that reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at 1M context length. GLM-5.2 delivers state-of-the-art performance across reasoning, coding, and agentic benchmarks, with multiple thinking effort levels to balance performance and latency. GLM-5.2 was developed by Z.ai (zai-org) as a part of the GLM model family. _This model is ready for commercial or non-commercial use._ ## Third-Party Community Consideration [...] ### Deployment Geography: Global ### Use Case: Developers and researchers can use GLM-5.2 for long-horizon reasoning tasks requiring extended context, complex software engineering and agentic workflows, mathematic

GLM-5.2 & GLM-5.1 & GLM-5github.com · supporting

# GLM-5.2 & GLM-5.1 & GLM-5 👋 Join our Wechat or Discord community. 📖 Check out the GLM-5.2 blog and GLM-5 Technical report. 📍 Use GLM-5.2 API services on Z.ai API Platform. 🔜 Try GLM-5.2 at z.ai. ## Introduction ### GLM-5.2 GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: [...] GLM-5.2's new capabilities include: Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency Improved Architecture: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP layer for specu

README.md · zai-org/GLM-5.2 at mainhuggingface.co · supporting

[Paper] [GitHub] ## Introduction We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: bench_52 bench_52 ## Benchmark

GLM-5.2 boosts app dev capabilities by 2x | Z.ai posted on the topic | LinkedInlinkedin.com · supporting

Agree & Join LinkedIn By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy. View organization page for Z.ai 13,900 followers GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an internal benchmark of 35 challenging mobile development tasks, each run twice for a total of 70 trials. We measured task completion, defined as core features working without major issues. Jen W., graphic [...] Jen W., graphic Impressive jump, especially given how long-horizon app-development tasks are. I’ve been analyzing the anti-hack section of your blog post, and it seems like app-building and debugging may have different reward-hacking modes: app-building can hack by downloading/copying the target solution, while debugging can

Z.ai on X: "Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way." / Xx.com · supporting

Log inSign up ## Post user avatar Z.ai @Zai\_org Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way. user avatar Zixuan Li Z.ai @ZixuanLi\_ Jun 19 GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Results: - GLM-5.1: 21/70 - GLM-5.2: 48/70 - Claude Fable 5: 56/70 That's more than a twofold improvement from GLM-5.1 to GLM-5.2. These come from an 00:00 3:05 AM · Jun 19, 202671KViews user avatar Himanshu @codingstark Jun 19 ngl, I tested glm 5.2 , and it's the best OSS model I've used so far. pretty impressive. 517 user avatar Revoxan @Revoxan [...] 517 user avatar Revoxan @Revoxan Jun 19 21 → 48 in one generation.Claude Fable 5 is still ahead at 56, but at this pace, the gap won’t last long. 418 user avatar SunX @SunX\

GLM-5.2: Built for Long-Horizon Tasksz.ai · supporting

Image 1 2026-06-16 · Research # GLM-5.2: Built for Long-Horizon Tasks Image 2Try it at Z.ai Image 3Call it at Z.ai Image 4Z.ai Coding Plan Image 5GitHub Image 6HuggingFace We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: [...] GLM-5.2 also introduces effort level control, enabling users to explicitly balance model capability against task execution speed and computational cost. As shown in the figure, GLM-5.2 delivers substantially stronger agentic coding performance than GLM-5.1 at comparable token budgets, with its capability roughly positioned between Claude Opus 4.7 and Claude Opus 4.8 under similar token consumption. Moreover, the Max effort level allows users to allocate additional computation when higher performance is required in challenging task