公告
先看证据,再决定买不买

频道每天最多 3 条价格异动与中转状态;具体商品请用机器人设置降价/补货提醒。交流群提问请带预算、模型、工具和使用频率。

查看
社群与联系Telegram 群点击加入Telegram 频道每天最多 3 条有效价格情报联系我们tgAIPricedb交流群979789483
返回资讯列表
product

智谱发布 GLM-5.3-Flash:320B MoE、百万级上下文并采用 MIT 许可证

Z.ai 宣布推出 GLM-5.3-Flash,这是一款总参数量 3200 亿、每次激活约 180 亿参数的混合专家模型,支持原生多模态输入和 1,048,576 个 token 的上下文窗口。

86% VERIFIED

据 Z.ai 公布的信息,GLM-5.3-Flash 此前曾以“Ox Alpha”这一名称进行预览,并被描述为完全运行在中国国产 AI 芯片上的模型。此次正式发布后,模型权重以 MIT 许可证开放,并可通过 Hugging Face、Z.ai API 及其官方产品使用。

官方列出的标准 API 价格为:每百万输入 token 0.15 美元、每百万输出 token 0.50 美元、每百万缓存输入 token 0.03 美元。多家行业媒体还将其定位为面向编码和智能体任务的低成本版本;相关性能与成本比较属于报道或厂商口径,实际结果会因任务和部署方式而异。

来源证据

GLM-5.3-Flash: Features, Benchmarks, and Pricingdatacamp.com · supporting

If you have been watching the race between open-weight Chinese labs and the Western frontier, this one is worth a look. The main claim is simple: near-frontier coding and agentic performance at a fraction of the cost. GLM-5.3-Flash scores 84.3 on Terminal-Bench 2.1, which puts it within striking distance of Claude Opus 4.8 (85.0) and GPT-5.6 Terra (87.4). It uses a mixture-of-experts design in a 320B-A18B configuration (320B aggregate parameters, 18B active per token), supports a 1M token context window, and is natively multimodal. [...] ## In a Nutshell GLM-5.3-Flash is Z.ai's cost-optimized model, codenamed "Ox Alpha," with a 320B-A18B MoE design and 1M token context. It scores 84.3 on Terminal-Bench 2.1, near Claude Opus 4.8 (85.0) and behind GPT-5.6 Terra (87.4). Blended price is roughly $0.10 per 1M tokens, about a tenth of GLM-5.3's ~$0.90. Trade-offs: ~49 tokens/sec output speed and a ~306 GiB FP8 checkpoint that rules out lightweight local hosting. Qwen3.8-Fl

GLM-5.3-Flash: Multimodal, MIT-Licensed, 1M Context - Eigent AIeigent.ai · supporting

## What is GLM-5.3-Flash? GLM-5.3-Flash is a Mixture-of-Experts model with 320 billion total parameters and 18 billion active parameters per token (TestingCatalog). Z.ai positions it as the lower-cost, higher-efficiency tier of the GLM-5 family: strong coding and agentic performance, but built to run cheaply and respond fast. Two things set it apart from earlier GLM releases. It's the first natively multimodal model in the line — vision is baked into the training, not bolted on — and it ships under the MIT license with open weights on Hugging Face. If the name "Ox Alpha" sounds familiar, that's because this is the same model: it ran anonymously as ox-alpha on OpenCode and OpenRouter for about a week before the reveal. ## The hybrid architecture: cheaper attention at 1M tokens [...] logo Blogs Industry|Aug 26, 2026 # GLM-5.3-Flash: Z.ai's Multimodal Model at One-Tenth the Price The first natively multimodal GLM-5 model — 320B-A18B, MIT-licensed, 1M context, and priced to run agen

Introducing GLM-5.3-Flash, a multimodal AI modelfacebook.com · supporting

## Md Al Mamun's Post ### Md Al Mamun AI content · 21h · OX Alpha is now GLM 5.3 Flash 🔥 Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips #AITools May be an image of text that says 'NEW UPDATE! OX Alpha z GLM 5.3 Flash is now GLM 5.3 Flash X Alpha Goodevening,AlMamun Good evening, Mamun How to use Ox Alpha for FREE today? Cmи3 Irage Cooleintirav WibSearth FASTER © DeepThirk acba SMARTER MORE ACCURATE FREE 100% FREE TO USE'

Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively ...marktechpost.com · supporting

By Asif Razzaq - Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with 320B total parameters and 18B active per token, a 1,048,576-token context window, and image and video input — released under an MIT license with weights on Hugging Face. According to Z.ai reports, it beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, while landing within half a point of Claude Opus 4.8 on its internal coding benchmark. The model spent its first week running anonymously as “Ox Alpha” on OpenCode and OpenRouter, served entirely on domestically produced Chinese AI chips. ## Is it deployable? [...] ## Key Takeaways 320B-A18B natively multimodal MoE, 1M context, MIT-licensed weights on Hugging Face. Hybrid KDA linear + NoPE sparse MLA attention: ~3× less attention compute, 4.4× smaller KV cache. 84.3 Terminal-Bench 2.1 and 63.4 DeepSWE

GLM-5.3-Flash: Specs, Benchmarks, and Running It Locally | MindStudiomindstudio.ai · supporting

## What is GLM-5.3-Flash? GLM-5.3-Flash is Zhipu AI’s (Z.ai’s) open-weight model in the GLM-5 lineup, and the first natively multimodal model in that series. It uses a mixture-of-experts (MoE) design with 320 billion total parameters but only 18 billion active per forward pass, paired with a hybrid sparse and linear attention architecture that keeps long-context inference affordable. It ships under the MIT license and supports a 1 million token context window, and it spent its early life running anonymously on public leaderboards under the codename “Ox Alpha” before Z.ai revealed its identity. ## TL;DR [...] GLM-5.3-Flash runs on a 320B-total, 18B-active MoE architecture, meaning it only computes a fraction of its parameters per token while keeping the full model’s knowledge in reserve. It introduces a hybrid sparse and linear attention scheme, the first time GLM has combined these approaches, specifically to make long-context serving cheaper without losing accuracy at scale. The m

Z.ai on X: "Introducing GLM-5.3-Flashx.com · supporting

980 @Zai_org Z.ai @Zai\_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-f… Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5.… API: docs.z.ai/guides/llm/glm… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai 2:12 PM · Aug 26, 20266.7MViews 980 @Zai_org Z.ai @Zai\_org Aug 26 Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03 74 @Zai_org Z.ai @Zai\_org Aug 26 [...] @Zai_org @Zai\_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously pre