Notice
先看证据,再决定买不买

频道每天最多 3 条价格异动与中转状态;具体商品请用机器人设置降价/补货提醒。交流群提问请带预算、模型、工具和使用频率。

View
Community & contactTelegram 群点击加入Telegram 频道每天最多 3 条有效价格情报联系我们tgAIPricedb交流群979789483
Back to news
Products

Z.ai Launches GLM-5.3-Flash with 1M-Token Context and MIT License

Z.ai has announced GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with approximately 18 billion active parameters per token, native multimodal support, and a context window of up to 1,048,576 tokens.

86% VERIFIED

The model was previously previewed under the codename “Ox Alpha,” which Z.ai says ran entirely on Chinese AI chips. With the public release, its weights are offered under the MIT License and the model is available through Hugging Face, Z.ai’s API, and other official services.

Z.ai lists standard API rates of $0.15 per million input tokens, $0.50 per million output tokens, and $0.03 per million cached input tokens. Industry coverage presents the model as a lower-cost option for coding and agentic workloads, while benchmark and cost comparisons should be understood as reported or vendor-supplied figures rather than universal guarantees.

Source evidence

GLM-5.3-Flash: Features, Benchmarks, and Pricingdatacamp.com · supporting

If you have been watching the race between open-weight Chinese labs and the Western frontier, this one is worth a look. The main claim is simple: near-frontier coding and agentic performance at a fraction of the cost. GLM-5.3-Flash scores 84.3 on Terminal-Bench 2.1, which puts it within striking distance of Claude Opus 4.8 (85.0) and GPT-5.6 Terra (87.4). It uses a mixture-of-experts design in a 320B-A18B configuration (320B aggregate parameters, 18B active per token), supports a 1M token context window, and is natively multimodal. [...] ## In a Nutshell GLM-5.3-Flash is Z.ai's cost-optimized model, codenamed "Ox Alpha," with a 320B-A18B MoE design and 1M token context. It scores 84.3 on Terminal-Bench 2.1, near Claude Opus 4.8 (85.0) and behind GPT-5.6 Terra (87.4). Blended price is roughly $0.10 per 1M tokens, about a tenth of GLM-5.3's ~$0.90. Trade-offs: ~49 tokens/sec output speed and a ~306 GiB FP8 checkpoint that rules out lightweight local hosting. Qwen3.8-Fl

GLM-5.3-Flash: Multimodal, MIT-Licensed, 1M Context - Eigent AIeigent.ai · supporting

## What is GLM-5.3-Flash? GLM-5.3-Flash is a Mixture-of-Experts model with 320 billion total parameters and 18 billion active parameters per token (TestingCatalog). Z.ai positions it as the lower-cost, higher-efficiency tier of the GLM-5 family: strong coding and agentic performance, but built to run cheaply and respond fast. Two things set it apart from earlier GLM releases. It's the first natively multimodal model in the line — vision is baked into the training, not bolted on — and it ships under the MIT license with open weights on Hugging Face. If the name "Ox Alpha" sounds familiar, that's because this is the same model: it ran anonymously as ox-alpha on OpenCode and OpenRouter for about a week before the reveal. ## The hybrid architecture: cheaper attention at 1M tokens [...] logo Blogs Industry|Aug 26, 2026 # GLM-5.3-Flash: Z.ai's Multimodal Model at One-Tenth the Price The first natively multimodal GLM-5 model — 320B-A18B, MIT-licensed, 1M context, and priced to run agen

Introducing GLM-5.3-Flash, a multimodal AI modelfacebook.com · supporting

## Md Al Mamun's Post ### Md Al Mamun AI content · 21h · OX Alpha is now GLM 5.3 Flash 🔥 Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips #AITools May be an image of text that says 'NEW UPDATE! OX Alpha z GLM 5.3 Flash is now GLM 5.3 Flash X Alpha Goodevening,AlMamun Good evening, Mamun How to use Ox Alpha for FREE today? Cmи3 Irage Cooleintirav WibSearth FASTER © DeepThirk acba SMARTER MORE ACCURATE FREE 100% FREE TO USE'

Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively ...marktechpost.com · supporting

By Asif Razzaq - Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with 320B total parameters and 18B active per token, a 1,048,576-token context window, and image and video input — released under an MIT license with weights on Hugging Face. According to Z.ai reports, it beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, while landing within half a point of Claude Opus 4.8 on its internal coding benchmark. The model spent its first week running anonymously as “Ox Alpha” on OpenCode and OpenRouter, served entirely on domestically produced Chinese AI chips. ## Is it deployable? [...] ## Key Takeaways 320B-A18B natively multimodal MoE, 1M context, MIT-licensed weights on Hugging Face. Hybrid KDA linear + NoPE sparse MLA attention: ~3× less attention compute, 4.4× smaller KV cache. 84.3 Terminal-Bench 2.1 and 63.4 DeepSWE

GLM-5.3-Flash: Specs, Benchmarks, and Running It Locally | MindStudiomindstudio.ai · supporting

## What is GLM-5.3-Flash? GLM-5.3-Flash is Zhipu AI’s (Z.ai’s) open-weight model in the GLM-5 lineup, and the first natively multimodal model in that series. It uses a mixture-of-experts (MoE) design with 320 billion total parameters but only 18 billion active per forward pass, paired with a hybrid sparse and linear attention architecture that keeps long-context inference affordable. It ships under the MIT license and supports a 1 million token context window, and it spent its early life running anonymously on public leaderboards under the codename “Ox Alpha” before Z.ai revealed its identity. ## TL;DR [...] GLM-5.3-Flash runs on a 320B-total, 18B-active MoE architecture, meaning it only computes a fraction of its parameters per token while keeping the full model’s knowledge in reserve. It introduces a hybrid sparse and linear attention scheme, the first time GLM has combined these approaches, specifically to make long-context serving cheaper without losing accuracy at scale. The m

Z.ai on X: "Introducing GLM-5.3-Flashx.com · supporting

980 @Zai_org Z.ai @Zai\_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-f… Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5.… API: docs.z.ai/guides/llm/glm… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai 2:12 PM · Aug 26, 20266.7MViews 980 @Zai_org Z.ai @Zai\_org Aug 26 Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03 74 @Zai_org Z.ai @Zai\_org Aug 26 [...] @Zai_org @Zai\_org Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously pre