公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
product

Kimi K3 在长上下文与智能体任务中表现突出,但全面领先仍有争议

现有材料显示,Kimi K3 拥有约 100 万 token 的上下文窗口,并在部分编码、设计和智能体评测中取得很强成绩。不过,公开数据并未证明它全面超过 Fable 5 和 GPT-5.6 Sol;有关 Fable 5 被悄悄降级的说法也缺乏直接证据。

45% VERIFIED

Kimi K3 的主要优势包括超长上下文、原生多模态能力,以及面向长流程编程和知识工作的定位。部分第三方评测和实测项目认为,它在前端开发、视觉设计和工具调用任务中表现出色,并具有较高的性价比。

但不同评测得出的结论并不一致。所提供的 GDPval v2 数据中,K3 的 1668 Elo 低于 Fable 5 的 1760 和 GPT-5.6 Sol 的 1748;另一些编码和设计测试则显示 K3 在特定场景中领先。因此,更稳妥的结论是 K3 已进入顶尖模型竞争范围,而不是在所有领域都排名第一。

关于 Fable 5 因网络安全问题被自动降级为 Opus 4.8 的说法,目前材料不足以确认。相关来源提到回退配置或潜在停用,但没有证明存在所描述的“悄悄降级”。

来源证据

Claude Fable 5 vs Kimi K3 vs GPT-5.6 Solcomposio.dev · supporting

At this price, it somehow feels even more impressive than Fable 5. Kimi K3 has now overshadowed GPT-5.6 Sol in both coding tests, which is honestly a little scary for an open model. Here’s a quick demo: You can find the generated code here: Kimi K3: Disasters and Composio Cost: $10.90 for this task, $25.92 cumulative Duration: ~1 hour Token Usage: 117,845 tokens for this task, 393,452 cumulative ## Conclusion So far, the two benchmarks are telling very different stories. On the coding side, Fable 5 delivered the strongest raw visuals, but Kimi K3 was the biggest surprise. It produced a far more polished and functional result than GPT-5.6 Sol, handled the disaster logic and Composio integration well, and did it at a fraction of Fable’s cost. [...] The biggest warning is still the frontier-kill set. All three models failed all five cases. These exact-state, cross-app workflows remain difficult enough that I would not deploy any of them without verification, retries, and strict a

Claude Opus 5 vs GPT-5.6 Sol vs Fable 5 vs Kimi K3 in 2026 — Fenxifenxi.fr · supporting

Claude Fable 5 remains a premium model for genuinely exceptional tasks: certain analytical, scientific or very long software jobs. Its price and retention constraints make it hard to generalise. Kimi K3 is the competitor to watch: a one million token context, native visual understanding and the lowest price of the four. Its value in production will depend on whether its weights and licence actually ship. We cover it in depth in our dedicated Kimi K3 article. [...] | Kimi K3 | $3 | $15 | 1M tokens | Frontend, native vision, long context, price, future open weights | API available; European residency guarantees and contract terms need auditing | [...] | Main need | Recommended model | Why | --- | General purpose frontier model | Claude Opus 5 | Good balance of capability, cost and output quality | | Agentic coding with Codex | GPT-5.6 Sol | Leads the Coding Agent Index in the Codex environment | | Longest and most complex tasks | Claude Fable 5 | Very strong in analysis, software engi

Kimi K3 对比 GPT-5.5 和 Opus 4.8(2026):更便宜的同级选手ofox.ai · supporting

| 模型 | GDPval v2 Elo(Artificial Analysis) | --- | | Claude Fable 5 | 1760 | | GPT-5.6 Sol (max) | 1748 | | Kimi K3 | 1668 | | Claude Opus 4.8 | 1600 | | GLM-5.2 | 1514 | | GPT-5.5 | 1494 | | Kimi K2.6 | 1190 | 这是一个实打实的结果,不是包装话术:在智能体工具调用和多步骤工作上,K3 的跑分超过本文拿来对比的这两款模型。GDPval v2 评的是真实、具经济价值的任务而非琐碎知识,所以这里的领先能映射到编码或运维智能体实际要做的那类工作。K3 还在 AA 的 AutomationBench(他们对 Zapier 智能体 SaaS 工作流评测的实现)上位居榜首,并在 AA-Briefcase(一个私有的长周期知识工作评测)上达到 1547,仅次于 Fable 5,比 Kimi K2.6 高出 732 分。三个不同的智能体测试,同一幅图景:K3 在每一项上都处于或接近顶端,而且是与它一同跻身高位的模型里最便宜的那个,遥遥领先。 为什么一个偏开放的 Kimi 能在智能体工作上领先两款闭源前沿模型,却在综合指数上与它们持平?因为智能体评测奖励的是规划、工具调用,以及在许多步骤里都不偏题的能力,而 K3 默认的最大思考预算加上它的 token 效率,恰恰是为此调优的。GPT-5.5 和 Opus 4.8 是强大的通用模型,但两者都不像这一代 Kimi 那样,是为了登顶自主智能体排行榜而打造的。如果你的用例是单次提示,那该信的读法是综合指数(三者在那里相当)。如果是一个跑很多轮的智能体,那就是 GDPval,而 K3 赢下它。 [...] `$5 / $30` `$2.50 / $15` `$1 / $6` TL;DR。 Artificial Analysis 把 Kimi K3 的综合智能水平与 GPT-5.5 和 Claude Opus 4.8 放在同一档,而在它的 GDPval v2 智能体评测上,K3(1668 Elo)实际得分高于 Opus 4.8(1600)和 GPT-5.5(1494)。只有 Claude Fable 5 和 GPT-5.6 Sol 明

Kimi K3 vs GPT-5.6 Sol vs Claude Fable 5 - AI Model Comparisonopenrouter.ai · supporting

Variant GPT-5.6 Sol (max) Claude Fable 5 Variant Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) #### Intelligence #### Coding #### Agentic ## Design Arena Table Graph Kimi K3 3D 1456 99% GPT-5.6 Sol ### No benchmarks available Claude Fable 5 3D 1373 98% Kimi K3 Asciiart — GPT-5.6 Sol ### No benchmarks available Claude Fable 5 Asciiart 1364 99% Kimi K3 Code Categories 1417 99% GPT-5.6 Sol ### No benchmarks available Claude Fable 5 Code Categories 1340 97% Kimi K3 Data Visualization 1380 98% GPT-5.6 Sol ### No benchmarks available Claude Fable 5 Data Visualization 1354 97% Kimi K3 Game Development — GPT-5.6 Sol ### No benchmarks available Claude Fable 5 Game Development 1393 99% Kimi K3 SVG — GPT-5.6 Sol [...] Weighted Average Input $1.698/ M tokens Claude Fable 5 Weighted Average Input $3.079/ M tokens Kimi K3 Cache Write — GPT-5.6 Sol Cache Write from$6.25/M tokens≤272K $6.25, >272K $12.50 Claude Fable

Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while ...the-decoder.com · supporting

Kimi is launching K3, a multimodal model with 2.8 trillion parameters and a context window of one million tokens. In the company's own benchmarks, it performs on par with leading proprietary models. According to Kimi, the new flagship model K3 has 2.8 trillion total parameters, processes images and video natively, and supports a context window of one million tokens. Kimi calls K3 the first open model in the roughly 3 trillion parameter range. Full model weights are scheduled for release by July 27. The model targets long-running programming tasks, knowledge work, and complex reasoning. [...] ## Key Points Kimi has released K3, a multimodal open-weight model built on a mixture-of-experts architecture with 896 experts, 2.8 trillion parameters, and a context window of one million tokens. Full weights are expected by the end of July. In Kimi's own benchmarks, K3 comes close to Claude Fable 5 and GPT 5.6 Sol but beats all other tested systems by a wide margin. Independent testing by Art

Kimi K3: An Open Source #1 — Coding Tests vs Fable 5, Opus 4.8 & GPT-5.6 Solyoutube.com · supporting

at the top of the leaderboard. And for the first time, an open source model, Kimi K3, tops this one right here. And it comes above Fable 5, which is to me incredible because Fable was just about to be discontinued because it was so good, it was dangerous by the US government. So amazing, amazing news that we have this open source model right here on top. And if you want to see other rankings right here from Arena AI, you also have HTML, you have React. On React, you also have Kimi K3 on top. So for React use cases, if you're building React applications, also you probably want to use Kimi. And also for other domains such as branding, reference-based design, data and analytics, consumer product etc in all of this Kimi K3 achieves the top performing results right here so to me that is just [...] days, K3 is only available on maximum thinking. But they will be rolling out lower levels of thinking in the next few days. That is for Pi. Tau is my favorite harness because of course it's my bab