Opus 5 用户评测:多数任务胜过 Fable 5,但 GPT-5.6 Sol 仍有长程优势
一位用户称其在一天内使用 Opus 5 约28亿 Token,并认为该模型在多数能力上超过调整后的 Fable 5。不过,他同时指出,GPT-5.6 Sol 在复杂后端逻辑和长程任务上更强。现有资料确认 Opus 5 强调效率和较低 Token 消耗,但不同来源对其与 GPT-5.6 Sol、Fable 5 的基准排名存在明显分歧。
这份早期体验认为,Opus 5 更像一名可以随时协作、提出方案的同事,在日常编码和通用任务中具备较强表现。多条社区反馈也提到,它完成相近任务时消耗的 Token 明显少于 Fable 5;部分二手资料称其价格约为后者的一半。
不过,Opus 5 是否全面领先仍无法确认。部分报道认为它在编码测试中大致追平或略胜 Fable 5,并领先 GPT-5.6 Sol;OpenAI 发布的官方资料则称 GPT-5.6 Sol 在编码及长程工程任务上领先 Fable 5。28亿 Token 的个人使用量和相关主观结论,仍应视为未经独立验证的用户观察。
来源证据
Anthropic's Opus 5 is about token efficiency, not a capability leap - Ars Technicaarstechnica.com · supportingA chart (made by Anthropic) with various benchmarks like Frontier-Bench and DeepSWE shows Opus 5 performing at about the same level or slightly ahead of the much-ballyhooed capabilities of Anthropic’s Fable model for coding tasks. It ostensibly beats Opus 4.8 and OpenAI’s competing GPT-5.6-Sol in just about every kind of task. The various benchmarks show an iterative increase in performance, but not a radical leap, while the main pitch is that it’s a model that offers something just shy of Fable at approximately half the cost. [...] Lately, companies like Cursor and Meta have been building “model routers,” systems that automatically select models of varying size and capability from an array of options based on the nature of the prompt. The idea is that you save a lot of tokens (and therefore compute and/or money) by not using something like Fable for every task. Companies like Anthropic are going to have to keep bringing token costs down, or at least (as is the case here) offering mo
Mediummedium.com · supportingSign up Sign in Sign up Sign in Unknown user ## No Time No Time Share through stories. Member-only story ## Claude Opus 5 review | Opus 5 vs Fable 5 | Claude Opus 5 benchmarks # Anthropic Just Launched Opus 5 (Here Is Everything You Need to Know) ## Smaller, Cheaper, But as Powerful as Fable 5? Pranit naik -- 1 Listen Share Read here for FREE On Frontier-Bench v0.1, a coding test where a model has to read a real repository, plan, edit files, run tests and recover from its own mistakes, Claude Opus 5 scored 43.3 percent. Claude Fable 5, the model Anthropic had spent six weeks describing as its most capable, scored 33.7 percent on the same test. Opus 5 costs $5 per million input tokens and $25 per million output. Fable 5 costs $10 and $50. [...] If you run a coding agent that chews through tokens all day, that is not an abstract benchmark question. It is a line item. Opus 5 went live on July 24, 2026. It is now the default model on Claude Max, the strongest model o
GPT-5.6: Frontier intelligence that scales with your ambitionopenai.com · supporting## Efficient by default, maximum performance on demand GPT‑5.6 Sol is our best coding model yet. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less. That advantage extends across the family: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost. It also sets new state-of-the-art results on Terminal‑Bench 2.1 and DeepSWE, which test complex command-line workflows and long-horizon engineering in real codebases. [...] general capabilities, GPT‑5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost. [...] 7, 8 ### Professional EvalGPT‑5.6 SolGPT‑5.6 TerraGPT‑5.6
Fable 5 Consumes Excessive Tokens on Opus | Reggie Chan, FRICS, CFA posted on the topic | LinkedInlinkedin.com · supportingI've been using Fable 5 for about a week now. As impressive as it is, I must admit I am disappointed. Some key insights after a week of use: 1. This is not the Fable 5 we had before. It looks like the changes that Anthropic made to the model degraded it noticeably. 2. Using any Effort setting below Extra seems to be a waste of tokens and almost guarantees that you will have to send several prompts for any medium complexity task. 3. Codex 5.5 is still be the best value by a long shot. Codex set to Extra High seems to out perform Opus on Ultracode and is a much more efficient use of tokens. 4. Fable 5's token usage is absurd especially considering that all token usage has been halved this last week. If anything, this makes me more excited to see what Codex 5.6 will look like. [...] This better turn out to be an ungodly plan if it ever gets finished. [...] A lot of this has been built previously in a bunch of third-party frameworks around Claude Code. All of them feel like forcing an unwi
Opus 5 Token Usage is Amazingreddit.com · supporting## Top Comments [...] # Opus 5 Token Usage is Amazing ## Post details r/ClaudeAI | 310 upvotes | 102 comments | Posted by u/Meme_Theory on 2026-07-25T14:00:40.562Z ## Content I am literally TRYING to use up my last 10% before my reset tonight, and I feel like the token counter is barely ticking. Using Fable for the same task took substantial amounts of my Max x20 usage. And I am very pleased with the results! (max effort) Just figured we could use a counterbalance to the flood of Opus 5 is terrible posts that have predictably already started. What have you liked in the last day with Opus 5? Edit: 3-hour update; 95%, I should slide right into the finish line - no tokens wasted, like a good Boy Scout. [...] - TL;DR of the discussion generated automatically after 80 comments. The overwhelming consensus in this thread is that Opus 5 is absurdly token-efficient. Many users, especially those on Max plans, are reporting that they're actively trying to burn through their usage limits and the
Fable 5 vs GPT 5.6 Sol: The Early Resultsyoutube.com · supportingland. Thank you so much for watching and have a wonderful day. [...] 422 comments ### Transcript: [...] of its kind. The world, in immediate response, rallied in sympathy with Anthropic, who have always been champions of never using even a line of copyrighted material for any of their models. And sarcasm. But wait, how does this story link to the corporate concentration point I was just making? Well, if this large-scale scraping to distill abilities into Chinese models becomes ever more sophisticated, successful, I think the incentives of the labs might switch. Better for them perhaps to serve their latest models to governments, approved businesses, and of course themselves for say 3 to 4 months, safe from this kind of distillation, and then only when they have a better internal model release the older one to you, the unwashed masses. And that theory is even before you get to geopolitics.