Qwen3.8-27B评测达52分:小型本地模型接近部分前沿系统,但仍需谨慎解读
多来源资料显示,Qwen3.8-27B在Artificial Analysis Intelligence Index最高推理档获得52分,进入与部分近期前沿模型相近的区间。该模型据称可在约3000美元级硬件上运行,因此引发了关于本地模型能力跃升的讨论。
Qwen3.8-27B的传播焦点是其在Artificial Analysis Intelligence Index上的52分成绩。该指数综合编码、科学、推理和专业任务等多项评测;资料显示,这一成绩与GPT-5.6 Luna的最高推理档相当,并接近DeepSeek V4 Pro和GLM-5.2等更大规模模型。另有报道在不同的Agentic Index上给出51分,因此这些数字不能直接相互比较。
不过,52分对应最高推理强度,通常意味着更长的思考过程、更高的token消耗和更低的响应速度。报道还指出,Qwen3.8-27B在Terminal-Bench、HLE等高难度任务以及实际体验上仍可能落后于大型专有模型。厂商自报成绩应与独立评测结合解读,现有证据更适合说明其性价比和本地部署潜力,而不是证明它全面取代前沿模型。
来源证据
The Ultimate Guide to Qwen3.8-27B - Linas's Newsletterlinas.substack.com · supportingThis AI release is unlike anything we’ve seen before, and the data proves it. Independent third-party benchmarking from Artificial Analysis puts Qwen3.8-27B at 52 on its Intelligence Index, a composite across reasoning, knowledge, math, and coding. That score ties GPT-5.6 Luna (max) and sits one point behind GLM-5.2 (max) and DeepSeek V4 Pro (max), despite both being vastly larger mixture-of-experts systems. On Qwen’s own reported agentic-coding benchmarks, it also jumps sharply over its dense predecessor, Qwen3.6-27B: from 63.4 to 73.0 on Terminal-Bench 2.1, for instance. Of course, we should treat the vendor-reported numbers as directional and the Artificial Analysis figure as the more independent read here, but both point in the same direction.
Qwen3.8 27B scores 52 on Artificial Analysisnews.ycombinator.com · supportinghow flawed AAII is. This Qwen model is nowhere close to the other models in that score range. reply | | | | | | | --- | | | Balinares 13 days ago | prev | next (javascript:void(0)) And once again, Qwen 3.8 27B beats Opus 4.6, what the hell. It's both funny and a bit terrifying and I still can't quite believe it. It runs decently on a gaming PC! Opus 4.6 came out only 6 months ago and was then broadly considered the new SOTA by a comfortable margin! How in hell did they package capability in the ballpark of a Feb 2026 frontier SOTA into 27B?! More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago? The coming months are going to be exciting, that's for [...] this is happening. While it doesn't mention this in the Artificial Analysis page, this is likely with Max reasoning, which has extremely long reasoning traces: It seems like the token usage is 2.3x
Qwen3.8-27B就是新的模型斩杀线 - 搜狐时间线timeline.sohu.com · supporting# Qwen3.8-27B就是新的模型斩杀线 新浪财经 (来源:硅星人) 最能动手折腾模型的开源社区开发者,往往也是最会玩梗造概念的一群人。最近半年,他们最爱的一个概念是“斩杀线”。 这个词借自游戏,后来在更广泛的场合里变得更加流行。它原指血量落到某个数字以下,对手下一回合就能把你带走,比赛实际上已经提前结束。 把它套到模型上,过往一段时间人们指代的是 DeepSeek 一直在扮演的角色:一个新模型出来,先和 DeepSeek 比能力,再比价格。如果性能上不去、价格又降不下来,发布当天基本就失去了讨论价值。 从 R1 开始,DeepSeek 一度成为模型世界默认的那条线。 但现在,这条线看起来要换人了。 8月14日,Qwen3.8-27B 发布,开源两天下载量破百万,登顶 Hugging Face 全球趋势榜,编程 Agent 工具 Cline 的开发者里,它只用四天就成了被选择最多的本地模型。Artificial Analysis 的 Intelligence Index 给了它 52 分,已经进入 GPT-5.6 Luna、DeepSeek V4 Flash 这一档模型所在的区间。开发者给它起了一个更直白的外号: “本地 Opus 4.6”。 当然,这个模型还是一个进化中的模型,一个综合 benchmark 的接近,不能直接等同于真实体验完全追平 Claude Opus 4.6。在 Terminal-Bench、HLE 等高难度任务上,Qwen3.8-27B 仍然存在差距。 另外,在很多评测里,它还有一个很突出的特点:特别能“想”,很多任务会生成远多于同类模型的 reasoning token,用时间换能力——Simon Willison 实测时,默认推理档让它画一只骑自行车的鹈鹕 SVG,它想了 21 分钟。 不过这些都没有改变真正重要的地方: [...] 而且单价也未必一直往下走。8月中,DeepSeek 完成了成立以来幅度最大的一轮提价,V4-Pro 高峰时段的输出价格从每百万 token 6 元调到 27 元,同步引入峰谷计费,智谱和 Kimi 此前也已相继调价。如果你的线画在报价单上,那就注定会跟着报价单移动。 也就是说,过去两年模型公司的竞争,基本是在回答:同样的智能,谁的 token 更便宜?但Qwen3.8-27B 把问
Qwen3.8-27B runs frontier-class coding agents and ... - VentureBeatventurebeat.com · supportingOn Artificial Analysis' Agentic Index measuring model performance on agentic tasks, meanwhile, Qwen3.8-27B scored 51, beating Claude Opus 4.8 on maximum reasoning effort — a frontier model Anthropic released less than three months ago. That doesn't mean these models are equivalent, but it helps explain why developers and AI power users stood up and took notice. As developer and AI podcaster/YouTuber Sero (@0xSero on X, real name Sharif Cherf) wrote on X: "A model that runs on 3k USD of hardware is beating everything from 4 months ago. Including Opus. Permanent underclass is cancelled." [...] ## Third-party results show a powerful, local model with performance equivalent to proprietary models from months ago The conversation changed Monday when third-party results began arriving. Third-party AI benchmarking outfit Artificial Analysis gave Qwen3.8-27B a score of 52 on its Intelligence Index, a composite of nine evaluations spanning coding, science, reasoning and professional tasks. Th
Qwen3.8-27B Benchmarks: Official and Independent Resultsqubrid.com · supporting| Setting | Intelligence Index | Output tokens across the index | Peer median | Speed | --- --- | `xhigh` | 52 | 160M | 48M | Notably slow | | `medium` | 44 | 75M | 45M | Notably slow | | Non-reasoning | 35 | 26M | 17M | 53.1 t/s | Intelligence Index v4.1.1 is a nine-evaluation composite: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. The 52 at `xhigh` is the number that circulated. As Simon Willison noted when the score landed, it matched GPT-5.6 Luna at maximum reasoning and sat one point behind GLM-5.2 and DeepSeek V4 Pro at max, both vastly larger mixture-of-experts systems. [...] Qubrid AI Logo AI Models Go to Platform Qubrid AI Logo Back to Blogs & News # Qwen3.8-27B Benchmarks: Official and Independent Results 9 min read Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index at maximum reasoning effort and 61.7 on SWE-bench Pro per Qwen's own evaluation. It leads its comparis
0xSero on X: "Qwen3.8-27b hits 52 on artificial analysis A ...x.com · supportingQwen3.8-27b hits 52 on artificial analysis A model that runs on 3k USD of hardware is beating everything from 4 months ago.