公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
product

GPT-5.6 Sol 优化自身推理效率,服务成本降低 20%

OpenAI 表示,GPT-5.6 Sol 参与优化了自身的生产内核和推测解码流程,使端到端服务成本降低 20%,令牌生成效率提升超过 15%。

97% VERIFIED

OpenAI 表示,GPT-5.6 Sol 在人工主导的流程中自主重写并优化了生产 GPU 内核,同时设计和运行了数百项实验,以提升模型的令牌生成效率。相关内核改进使模型端到端服务成本降低了 20%。

该模型还改进了推测解码流程,使令牌生成效率提升超过 15%。OpenAI 称,这类模型、推理系统和工具链之间的反馈机制将持续帮助其推进 AI 的效率边界。

来源证据

OpenAI autonomously improved the inference efficiency of GPT-5.6 using GPT-5.6 itself. - GIGAZINEgigazine.net · supporting

According to OpenAI, the high cost-effectiveness of GPT-5.6 Sol is achieved through self-improvement using GPT-5.6 Sol itself. In inference, the inference engine kernel itself was improved with GPT-5.6 Sol. GPT-5.6 Sol has learned syntax that contributes to the efficiency of ' Triton ' and ' Gluon, ' the programming languages used in OpenAI's kernel, and it is said that kernel optimization has succeeded in reducing end-to-end service costs by as much as 20%. Improvements to speculative decoding were also implemented autonomously. Speculative decoding is an inference acceleration technique that uses a small 'draft model' to predict the next token, and token generation efficiency has been improved by 15% by improving the draft model using GPT-5.6 Sol.

Advancing the price-performance frontier with GPT-5.6openai.com · supporting

GPT‑5.6 Sol is increasingly helping us find and deliver the next round of gains. Within a human-led process, Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training, intervening when problems arose. The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. This work continues, creating a tighter feedback loop: as our models improve and are able to work more autonomously, our ability to improve efficiencies accelerates. Read more⁠ about the engineering behind GPT‑5.6. ## A compute strategy built for scale [...] ## How we advance the efficiency frontier Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware product

How GPT-5.6 fuses frontier intelligence with frontier efficiencyopenai.com · supporting

by OpenAI. These efforts, combined with broader kernel advancements from GPT‑5.6 Sol, reduced end-to-end serving costs by 20%. We’ve also heavily invested in verification tooling, such as the open-source tool FpSan⁠(opens in a new window) (Floating-Point Sanitizer), to help validate the correctness of the kernels written by GPT‑5.6 Sol. [...] training instability. The resulting improvements increased token-generation efficiency by more than 15%. [...] We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-token predictions. Even when individual operations are fast, excess memory movement, synchronization, and inefficient data layouts can leave GPUs idle. To avoid this, GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model. This worked in part becaus

OpenAI used GPT-5.6 Sol to make it efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding. https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/threads.com · supporting

# Thread 352 views sung.kim.mw's profile picture sung.kim.mw OpenAI used GPT-5.6 Sol to make it efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. - 15%+ better token-generation efficiency from improved speculative decoding. openai.com/index… How GPT-5.6 fuses frontier intelligence with frontier efficiency openai.com How GPT-5.6 fuses frontier intelligence with frontier efficiency 5 1 Log in or sign up for ThreadsSee what people are talking about and join the conversation. Log in with username instead

GPT-5.6 just made itself better...youtube.com · supporting

a fraction of the price, and it's slightly ahead. Here's Claude Opus 5 low, coming in slightly lower than GLM 5.2, slightly lower than 5.6 Luna, but it is really expensive at about 40 cents per task completion. So, Luna is just going to be an absolute workhorse of a model. And all of that is great. But, that's not the crazy part. The crazy part is they used their Frontier model GPT-5.6 Soul to find all of these performance gains. That allowed them to drop the price of their other models. So, just yesterday after deployment, we applied GPT-5.6 Soul to advance the frontier of efficiency by making itself more efficient to run. The results: 20% lower serving costs from production GPU kernel improvements, 15% better token generation efficiency from improved speculative decoding. And if you [...] with the marketplace, it's basically like an App Store, but for agents. And of course, you can find my loopy skill in the marketplace for free. So go find my skill on the marketplace, link down belo

GPT-5.6 Sol Just Optimized Itself & Anthropic and OpenAI Want to SLOW AI Development!youtube.com · supporting

a lot of labs do this. We're seeing that more models are getting smarter at doing this. But, what they're saying is that instead of just focusing on creating a smarter model, OpenAI focused on creating an efficient model, and there's a reason for this. As AI deployment gets more and more broader and more and more people are adopting AI, efficiency becomes a critical thing for most labs because now all of these labs have to serve their strongest models, which are way more energy demanding compared to previous models to a broader audience, so efficiency becomes important. And through applying GPT 5.6 soul to make more efficiency improvements, The model is now 20% lower when it comes to serving cost from production GPU kernel improvements. So, that is a big number for a big company like [...] So, that is a big number for a big company like OpenAI. 20% lower servicing cost basically means millions or I don't know, it could be in the billions. I don't think so, probably millions, hundreds o