OpenAI заявила о снижении стоимости обслуживания GPT-5.6 Sol и росте эффективности вывода
По данным OpenAI, GPT-5.6 Sol помог оптимизировать производственную инфраструктуру и стек вывода, на котором он работает. Изменения затронули GPU-ядра, маршрутизацию, планирование, кэширование и реализацию модели; вторичный источник сообщает примерно о 20%-ном снижении стоимости обслуживания и росте эффективности генерации токенов более чем на 15%.
OpenAI описывает GPT-5.6 Sol как участника оптимизации производственной системы, а не только как модель для демонстраций или тестов. Улучшения включают маршрутизацию и планирование запросов, GPU-ядра, кэширование, перемещение данных и реализацию модели, что позволяет получать больше полезного результата на том же оборудовании.
Согласно представленному отчету, доработка производственных GPU-ядер снизила стоимость обслуживания примерно на 20%, а улучшенная спекулятивная декодировка и другие методы повысили эффективность генерации токенов более чем на 15%. Официальные выдержки подтверждают общую программу повышения эффективности, но не позволяют независимо подтвердить весь показатель в 15%.
Такой подход основан на идее использования более мощной модели для улучшения систем, которые обеспечивают ее работу. Для пользователей API это может означать меньшую стоимость завершенного задания, более быстрые ответы и увеличение пропускной способности при прежнем вычислительном бюджете.
Источники
GPT-5.6 Rewrites Kernel to Cut Service Costs by 20%eu.36kr.com · supportingWhat GPT-5.6 modified this time is not a demo running in a demonstration environment, but the production system that OpenAI is running online, which handles billions of user requests every day. In the words of Greg Brockman, President of OpenAI: Letting GPT-5.6 Sol improve the efficiency of production services is exactly one of the reasons why it is both more powerful and lower-cost. Inside OpenAI, a complete set of methodologies has long been formed, which Tibo (Thibault Sottiaux), head of Codex, summed up into two steps: The first step is to train a sufficiently powerful model; The second step is to use this model to improve everything, including the infrastructure, inference stack and kernel that runs itself. ## How was the 20% cost saving achieved? [...] But in a world where computing power is never enough and demand grows faster than production capacity, another key variable has emerged: with the same batch of GPUs, who can squeeze out more tokens? Inference efficiency is be
Advancing the price-performance frontier with GPT-5.6openai.com · supporting## How we advance the efficiency frontier Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let us complete more useful work with the same compute, reducing the time, tokens, and cost required for each result. [...] In practice, businesses can define the outcome and quality standard they need, then use evaluations to determine where additional intelligence materially improves the result and where faster, lower-cost processing can deliver the same quality. A coding workflow, for example, might use Sol to resolve uncertainty and define the plan, then use Luna to implement well-specified changes, write and run tests, and eva
GPT-5.6: Frontier intelligence that scales with your ambitionopenai.com · supportingGPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost. The result is stronger performance per dollar: more successful work for the same spend, or comparable results at a lower total cost. We also introduce a new way to accelerate the most demanding work: `ultra` is our highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster. Stronger computer use and design judgment make GPT‑5.6 Sol our most polished collaborator yet, helping it inspect, refine, and deliver ready-to-use results. [...] —Angel Faus, VP of Engineering at Clio > “GPT‑5.6 delivered the best efficiency profile we’ve seen for complex financial research. In our evals, it performed at a top-tier level while being 1.72x more token-efficient, leading in three
How GPT-5.6 fuses frontier intelligence with frontier efficiencyopenai.com · supportingAchieving this requires optimizing the entire system. A model can be highly efficient in isolation, but still be expensive to serve if requests are distributed poorly, hardware sits idle, or data movement slows down computation. Improvements at every layer compound, with gains coming from optimizations in routing (where requests are sent), scheduling (when requests are sent), kernels (software that runs on GPUs), caching (saved and reused work), and model implementation (the ordering of GPU code). GPT‑5.6 Sol in Codex played an instrumental role in all of these optimizations. [...] As we’ve scaled our models to 1 billion active users and more than 2 million businesses over the past four years, efficiency has been central to distributing the benefits of intelligence to everyone. Our mission is to ensure that artificial general intelligence benefits all of humanity. Over these years, we’ve worked to continuously unlock greater optimizations across our stack in order to offer the most per
Previewing GPT-5.6 Sol: a next-generation modelopenai.com · supportingGPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount. We’re also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.
For what it's worth, I let GPT-5.6 + Claude Opus 5 factfacebook.com · supportingBy 2031, the model computation behind mathematical research quoted at $2,000 today may realistically cost somewhere between a cup of coffee and a restaurant meal—even while researchers spend millions pushing forward the new frontier. SOURCES OpenAI, “Ten advances in mathematics and theoretical computer science,” August 1, 2026. Hans Gundlach, Jayson Lynch, Matthias Mertens and Neil Thompson, “The Price of Progress: Price Performance and the Future of AI,” November 2025, revised March 2026. Stanford Institute for Human-Centered AI, 2025 AI Index Report, Research and Development. OpenAI, “How GPT-5.6 fuses frontier intelligence with frontier efficiency,” July 29, 2026. Epoch AI, “Global AI computing capacity is doubling every 7 months,” [...] At 3× per year, $2,000 falls to about $8 after five years. Hardware is improving simultaneously. Stanford’s 2025 AI Index estimated that machine-learning hardware price-performance improved by around 30% annually, while energy efficiency