OpenAI Says GPT-5.6 Sol Cuts Serving Costs and Improves Inference Efficiency
OpenAI says GPT-5.6 Sol helped optimize the production infrastructure and inference stack used to run it. The work spans GPU kernels, routing, scheduling, caching, and model implementation; a secondary report attributes roughly 20% lower serving costs and more than 15% better token-generation efficiency to the changes.
OpenAI describes GPT-5.6 Sol as an active contributor to production-system optimization, rather than a model used only for demonstrations or evaluations. The improvements cover request routing and scheduling, GPU kernels, caching, data movement, and model implementation, allowing the same hardware to produce more useful work.
A supplied report says production GPU-kernel improvements reduced serving costs by about 20%, while techniques such as improved speculative decoding increased token-generation efficiency by more than 15%. The official excerpts confirm the broader efficiency program, but do not independently establish the full 15% figure in the material provided.
The approach reflects a broader strategy: use a more capable model to improve the systems that run it. For API customers, the potential benefits are lower cost per completed task, faster responses, and greater request capacity from the same compute budget.
Source evidence
GPT-5.6 Rewrites Kernel to Cut Service Costs by 20%eu.36kr.com · supportingWhat GPT-5.6 modified this time is not a demo running in a demonstration environment, but the production system that OpenAI is running online, which handles billions of user requests every day. In the words of Greg Brockman, President of OpenAI: Letting GPT-5.6 Sol improve the efficiency of production services is exactly one of the reasons why it is both more powerful and lower-cost. Inside OpenAI, a complete set of methodologies has long been formed, which Tibo (Thibault Sottiaux), head of Codex, summed up into two steps: The first step is to train a sufficiently powerful model; The second step is to use this model to improve everything, including the infrastructure, inference stack and kernel that runs itself. ## How was the 20% cost saving achieved? [...] But in a world where computing power is never enough and demand grows faster than production capacity, another key variable has emerged: with the same batch of GPUs, who can squeeze out more tokens? Inference efficiency is be
Advancing the price-performance frontier with GPT-5.6openai.com · supporting## How we advance the efficiency frontier Our efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT‑5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let us complete more useful work with the same compute, reducing the time, tokens, and cost required for each result. [...] In practice, businesses can define the outcome and quality standard they need, then use evaluations to determine where additional intelligence materially improves the result and where faster, lower-cost processing can deliver the same quality. A coding workflow, for example, might use Sol to resolve uncertainty and define the plan, then use Luna to implement well-specified changes, write and run tests, and eva
GPT-5.6: Frontier intelligence that scales with your ambitionopenai.com · supportingGPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost. The result is stronger performance per dollar: more successful work for the same spend, or comparable results at a lower total cost. We also introduce a new way to accelerate the most demanding work: `ultra` is our highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster. Stronger computer use and design judgment make GPT‑5.6 Sol our most polished collaborator yet, helping it inspect, refine, and deliver ready-to-use results. [...] —Angel Faus, VP of Engineering at Clio > “GPT‑5.6 delivered the best efficiency profile we’ve seen for complex financial research. In our evals, it performed at a top-tier level while being 1.72x more token-efficient, leading in three
How GPT-5.6 fuses frontier intelligence with frontier efficiencyopenai.com · supportingAchieving this requires optimizing the entire system. A model can be highly efficient in isolation, but still be expensive to serve if requests are distributed poorly, hardware sits idle, or data movement slows down computation. Improvements at every layer compound, with gains coming from optimizations in routing (where requests are sent), scheduling (when requests are sent), kernels (software that runs on GPUs), caching (saved and reused work), and model implementation (the ordering of GPU code). GPT‑5.6 Sol in Codex played an instrumental role in all of these optimizations. [...] As we’ve scaled our models to 1 billion active users and more than 2 million businesses over the past four years, efficiency has been central to distributing the benefits of intelligence to everyone. Our mission is to ensure that artificial general intelligence benefits all of humanity. Over these years, we’ve worked to continuously unlock greater optimizations across our stack in order to offer the most per
Previewing GPT-5.6 Sol: a next-generation modelopenai.com · supportingGPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount. We’re also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.
For what it's worth, I let GPT-5.6 + Claude Opus 5 factfacebook.com · supportingBy 2031, the model computation behind mathematical research quoted at $2,000 today may realistically cost somewhere between a cup of coffee and a restaurant meal—even while researchers spend millions pushing forward the new frontier. SOURCES OpenAI, “Ten advances in mathematics and theoretical computer science,” August 1, 2026. Hans Gundlach, Jayson Lynch, Matthias Mertens and Neil Thompson, “The Price of Progress: Price Performance and the Future of AI,” November 2025, revised March 2026. Stanford Institute for Human-Centered AI, 2025 AI Index Report, Research and Development. OpenAI, “How GPT-5.6 fuses frontier intelligence with frontier efficiency,” July 29, 2026. Epoch AI, “Global AI computing capacity is doubling every 7 months,” [...] At 3× per year, $2,000 falls to about $8 after five years. Hardware is improving simultaneously. Stanford’s 2025 AI Index estimated that machine-learning hardware price-performance improved by around 30% annually, while energy efficiency