DeepSeek Launches V4 Pro Production Build with Major Agent Upgrades
DeepSeek's production release of V4 Pro brings major agentic workflow improvements, flexible reasoning effort levels, and native OpenAI Responses API support. The model is now available in the app and web via Expert Mode, as well as through the API with unchanged model names.
DeepSeek today announced the production launch of V4 Pro, describing it as a major upgrade for agent workflows with strong production gains. Users can access the model immediately through Expert Mode on the app and web, while developers can connect via the API; DeepSeek emphasized that model names are unchanged and pointed developers to the API docs for setup.
The release adds configurable reasoning effort for both V4 Pro and V4 Flash: low is intended for simple tasks, high for everyday agent workflows, and max for complex assignments. DeepSeek also introduced native OpenAI Responses API support, optimized for Codex with one-click configuration.
Independent coverage places V4 Pro as DeepSeek's largest model to date at 1.6 trillion total parameters with 49 billion active, leading open-weight agentic benchmarks, while V4 Flash is a smaller, efficiency-focused variant. Reports note that an April release was a preview; today's announcement is the production build.
Source evidence
DeepSeek is back among the leading open weights models with V4 Pro and V4 Flashartificialanalysis.ai · supportingDeepSeek V4 Pro scales DeepSeek’s architecture substantially, while V4 Flash is positioned for size efficiency: V4 Pro is DeepSeek’s largest model to date at 1.6T total parameters / 49B active, a major step up from the V3 family’s 671B total / 37B active architecture. V4 Flash is far smaller at 284B total / 13B active, but sits strongly on the Intelligence vs Size frontier, near MiniMax-M2.7. DeepSeek V4 Pro leads open weights models on GDPval-AA, our agentic real-world work tasks benchmark. V4 Pro (Max) scores 1554, ahead of Kimi K2.6 (1484), GLM-5.1 (1535), GLM-5 (1402), and MiniMax-M2.7 (1514). V4 Flash (Reasoning, Max Effort) scores 1388. [...] DeepSeek has released DeepSeek V4 Pro and V4 Flash. V4 is the first new architecture from DeepSeek since V3. V4 introduces a new architecture with V4 Pro at 1.6T total / 49B active parameters and V4 Flash at 284B total / 13B active parameters, and is DeepSeek's first two-tier lineup, with Pro positioned for maximum capability and Flash for
DeepSeek V4 Pro: Model Overview, Features & ...deepinfra.com · supportingDeepSeek V4 Pro is a 1.6-trillion parameter Mixture-of-Experts (MoE) model from DeepSeek, released on April 24, 2026 under the MIT license. It is designed for advanced reasoning, complex software engineering, and long-running agentic tasks, and arrives alongside DeepSeek-V4-Flash, a lighter 284B-parameter variant built for faster, lower-cost inference. The V4 series is DeepSeek’s first two-tier lineup and introduces a new architecture — the first from the lab since V3. Both models are hybrid thinking/non-thinking and support a 1 million token context window. ## Architectural Innovations The V4 series is built on several technical advances over DeepSeek-V3.2: [...] ## Getting Started with the API DeepSeek-V4-Pro is available for immediate integration via the DeepInfra platform under the model identifier deepseek-ai/DeepSeek-V4-Pro. Access the model at deepinfra.com/deepseek-ai/DeepSeek-V4-Pro. Reasoning Modes A key feature of DeepSeek V4 is configurable reasoning depth. Developers
DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major ...marktechpost.com · supporting## Serving it DSpark is enabled with one vLLM flag: `--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'`. The DSpark paper reports 60–85% faster per-user generation on V4-Flash versus the MTP-1 baseline at matched aggregate throughput. There is no Jinja chat template. DeepSeek ships an `encoding/` folder with `encode_messages` and `parse_message_from_completion_text` instead. `reasoning_effort` takes `low`, `high`, or `max`. DeepSeek recommends `temperature = 1.0`, `top_p = 0.95` for agentic use and `1.0` otherwise, with up to 384K output tokens at `high` and `max`. ## Key Takeaways [...] | Benchmark | V4-Flash-0731 | V4-Flash (Preview) | V4-Pro (Preview) | GLM-5.2 | Opus-4.8 | --- --- --- | | Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 | | NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 | | Cybergym | 76.7 | 38.7 | 52.7 | — | 83.1 | | DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 | | Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59
DeepSeek V4 Pro vs Flash: What Launched, What Changed, and the Huawei Chipremio.ai · supportingBoth models support thinking mode (with a reasoning\_effort parameter accepting high or max) and non-thinking mode. For complex agent workflows, DeepSeek recommends thinking mode at max intensity. API model names: deepseek-v4-pro and deepseek-v4-flash. The old model names deepseek-chat and deepseek-reasoner currently map to V4-Flash non-thinking and V4-Flash thinking mode respectively, and will be deprecated on July 24, 2026. DeepSeek has explicitly optimized V4 for integration with major agent frameworks including Claude Code, OpenClaw, OpenCode, and CodeBuddy. Document and code generation tasks are noted as areas of meaningful improvement. ## How to Access DeepSeek V4 Right Now [...] V4-Pro is the flagship model optimized for maximum capability: complex reasoning, agentic coding, and tasks where quality matters more than speed. V4-Flash is smaller and faster with lower API cost, matching V4-Pro on simple tasks and approaching it on most reasoning tasks. For high-volume or latency-
DeepSeek V4 Pro 0813 with Major Agent Upgradeyoutube.com · supportingDeepseek is not stopping. They have just released the production build of their V4 Pro model and the Agentic benchmarks they have jumped hard. Terminal Bench 2.1 went from 72 to 88. Cyber Gym from 53 to 83. The coding and agentic race is heating up fast. Quen 3.8 Max 2.4 4 trillion model is also out on hugging face. In this video we are going to check out this new update from deepseek. We will also be covering lot of other updates which have dropped today. This is Fad Miza and I welcome you to the channel. So what I'm going to do I'm going to use this model again with Hermes agent and we are going to give it a very complex real world task. Let me quickly launch the Hermes agent. And the task which I'm going to give this model is this broken buggy fullstack application. So there is a back [...] has done well. Let me know your thoughts in the comments. I will be covering more models because there's a lot to cover today. Thank you for all the support. [...] # DeepSeek V4 Pro 0813 with Maj
deepseek/deepseek-v4-prozenmux.ai · supporting\\Only the DeepSeek provider has the official 0731 version; other providers are still using the old 0424 version\\.DeepSeek-V4-Flash is the efficiency-oriented variant of the DeepSeek V4 series, released as a preview and open-sourced alongside the flagship V4-Pro. It is designed for developers who need the V4 generation's long-context and reasoning capability at a faster, more economical API tier. Compared with V4-Pro, V4-Flash uses smaller total parameters and active parameters, resulting in faster response times and lower API cost. It retains reasoning capability close to V4-Pro and matches V4-Pro on simple agent tasks, with a measurable gap appearing only on the most demanding agent workflows. World knowledge is slightly below V4-Pro but remains competitive within the open-source [...] \\Only the DeepSeek provider has the official 0731 version; other providers are still using the old 0424 version\\ . DeepSeek-V4-Flash is the efficiency-oriented variant of the DeepSeek V4 series, rel