8/6/2026, 1:04:26 PM · foundation-models

DeepSeek Launches V4-Flash-0731 as Official Fast-Inference Model, Prioritizing Post-Training Over Architecture

DeepSeek's July 31 release graduates its compact Flash model from preview to general availability via re-post-training alone, achieving agentic benchmark gains without altering the underlying architecture.

Overview

<cite index="9-1">On July 31, 2026, DeepSeek officially announced DeepSeek-V4-Flash-0731, the production release of its Flash model.</cite> The release advances the competitive large language model (LLM) landscape by demonstrating that performance improvements can be achieved without redesigning a model's core architecture—a signal with implications for the entire open-weight model ecosystem.

Architecture and Specifications

<cite index="9-5">The model retains 284 billion total parameters, activates only 13 billion parameters per token using a Mixture-of-Experts (MoE) architecture, and supports a 1 million token context window.</cite> <cite index="15-1,15-4,15-5">DeepSeek-V4-Flash-0731 is the official release superseding the preview version, and it comes with a speculative decoding module attached.</cite> Including this DSpark speculative decoding module, <cite index="11-6">Hugging Face reports 304 billion parameters for the repository in total.</cite>

What Changed: Post-Training, Not Architecture

<cite index="9-7,9-8">Instead of investing in a larger model, DeepSeek focused entirely on post-training. According to the company, DeepSeek-V4-Flash-0731 is simply the April Preview model retrained using a significantly improved post-training pipeline focused on coding, AI agents, reasoning, and tool use.</cite> <cite index="18-1">DeepSeek confirmed via its official X account that V4-Flash-0731 keeps the exact same model architecture and size as the preview version.</cite>

<cite index="9-9">The result is a model that reportedly delivers substantially better performance while maintaining the same architecture, application programming interface (API) endpoint, latency profile, and pricing.</cite> <cite index="9-15">For developers, this means applications already using `deepseek-v4-flash` automatically benefit from the new version without changing API calls.</cite>

Benchmark Performance

<cite index="15-6">DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed in its model card despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.</cite> On specific vendor-published evaluations, <cite index="14-7,14-8,14-9">Terminal Bench 2.1 rises to 82.7, DeepSWE reaches 54.4, and Toolathlon Verified lands at 70.3.</cite> <cite index="14-11">An independent evaluation from Artificial Analysis also reports a roughly ten-point jump on its Intelligence Index.</cite>

<cite index="14-13">DeepSeek's broader V4 technical report describes a two-stage post-training process for the family: domain-specific experts are cultivated with supervised fine-tuning and reinforcement learning, then consolidated through on-policy distillation.</cite> However, <cite index="14-14">the company has not yet published an equivalent ablation for 0731, so the specific data or reward changes that produced the new agent behavior remain undisclosed.</cite>

Distribution and API Changes

<cite index="11-1">DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta on July 31, 2026.</cite> <cite index="10-7">The weights were released MIT-licensed on the same day, superseding the preview checkpoint.</cite> <cite index="11-7">On the API side, `deepseek-v4-flash` now natively supports the Responses API format and is adapted for Codex.</cite>

<cite index="18-2,18-3">The July 31 upgrade applies only to the DeepSeek-V4-Flash API; the DeepSeek-V4-Pro API and App/Web models remain unchanged.</cite> <cite index="18-7">DeepSeek stated that the official release of DeepSeek-V4-Pro is "coming ASAP."</cite>

Industry Context

<cite index="13-7,13-8">The release carries a notable sequencing detail: DeepSeek shipped the official version of its smaller model before the 1.6-trillion-parameter V4-Pro flagship, and its changelog markets the small model's agent-benchmark scores as exceeding the larger one's on DeepSeek's own evaluation harness.</cite> <cite index="14-15">The release is consistent with a pattern visible across the industry: once a base model is capable enough, the quality of tool trajectories and feedback can matter as much as adding more raw capacity.</cite>

The V4-Flash-0731 launch underscores a broader shift in open-weight model development—where inference efficiency, agentic performance, and post-training methodology increasingly differentiate competitors more than raw parameter counts alone.

Cross-references

Sources

  1. [1]
    DeepSeek-R2: China's Powerful New AI Model for 2025
  2. [2]
    The Complete Guide to DeepSeek Models: V3, R1, V4 and Beyond
  3. [3]
    DeepSeek V4 Preview: What the Fast, Expert, and Vision Modes Suggest
  4. [4]
    Cerebras Launches World Fastest DeepSeek R1 Llama-70B Inference - Cerebras
  5. [5]
    Build with DeepSeek V4 Using NVIDIA Blackwell and GPU-Accelerated Endpoints | NVIDIA Technical Blog
  6. [6]
    DeepSeek-R1 Overview: Features, Capabilities, Parameters
  7. [7]
    Models | Machine Learning Inference | DeepInfra
  8. [8]
    DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles - LMSYS Org
  9. [9]
    New DeepSeek-V4-Flash-0731 is Here !! | by Mehul Gupta | Data Science in Your Pocket | Jul, 2026 | Medium
  10. [10]
    DeepSeek V4 Flash Is Now Official: What Changed in the 0731 Build
  11. [11]
    DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains - MarkTechPost
  12. [12]
    deepseek-ai/DeepSeek-V4-Flash-0731
  13. [13]
    DeepSeek V4 Flash 0731: Official Release, Agent Benchmarks
  14. [14]
    DeepSeek V4 Flash 0731: The Update That Repriced Coding… | NxCode
  15. [15]
    deepseek-ai/DeepSeek-V4-Flash-0731 · Hugging Face
  16. [16]
    DeepSeek (chatbot)
  17. [17]
    Change Log | DeepSeek API Docs
  18. [18]
    DeepSeek on X: "⚠️ Note 🔷 DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version. 🔷 Today's upgrade applies ONLY to the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official release of DeepSeek-V4-Pro" / X