Overview
Chinese artificial intelligence (AI) startup DeepSeek on August 13, 2026, formally released the general-availability build of its V4 Pro large language model (LLM) and announced sweeping application programming interface (API) price increases effective August 17—capping a week that marked one of the company's most significant strategic pivots since its R1 model went viral in early 2025.
The Model: Architecture and Benchmark Position
<cite index="2-7">The build, designated V4-Pro-0813, went live on the app, web, and API platforms with what DeepSeek describes as substantially upgraded agent capabilities.</cite> <cite index="9-1">V4 Pro is a Mixture-of-Experts (MoE) architecture with 1.6 trillion total parameters and 49 billion activated per token, supporting a one-million-token context window.</cite> <cite index="9-3">In the one-million-token context setting, V4 Pro requires only 27% of single-token inference floating-point operations (FLOPs) and 10% of key-value (KV) cache compared with DeepSeek-V3.2.</cite>
<cite index="23-1">DeepSeek V4 Pro 0813 (Reasoning, Max Effort) scores 53 on the Artificial Analysis Intelligence Index, placing it well above average among other open-weight models of similar size, whose median score is 27.</cite> <cite index="27-1">That score is up from 45 for the April V4 Pro preview release, representing a meaningful capability gain.</cite> <cite index="22-4">It places V4 Pro just one point ahead of DeepSeek V4 Flash 0731, the company's cheaper, faster model that had briefly matched or beaten Pro on the same index.</cite>
<cite index="24-6">The 53 score is lower than Moonshot's Kimi K3 at 60 and Anthropic's Claude Opus 5 at 63.</cite> <cite index="22-11">On GDPval-AA v2, Artificial Analysis's test for agentic real-world work tasks, DeepSeek V4 Pro scores 55%, ahead of GPT-5.6 Luna's 53% and Gemini 3.6 Flash's 50%, though still behind Claude Opus 5's 67% and Grok 4.6's 62%.</cite> <cite index="24-3">DeepSeek reports that V4 Pro scored 62.7 on the DeepSWE benchmark for real-world software engineering, a notable increase from 12.8 in earlier testing.</cite>
API Pricing: Structure and Magnitude
<cite index="11-2,11-3">DeepSeek will raise API prices for its V4-Pro and V4-Flash models and introduce peak and off-peak pricing; the new rates—ranging from 50% to 1,100% above current prices depending on the model, token type, and time of use—take effect on August 17, according to a company statement.</cite>
<cite index="18-1,18-2">The V4 Pro model will cost $3.96 per one million output tokens at peak hours, more than four times the previous rate of $0.87, with off-peak hours priced at $1.98.</cite> <cite index="18-6">The company stated it is adopting the new peak and off-peak structure to "allocate resources more reasonably."</cite>
<cite index="24-9">For comparison, Moonshot's Kimi K3 costs between $14 and $15 per million output tokens, and competitors such as Anthropic's Claude Fable 5 can reach $50 per million.</cite> Even at the new rates, DeepSeek's pricing remains well below frontier closed-model alternatives.
Commercial Context
<cite index="15-7,15-8">Ultra-low pricing attracted a large number of developers and enterprise users, leading to a surge in requests that exceeded platform capacity and caused frequent instability; the price adjustment is widely interpreted as a necessary measure to balance resource pressure and service quality.</cite>
<cite index="13-7">The August move marks another strategic shift following the introduction of peak/off-peak pricing in mid-July.</cite> <cite index="1-5">DeepSeek's early lead has since been challenged by a string of releases from Chinese competitors including Moonshot AI, Zhipu AI, MiniMax, Alibaba, and ByteDance.</cite>
<cite index="1-7">Reuters reported in July that the company was planning a new fundraising round at a valuation of approximately $74 billion, weeks after raising about $7.4 billion in its first outside financing round.</cite> <cite index="4-10">DeepSeek has also said it aims to at least double staffing across departments, including data-center and AI-agent teams.</cite>
Open Weights and Ecosystem Access
<cite index="8-4,8-5">The model supports both OpenAI ChatCompletions and Anthropic APIs, and both V4 Pro and V4 Flash support one-million-token context windows along with dual Thinking and Non-Thinking modes.</cite> The weights for the V4 Pro family remain publicly accessible on Hugging Face, preserving the open-model positioning that has driven developer adoption—even as the hosted API shifts to a more commercial rate card.