7/24/2026, 1:02:29 PM · foundation-models

Google Launches Gemini 3.6 Flash at $1.50/$7.50 Per Million Tokens, Pressing Open-Weight Rivals

Google's surprise July 21 release of Gemini 3.6 Flash cuts output-token pricing and improves agentic efficiency just as open-weight challengers DeepSeek V4 and Kimi K3 intensify competitive pressure on the Large Language Model market.

Google Releases Gemini 3.6 Flash With Lower Output Pricing

<cite index="21-1">Gemini 3.6 Flash launched on July 21, 2026, priced at $1.50 per million input tokens and $7.50 per million output tokens, with a 1-million-token context window and day-one availability in Google AI Studio, the Gemini Application Programming Interface (API), and the Gemini app.</cite> <cite index="21-3">The model surfaced publicly after its internal identifier first appeared inside the Google Antigravity Integrated Development Environment (IDE) earlier the same day.</cite>

<cite index="21-4">Google confirmed the release with an official blog post from the Gemini team, a model card, and simultaneous availability across AI Studio, the Gemini API, Android Studio, Antigravity, and Vertex AI.</cite>

Pricing and Efficiency Improvements

<cite index="22-1,22-2">Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor, Gemini 3.5 Flash, and cuts the output price from $9.00 to $7.50 per million tokens, while holding the input price steady at $1.50 per million tokens.</cite> <cite index="29-3">Combined, those two factors mean the per-task bill for output-heavy workloads falls by roughly one third.</cite>

<cite index="25-2">Official model cards released by Google confirm that Gemini 3.6 Flash features a 1-million-token input context window alongside a maximum output limit of 64,000 tokens, with a knowledge cutoff date of March 2026.</cite> The knowledge cutoff represents a fourteen-month leap forward from the January 2025 cutoff carried by Gemini 3.5 Flash.

Alongside Gemini 3.6 Flash, <cite index="26-7,26-8">Google simultaneously shipped Gemini 3.5 Flash-Lite and a security-tuned Gemini 3.5 Flash Cyber.</cite> <cite index="23-11,23-12">Flash-Lite is designed for latency-sensitive applications such as agentic search, document processing, and translation, priced at $0.30 per million input tokens and $2.50 per million output tokens.</cite> <cite index="28-6,28-7">Flash Cyber, meanwhile, powers a multi-agent vulnerability scanning tool called CodeMender and is gated to governments and trusted partners under a limited-access pilot, reflecting dual-use security concerns.</cite>

Benchmark Performance

<cite index="25-4,25-5">Gemini 3.6 Flash scores 49% on the DeepSWE coding benchmark, up from 37% achieved by version 3.5 Flash, and pushes machine-learning engineering performance to 63.9% on MLE-Bench compared to 49.7% previously.</cite> <cite index="25-6">Google also integrates computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, with an OSWorld-Verified score of 83.0%, up from 78.4%.</cite> <cite index="26-17,26-18,26-19">The model outscores Google's own Gemini 3.1 Pro on the Artificial Analysis Intelligence Index, 50 to 46, though independent testing finds it equally intelligent to Gemini 3.5 Flash — faster and cheaper, but not meaningfully smarter.</cite>

Competitive Context: Open-Weight Pressure

The release arrives as open-weight Large Language Models (LLMs) from China press hard on the frontier. <cite index="12-3">On April 24, 2026, DeepSeek released DeepSeek V4 and V4-Pro.</cite> <cite index="17-1">Reuters reported that the V4 launch was treated as a preview phase, and DeepSeek did not provide a timeline for finalization.</cite> <cite index="15-14">Given DeepSeek's history with V3, V4 pricing is expected to land well below the $5/$30 per million tokens charged by comparable closed models from Western competitors.</cite>

Moonshot AI's Kimi K3 poses an additional challenge. <cite index="30-5,30-6">Released July 16, 2026, Kimi K3 is a 2.8-trillion-parameter open-weight, multimodal reasoning model — the largest open-weight model shipped to date — with a 1-million-token context window and API pricing of $3 per million input tokens and $15 per million output tokens.</cite> <cite index="33-1,33-2,33-3">The full model weights are scheduled for public release by July 27, 2026; further architectural and training details will accompany a forthcoming technical report, and Kimi K3 is the first open model to reach 2.8 trillion parameters.</cite>

Flagship Delay Adds Strategic Pressure

<cite index="20-9,20-10">Despite the arrival of these Flash variants, Gemini 3.5 Pro remains absent. Google originally promised the flagship model for a June release during its Google I/O conference, but that deadline passed without a public launch.</cite> <cite index="22-8">The absence of Gemini 3.5 Pro leaves Google without a direct competitor to frontier-class models from its rivals at the top of the capability spectrum.</cite> <cite index="26-9">Google has also quietly revealed that pre-training has already begun on Gemini 4.</cite>

For developers, Gemini 3.6 Flash's combination of lower per-token output costs, a refreshed knowledge cutoff, and improved agentic benchmark scores offers a meaningful efficiency gain within Google's managed infrastructure, even as open-weight alternatives promise self-hostable alternatives at potentially lower long-run costs.

Sources

  1. [1]
    Gemini 2.5 Flash by Google — Pricing, Specs & API Access
  2. [2]
    Gemini API Pricing May 2026: 3.5 Flash, 3.1 Pro, 2.5 Lite
  3. [3]
    Gemini 2.5 Flash - API Pricing & Benchmarks | OpenRouter
  4. [4]
    Gemini Developer API pricing | Gemini API | Google AI for Developers
  5. [5]
    Gemini 2.5 Flash API Pricing 2026 - Costs, Performance & Providers
  6. [6]
    Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash ... | Hacker News
  7. [7]
    Gemini 2.5 Flash Lite API Pricing 2026 - Costs, Performance & Providers
  8. [8]
    Agent Platform Pricing | Google Cloud
  9. [9]
    Gemini API Pricing Calculator & Cost Guide (Jul 2026)
  10. [10]
    Gemini 2.5 Flash API Pricing — Token Cost Calculator | AI Pricing Guru
  11. [11]
    V4 (DeepSeek) release date - Manifold Markets
  12. [12]
    DeepSeek (chatbot)
  13. [13]
    Ray Wang on X: "DigitTimes: DeepSeek V4 expected to launch in July; R2 likely to follow in August. As the US accelerates its global AI expansion, China's tech sector is watching closely for DeepSeek's next flagship release: R2. Liang, true to form, remains low-key—reportedly sticking to https://t.co/oL4Xxt0tOF" / X
  14. [14]
    Change Log | DeepSeek API Docs
  15. [15]
    DeepSeek V4 Released: Everything You Need to Know (April 2026)
  16. [16]
    DeepSeek V4 Targets Coding Dominance with Mid-February | Introl Blog
  17. [17]
    DeepSeek V4 Preview: API Models, Pricing and Changes
  18. [18]
    c 215329073
  19. [19]
    www.mexc.com
  20. [20]
    Google Starts Surprise Gemini 3.6 Flash and 3.5 Flash Lite Rollout Amid Pro Model Delays
  21. [21]
    What Is Gemini 3.6 Flash? Pricing, Benchmarks & Availability
  22. [22]
    Google Launches Gemini 3.6 Flash With 17% Token Savings, but Flagship 3.5 Pro Remains Missing | MLQ News
  23. [23]
    Google Launches Cheaper Gemini 3.6 Flash and Flash-Lite, But Flagship 3.5 Pro Remains Delayed — BigGo Finance
  24. [24]
    Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4
  25. [25]
    Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way | VentureBeat
  26. [26]
    Gemini 3.6 Flash: Pricing, Benchmarks & What's New
  27. [27]
    Google Launches Gemini 3.6 Flash with 17% Token Reduction, Igniting 'Cost Per Task' Competition — BigGo Finance
  28. [28]
    Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads - MarkTechPost
  29. [29]
    Gemini 3.6 Flash: Pricing, Benchmarks & API Access
  30. [30]
    Kimi K3: Moonshot AI’s 2.8T Open-Weight Model — Release, Specs & Pricing (2026)
  31. [31]
    Kimi K3 - Kimi API Platform
  32. [32]
    What Is Kimi K3? Moonshot's 2.8T, 1M-Context Flagship
  33. [33]
    Kimi K3 Tech Blog: Open Frontier Intelligence
  34. [34]
    Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community
  35. [35]
    Kimi K3, and what we can still learn from the pelican benchmark
  36. [36]
    Kimi (chatbot)
  37. [37]
    Kimi K3 - API Pricing & Benchmarks | OpenRouter
  38. [38]
    Kimi K3 Status - Release Date, Official Updates and 2026 News
  39. [39]
    Kimi K3 Review: 2.8T Open Model vs Fable 5 | TECHSY