8/17/2026, 1:04:18 PM · infrastructure

OpenAI Previews Ultrafast API Tier Running GPT-5.6 Sol at Up to 14× Standard Speed

OpenAI and Cerebras have unveiled a limited-preview inference service tier delivering up to 750 output tokens per second from GPT-5.6 Sol, signalling a strategic push toward latency-critical agentic and production workloads.

Overview

On August 13, 2026, OpenAI previewed Ultrafast, a new application programming interface (API) service tier that runs GPT‑5.6 Sol — the company's most capable large language model (LLM) — at up to 14 times the speed of its Standard processing tier. <cite index="4-3,4-4">The new tier runs GPT‑5.6 Sol up to 14× faster than Standard processing and is powered by Cerebras, generating up to 750 output tokens per second.</cite>

For reference, <cite index="6-2,6-4">Standard processing generates approximately 53 tokens per second.</cite>

Hardware Architecture

The speed differential is rooted in the underlying silicon. <cite index="1-3">Cerebras' Wafer-Scale Engine chips keep 44 GB of SRAM on-chip, avoiding the memory bandwidth bottlenecks typical graphics processing unit (GPU) setups face.</cite> <cite index="8-9">Traditional GPU-based systems must continuously shuttle massive model weights between external HBM and processing cores, while multi-GPU setups add cross-chip interconnect latency through PCIe or NVLink.</cite>

Ultrafast is not a new model. <cite index="9-4">It is a new way to serve an existing one</cite>, made possible by the compute architecture Cerebras brings to OpenAI's platform.

Tiered Inference Pricing Strategy

<cite index="3-3,3-4,3-5">OpenAI already monetizes inference speed in tiers. Through the API, it offers a "Fast Mode" that promises up to 2.5× speed with lower latency for GPT-5.6 Sol at roughly double the price. Ultrafast adds a third, faster, and likely pricier tier.</cite> <cite index="6-7">Ultrafast currently has no published pricing.</cite>

Underlying Partnership

Ultrafast is the first public product to emerge from a major infrastructure deal signed earlier in the year. <cite index="19-1">In January 2026, OpenAI signed a landmark $10+ billion agreement with Cerebras Systems to secure 750 megawatts of computing power through 2028.</cite> <cite index="22-12">"Cerebras adds a dedicated low-latency inference solution to our platform," Sachin Katti, who works on compute infrastructure at OpenAI, wrote in the blog.</cite>

Availability and Early Customers

<cite index="4-12,4-13">GPT‑5.6 Sol on Ultrafast mode is available in a limited preview to a select group of customers, with access to expand as capacity grows.</cite> <cite index="8-4">Early adopters include Jane Street, Basis, Rogo, and Podium.</cite> <cite index="14-9">OpenAI sees Ultrafast supporting live or near-production tasks including voice, customer support, commerce, developer agents, financial research, and security response.</cite>

Benchmark Performance

Cerebras published independent benchmark data alongside the announcement. <cite index="8-5">On the Humanity's Last Exam benchmark, Ultrafast completed all 2,500 questions in 11 hours and 11 minutes, compared with 78 hours and 27 minutes for Claude Fable 5.</cite> <cite index="2-1">On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6× end-to-end speedup with no quality degradation.</cite>

Context: GPT-5.6 Family

<cite index="11-1,11-2">GPT-5.6 is an LLM developed by OpenAI and released on July 9, 2026, as a family of models in three variants ranked from least to most capable: Luna, Terra, and Sol.</cite> <cite index="16-2">GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science.</cite> Ultrafast does not alter the model's weights or capabilities; it addresses only the serving layer.

OpenAI has not announced a general availability date or pricing for Ultrafast. Interested developers can register for updates through a waitlist form on OpenAI's website.

Cross-references

Sources

  1. [1]
    OpenAI's new Ultrafast mode runs GPT-5.6 Sol 14x faster using Cerebras chips | daily.dev
  2. [2]
    Accelerating GPT-5.6 Sol Ultrafast with OpenAI
  3. [3]
    GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
  4. [4]
    Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed | OpenAI
  5. [5]
    OpenAI previews 'Ultrafast' GPT-5.6 Sol running up to 14 times faster - 9to5Mac
  6. [6]
    GPT-5.6 Sol Now Runs at Real-Time Speed: OpenAI's Ultrafast Preview Offers No Price or Date
  7. [7]
    GPT-5.6 Sol Ultrafast: OpenAI Previews 14x Inference
  8. [8]
    OpenAI Launches Ultrafast Mode for GPT-5.6 Sol, Delivering Up to 14x Faster Output — BigGo Finance
  9. [9]
    OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips
  10. [10]
    OpenAI Previews Ultrafast Mode For GPT-5.6 Sol, Running Up To 14X Faster - TechDogs
  11. [11]
    GPT-5.6 - Wikipedia
  12. [12]
    GPT-5.6 Sol: Benchmarks, Pricing & API Access Guide 2026
  13. [13]
    GPT-5.6 Sol Model | OpenAI API
  14. [14]
    Previewing GPT-5.6 Sol: a next-generation model | OpenAI
  15. [15]
    GPT-5.6: Frontier intelligence that scales with your ambition | OpenAI
  16. [16]
    GPT-5.5
  17. [17]
    OpenAI signs $10 billion deal with Cerebras, with 750MW of big-chip compute - DCD
  18. [18]
    OpenAI's $10 Billion Cerebras Deal: How Wafer-Scale Computing Is Solving the AI Infrastructure Crisis | Programming Helper Tech
  19. [19]
    OpenAI Signs $10 Billion Deal With Cerebras for AI Computing - Bloomberg
  20. [20]
    OpenAI chip deal with Cerebras adds to roster of Nvidia ...
  21. [21]
    Cerebras scores OpenAI deal worth over $10 billion ahead of AI chipmaker's IPO
  22. [22]
    OpenAI Secures Multi-Year Compute Agreement With Cerebras Valued at Over $10B
  23. [23]
    openai signs deal reportedly worth 10 billion for compute from cerebras
  24. [24]
    OpenAI signs $10 billion computing deal with Nvidia challenger Cerebras