Overview
<cite index="6-2">Nvidia released a white paper on July 21, 2026, providing an architectural overview of the Vera central processing unit (CPU) and the technologies underpinning its performance, efficiency, and scalability.</cite> The publication marks the most detailed public disclosure yet of <cite index="9-3">Nvidia's first in-house CPU core design</cite>, and arrives as the broader Vera Rubin platform ramps toward commercial availability.
Architecture and Core Design
<cite index="3-12,3-13">Nvidia has published new architectural details for its Vera data center CPU and the custom Olympus core. Vera combines 88 Olympus cores, 176 hardware threads, and a 164 MB unified L3 cache on a monolithic compute die.</cite> <cite index="1-5">The chip was designed to prioritize single-thread performance and memory bandwidth rather than raw core density.</cite>
<cite index="3-14,3-15">Each Olympus core uses a wide front end with a neural branch predictor, a 64 KB four-way L1 instruction cache, and a 48-instruction decode queue. The core can fetch up to 16 instructions per cycle and includes a 10-wide decoder capable of processing ten fused instructions per cycle.</cite>
<cite index="3-3,3-4,3-5">Vera uses Nvidia Spatial Multithreading, which provides two hardware threads per Olympus core. Nvidia says the design can partition core resources between threads to reduce contention compared with conventional simultaneous multithreading (SMT), allowing a core to prioritize one performance-sensitive thread while its second thread handles system and management tasks.</cite>
<cite index="7-3,7-4">Rather than using chiplets, like many competitive high core-count server processors, Vera uses a monolithic die with Nvidia's second-generation Scalable Coherent Fabric (SCF) linking the various intellectual property blocks. Using a monolithic design eliminates the latency penalties associated with traversing multiple chiplets, while simultaneously delivering roughly three times greater core-to-core bandwidth.</cite>
Memory and Interconnect
<cite index="3-8">Nvidia lists up to 3.4 TB/s of core-to-core bandwidth, 1.2 TB/s of SOCAMM2 Low-Power Double Data Rate 5X (LPDDR5X) bandwidth, and up to 1.5 TB of memory capacity per CPU.</cite> <cite index="7-7">The company claims the design provides roughly three times more memory bandwidth per core, five times greater bandwidth-per-watt efficiency, and approximately 40% lower memory latency under load than competing server platforms.</cite>
<cite index="5-6">The most strategically significant specification is the integrated NVLink-C2C interface, which provides up to 1.8 TB/s of coherent bandwidth between the CPU and Rubin Graphics Processing Units (GPUs).</cite> <cite index="9-9">The CPU and GPU share a coherent address space via NVLink-C2C.</cite>
Target Workloads and Performance
<cite index="4-9">Nvidia said the chip was designed for agentic artificial intelligence (AI) workloads, which depend on strong per-thread progress through branch-heavy code, pointer-heavy data structures, and long dependency chains.</cite> <cite index="8-3,8-4">Vera is built for the CPU work behind agentic AI and reinforcement learning (RL), including code execution, tool use, sandboxing, analytics, data pipelines, and orchestration beyond the model — serving as both a host CPU for accelerated systems and a standalone CPU for AI factory workloads.</cite>
<cite index="4-2,4-3,4-4">Nvidia published the white paper alongside unofficial SPEC CPU 2026 results comparing Vera against AMD's EPYC 9755, with both chips running in a dual-socket configuration. In the SPECrate integer suite, Vera's overall base score came to 925, edging out 898 for a dual-socket EPYC 9755 system — a roughly 3% margin achieved while operating with a smaller thread count.</cite>
Platform Context and Vertical Integration
<cite index="16-2">The broader Vera Rubin platform brings together the Nvidia Vera CPU, Nvidia Rubin GPU, Nvidia NVLink 6 Switch, Nvidia ConnectX-9 SuperNIC, Nvidia BlueField-4 Data Processing Unit (DPU), and Nvidia Spectrum-6 Ethernet switch, as well as the newly integrated Nvidia Groq 3 Low-Power Unit (LPU).</cite>
<cite index="14-5">Each Vera CPU rack integrates 256 Vera CPUs and supports more than 22,500 concurrent sandbox environments, providing scalable, energy-efficient CPU capacity for tool calls, evaluation, data processing, and orchestration.</cite>
<cite index="11-7,11-8,11-9">Los Alamos National Laboratory has selected Vera Rubin technology for three new supercomputers — Mission, Vision, and Veritas. Mission will focus on national security workloads, while Vision will support open scientific research and AI-driven discovery. Veritas is specifically designed to enable agentic AI applications in scientific research, combining Rubin GPUs with standalone Vera CPU partitions.</cite>
<cite index="9-12">Commercial availability is expected in fall 2026; no pricing has been disclosed.</cite>