« All posts

Inside Nvidia's Vera CPU and its custom Olympus cores

A technical breakdown of Nvidia's Vera CPU and custom Olympus cores, covering monolithic design, branch prediction, and spatial multithreading.

With Vera, Nvidia is making its first serious push into standalone server CPUs, directly challenging Intel and AMD. The chip packs 88 custom Armv9.2 cores, 176 threads, and up to 1.5TB of LPDDR5X memory, and unlike its Grace predecessor it can be sold independently of Nvidia GPUs. Major cloud providers including Alibaba, ByteDance, Meta, Oracle, CoreWeave, Lambda, Nebius, and NScale have already committed to deploying it.

A recently published Nvidia whitepaper reveals the chip's architecture in detail for the first time. Rather than following AMD's multi-die chiplet playbook, Vera houses all 88 cores on a single monolithic compute die fabricated on TSMC's 3nm process, while memory and I/O functions are handled by separate chiplets. In dual-socket form, the Vera CPU Superchip links two dies over a 1.8TB/s NVLink-C2C interconnect for 176 cores and 352 threads, fed by 16 SOCAMM2 modules delivering 2.4TB/s of aggregate memory bandwidth — roughly double what AMD's Turin Epyc offers.

At the core level, Olympus builds on existing Arm IP but adds a fully custom neural branch predictor, memory renaming, and value prediction — all aimed at eliminating pipeline stalls and boosting instructions per clock. It also introduces a rare feature for Arm designs: a form of simultaneous multithreading Nvidia calls spatial multithreading. Together, these choices are meant to make Vera equally suited to managing GPU clusters and hosting the agentic AI workloads that increasingly run outside the GPU.