Skip to main content
TechnologyJul 22, 2026· 4 min read

Does NVIDIA Want to Retire x86 CPUs? Vera Promises Performance Up to 2 Times Greater

NVIDIA has published a detailed analysis of the new Vera CPU, the Arm processor designed for AI data centers and specifically built for Agentic AI and reinforcement learning workloads. According to the company, the evolution of artificial intelligence models makes the central processor an increasingly important element in executing, orchestrating, and managing the context needed to power GPUs, reducing bottlenecks that limit overall system performance.

Vera is based on 88 proprietary Olympus cores, developed on the Armv9.2 architecture. The chip adopts a monolithic design, integrates the second generation of NVIDIA's Scalable Coherency Fabric, supports NVLink-C2C, PCI Express 6.4, CXL 3.1, and features SOCAMM2 LPDDR5X memory with capacities of up to 1.5 TB.

The configuration includes 88 cores and 176 threads thanks to Spatial Multi-Threading technology, a solution NVIDIA proposes as an alternative to the traditional Simultaneous Multi-Threading used by x86 CPUs. Instead of relying solely on opportunistic resource sharing, the Olympus core has been designed to more effectively distribute resources between two hardware threads. This allows Olympus to operate like a high-throughput single-threaded core when maximum per-thread performance is required, while the twin thread handles management tasks, or as two more isolated execution contexts when greater thread density is necessary.

According to NVIDIA, the result is high performance for single threads, better resource utilization, and more predictable behavior under load. The company states that Olympus was designed with a single goal: to maximize IPC in applications characterized by irregular execution flows, numerous conditional jumps, and a high level of concurrency typical of modern AI agents. Compared to the previous Grace architecture, Vera promises an IPC increase of up to 50%.

Among its distinguishing features are an advanced Neural Branch Prediction system, a decoding pipeline that is 10 instructions wide, extensive out-of-order execution resources, memory renaming techniques and value prediction to increase parallelism, along with an execution engine equipped with eight simple ALUs, two complex ALUs, and six 128-bit SVE vector units also compatible with FP8 calculations. Each core also has 2 MB of L2 cache, while the CPU integrates 164 MB of shared L3 cache.

Particular attention has been given to the memory subsystem. Vera uses eight LPDDR5X controllers operating up to 9,600 MT/s, with an aggregate bandwidth that can reach 1.2 TB/s. NVIDIA claims that each core can have about 14 GB/s of memory bandwidth, a value significantly higher compared to traditional x86 platforms. The use of LPDDR5X memory in the SOCAMM2 format is also indicated as a key element for reducing energy consumption: according to the company, a complete system would require about 30-40 W for the memory subsystem, against values that can exceed 100 W with DDR5 RDIMM platforms and over 200 W with some MRDIMM implementations.

On the interconnection front, Vera provides up to 88 PCIe 6.4 lanes, which become 176 in dual-socket configurations, while NVLink-C2C achieves a coherent bandwidth of up to 1.8 TB/s between CPUs and GPUs. The platform also integrates advanced Confidential Computing features, with virtual machine isolation, memory encryption, and support for Arm extensions dedicated to secure execution environments.

Alongside the presentation of the architecture, NVIDIA has released a series of internal benchmarks that compare Vera with AMD EPYC Turin processors based on the Zen 5 architecture. As always in these cases, the results were produced by the manufacturer itself and will need to be verified through independent testing (here's a preview) when the processor becomes available.

In NVIDIA's architectural tests, they claim IPC increases of up to 1.9 times compared to Zen 5, branch prediction up to 2.3 times faster, and backend operation throughput that would see increases of up to 4.3 times. The frontend would also benefit from the new Olympus architecture, with performance up to 2.4 times better in instruction fetch operations. AMD does not fear NVIDIA: in its tests, EPYC Venice beats the Vera CPU.

The company attributes a good portion of these results to the choice of a monolithic design. According to NVIDIA, the chiplet architectures used by modern x86 CPUs introduce additional latencies in communication between different dies and in the memory subsystem. In the published benchmarks, Vera maintains an almost constant latency even with high levels of memory usage, offering up to three times the available bandwidth and a maximum claimed latency of up to 40 times lower compared to the AMD EPYC 9755 under the most demanding load conditions. Core-to-core latency would also be about 50% lower, with a more uniform distribution among all core pairs.

In application loads targeted at Agentic AI, NVIDIA reports performance up to 1.8 times better in executing Python code, increases of up to 1.7 times in other software workloads, a 2.6 times advantage in graph traversal algorithms, and an improvement of up to 20% in ClickHouse-based analytical databases. In a specific data processing test, the company also claims a performance increase of up to 6 times with a 40% reduction in latency compared to x86 CPUs.

According to NVIDIA, Vera aims at an addressable market estimated at around 200 billion dollars. The company also claims that the CPU has already been adopted in the initial phase by entities such as OpenAI, Anthropic, SpaceX, and Perplexity, in addition to being planned for future supercomputers and cloud infrastructures dedicated to artificial intelligence.