On September 18, 2026, AMD released performance data for its 6th-generation server CPU, the EPYC 9006 series. The flagship EPYC 9996 reportedly achieved 2.9x throughput on Redis, 3.7x on NGINX, and 3.5x on MongoDB compared to Intel's Xeon 6980P. Against NVIDIA's Vera, AMD claimed a 2.24x advantage per two-socket configuration and a 1.2x advantage per core on SPECrate 2026 Integer, a benchmark measuring integer parallel throughput.

These are big numbers, but they don't all rest on the same kind of evidence. The enterprise and cloud workload comparisons against Xeon are based on AMD's own physical measurements, while the SPEC figures for Vera are AMD estimates. The 100kW rack comparison isn't a measurement of physical racks lined up side by side—it's a model combining per-node performance with the number of nodes that could theoretically fit. Furthermore, EPYC and Xeon differ in core count and memory configuration, and the compilers and power settings used aren't aligned either. What this data shows isn't a uniform measure of raw CPU speed, but the throughput of platforms configured for specific workloads.

AD

The 3.7x Figure Compares 256 Cores Against 128

For enterprise and cloud workloads, AMD compared a single-socket EPYC 9996 against a single-socket Xeon 6980P. The EPYC chip has 256 cores and 512 virtual CPUs with simultaneous multithreading enabled, while the Xeon has 128 cores and 256 virtual CPUs. Memory configurations also differ: EPYC uses DDR5-8000 across all channels, while Xeon uses DDR5-6400. Both systems run Ubuntu and use KVM for virtualization, but the server hardware and BIOS settings differ as well.

Workload EPYC 9996 Xeon 6980P AMD's Claimed Multiplier
Redis 124,130,835.50 42,636,320 2.9x
NGINX 38,451,232 10,526,326 3.7x
MongoDB 6,195,941 1,794,003 3.5x

Source: AMD internal measurements as of July 2026. Redis figures reflect SET/GET requests with one server instance per 8 virtual CPUs; NGINX figures reflect 1KB file requests via WRK; MongoDB figures come from AMD's own throughput testing. The units are only valid for comparison within each workload—absolute values cannot be compared across different workloads.

The key takeaway from this table is that EPYC, packing twice the cores and virtual CPUs into the same single socket, showed major gains in workloads that can exploit that parallelism. NGINX and Redis tests run many server instances—each handling 8 virtual CPUs—and sum the results. This setup approximates environments where many concurrent agents simultaneously generate web requests, caching operations, and database queries. However, it doesn't mean that a single request's response time becomes 3.7 times faster. Neither the cost of per-core-licensed software nor total server power consumption—which includes AMD's 600W "Default CPU Power" baseline—can be judged from this multiplier alone.

The MySQL test also requires caution. AMD used TPROC-C, an open-source workload based on TPC-C, but results from it cannot be compared to officially published TPC-C-compliant benchmarks. While AMD claims a 2.6x improvement, this figure cannot be translated into an official TPC-C record.

The 2.24x Vera Comparison Pits Two Estimates Against Each Other

For the Vera comparison, AMD used different Venice configurations depending on the question being asked. The per-core performance figure comes from a high-clock 96-core variant: an estimated two-socket score of 1210 divided by total core count yields 6.3. For Vera, an estimated score of 925 across two sockets with 88 cores each, divided by 176 cores, yields 5.3—giving AMD's claimed ratio of roughly 1.2x.

For platform-wide throughput, AMD compared an estimated score of 2070 using two 256-core EPYC 9996 chips against Vera's estimated 925. Dividing 2070 by 925 yields approximately 2.24. AMD states that both sides used GCC 15.2 as the compiler, but AMD itself characterizes both figures as "internal estimates" and "preliminary engineering projections." These are not results from third-party testing of two commercially available servers under identical conditions.

According to NVIDIA's official specifications, Vera features 88 Olympus cores, 176 threads, up to 1.2TB/s of LPDDR5X memory bandwidth, and a 250–450W power range. Vera emphasizes per-core performance and memory bandwidth, while the EPYC 9996 leverages its 256-core density. The 2.24x figure reflects the latter's parallel throughput advantage, but this gap may not hold for workloads that rely heavily on single-thread response times or memory bandwidth. NVIDIA itself also describes Vera's specifications as provisional.

AD

The 3.4x Rack Claim Isn't a Physical Measurement

In its September announcement text, AMD stated that the EPYC 9996 delivers 3.4 times the throughput of Vera per 100kW rack, based on estimates. This comparison assumes all systems use two-socket nodes, multiplying per-node performance by the number of nodes that could fit within a 100kW power budget. The workloads covered include SPEC integer performance, Java, NGINX, Redis, Memcached, and TPROC-C—six in total. Vera's per-node performance figure was also an AMD estimate derived from existing data.

However, AMD's own published materials contain an inconsistency: the same 100kW rack comparison yields 3.4x in the September announcement text but 3.30x according to the methodology described in a June document.

AMD Publication Venice-to-Vera Ratio Verifiable Methodology
June 2026 rack comparison 3.30x Geometric mean of 6 workloads, two-socket, 100kW, normalized node count
September 2026 Newsroom text 3.4x Footnote refers back to the June methodology document

Rounding 3.30 to one decimal place gives 3.3, not 3.4. Whether AMD updated its inputs for the September version, changed the workload mix, or simply made an error in the text, could not be determined from the publicly available materials we reviewed. As such, the 3.4x figure should be treated not as a reproducible physical rack measurement, but as an AMD model output whose discrepancy remains unresolved.

The idea of establishing a common 100kW power envelope for comparison is itself reasonable, closely mirroring real data center constraints. But CPU specifications alone aren't sufficient to judge this. Node-level power consumption also involves memory, motherboards, cooling, and networking—all of which affect both sustainable performance under load and how many nodes can fit in a rack. AMD itself explicitly notes that its modeled results may not reflect actual deployment performance.

MRDIMM Gains Vary Significantly by Workload

For HPC benchmarks, AMD used two EPYC 9996 chips (512 cores total) against a Xeon 6980P setup with 256 total cores. Both systems had simultaneous multithreading disabled, and results represent averages of three runs. The EPYC system was equipped with 2,048GiB of memory, while the Xeon system had 1,536GiB. Compilers also differed: AMD used AOCC and AOCL, while Intel used its own compiler and MKL—meaning this isn't a test that isolates CPU performance under identical conditions.

Using MRDIMM configurations, AMD reported gains over Xeon of 3.13x for the molecular dynamics tool GROMACS, 3.08x for NAMD, 1.80x for the materials science tool Quantum ESPRESSO, and 2.90x for the weather modeling tool WRF. Here too, throughput, elapsed time, and average time metrics are mixed together, requiring careful conversion to a consistent direction for proper interpretation.

Looking specifically at memory performance on the same EPYC 9996 chip, GROMACS improved from 8.249 ns/day with standard DDR5-8000 RDIMM to 9.195 ns/day with faster MRDIMM—an 11.5% gain. NAMD, however, actually declined slightly, from 1.72 to 1.71 ns/day, a 0.6% decrease. While measurement variance wasn't disclosed, the published figures at minimum don't support a claim that increased bandwidth uniformly speeds up all workloads by the same margin.

AD

Evaluating CPU-Adjacent Infrastructure, Not Full Agent Workflows

AMD explains that as agentic AI systems increasingly perform search, planning, and calls to tools and databases, the workload burden on CPUs surrounding GPUs grows correspondingly. The Redis, NGINX, encryption, and database tests in this release represent individual workloads that make up that supporting infrastructure layer. In environments that can exploit high parallelism, the EPYC 9996's core density becomes a compelling option.

However, AMD's own evaluation explicitly excludes GPU-based inference. It doesn't measure the full agent workflow—including model "thinking" time, CPU-to-GPU data transfer, external API calls, or storage latency. Which CPU is appropriate depends first on the number of concurrent workloads and single-thread latency requirements. Memory bandwidth, software parallelization capability, and whether licensing is core-based also factor into the decision.

The next step for evaluation would be independent testing on commercially available OEM servers. This would involve matching power limits and price points, aligning compilers and memory capacity, and measuring not just throughput but response times under heavy load. If end-to-end agent workflows involving GPUs could also be measured, it would become possible to judge how much of Venice's core density advantage actually translates into real service capacity. Rather than fixating on headline figures like the claimed 3.7x, the deciding factor in procurement should be whether your actual workloads can reproduce the specific conditions behind these benchmarks.