At "Advancing AI 2026" on July 23, 2026, AMD unveiled its 6th-generation server CPU, the "EPYC 9006." Leading the lineup, the EPYC 9996 reaches 256 cores and 512 threads, a 33% increase in core count over the previous generation's flagship, the EPYC 9965. The SP7 platform raises memory bandwidth to up to 1.6TB/s and adopts PCIe 6.0. SP7 handles high density and high clock speeds, while SP8 targets general-purpose enterprise use. There's also the 9006X with a large cache, and an LP variant dedicated to AI hosts.

This branching reflects how server CPUs in the AI era face simultaneously different demands. Running many agents in parallel benefits from core density. A host CPU that feeds data to GPUs needs fast single-core performance along with responsive memory and I/O. Rather than assigning the same CPU to every use case, AMD has made it possible to choose a design point within the same generation.

AD

Four EPYC 9006 Families, Bigger Than 256 Cores

The EPYC 9996's 256 cores mark an increase of 64 cores over the current EPYC 9965's 192 cores. With SMT (Simultaneous Multithreading) enabled, it reaches 512 threads. As of July 2026, AMD describes this as the highest core count among announced single-socket CPUs. Comparing the upper limits against the previous generation reveals the broader scope of the platform as a whole.

Item 5th-Gen EPYC 9965 6th-Gen EPYC 9006 SP7 Max
Cores/Threads 192/384 256/512
L3 Cache 384MB 1,024MB
Memory Bandwidth 614GB/s 1.6TB/s
PCIe 5.0, 128 lanes 6.0, 64Gbps/lane

The right-hand column represents the maximum values across the EPYC 9006 SP7 product family, not a single SKU that satisfies all of them simultaneously. Even so, core count grew roughly 1.33x, the L3 cache ceiling roughly 2.67x, and memory bandwidth roughly 2.61x. PCIe has also advanced a generation, with AMD stating that per-lane bandwidth has doubled. In other words, AMD has simultaneously widened the pathways for sending data to GPUs, networking, and storage.

On the manufacturing side, AMD announced in May that it had begun mass-production ramp-up of Venice in Taiwan. Venice is manufactured on TSMC's 2nm process, and AMD states it is the first HPC product to enter mass-production ramp-up on that process. Future plans call for ramping up production at TSMC's Arizona facility as well.

Density Backed by 1.6TB/s and PCIe 6.0

Turning 256 cores into actual processing throughput requires continuously feeding data to each core. The EPYC 9006 SP7 employs a 16-channel memory configuration, with AMD claiming bandwidth of up to 1.6TB/s—a figure aimed at suppressing memory wait times even as core counts increase.

Within the same SP7 family, the high-clock-speed variant tops out at 96 cores and 5.0GHz. This is because GPU servers need to quickly complete preprocessing and control tasks to keep accelerators from sitting idle. The EPYC 9996, which lines up massive numbers of independent tasks, and the 96-core variant that accelerates the progress of a single thread, serve different roles despite both being built on Venice.

For technical computing workloads that demand even more cache, AMD offers the EPYC 9006X SP7. The EPYC 9684X, equipped with 3D V-Cache, reaches an L3 cache of 1,152MB, with 16-channel MRDIMM support up to 12,800MT/s and clock speeds of up to approximately 5.15GHz. By separating the density of the standard SP7, the responsiveness of the high-clock variant, and the data locality of the 9006X, users can select a CPU based on memory latency or response-time requirements.

AD

How to Interpret the Claimed 3.45x

AMD compared the EPYC 9996 against the Intel Xeon 6980P and the EPYC 9965 across five workload groups that support agentic AI. Normalized values, with the Xeon 6980P set at 1.00, are as follows.

Workload Group EPYC 9965 EPYC 9996
Gateway/Streaming Response 2.39 2.83
Context Building/Planning/Routing/Verification 2.01 3.45
Search/Similarity Vector Search 1.49 2.38
Enterprise Tools 1.69 2.65
Short-Duration Tool Execution 1.67 2.50

The maximum figure of 3.45x corresponds to the workload group that aggregates context building through verification. Across generations, this same category rose from 2.01 to 3.45, which AMD expresses as up to a 1.7x improvement. This is not a claim that every application will run 3.45x faster than on Xeon.

Within each workload group, separate individual tests are blended together—Web processing via NGINX and WRK, similarity search via FAISS, TPCx-AI, and so on. The enterprise tools column is also a geometric mean of multiple database benchmark results. In other words, this chart spans the peripheral CPU-side processing that AI agents invoke, rather than representing an end-to-end composite score for running a single agent from start to finish.

The figures are based on AMD's own measurements and estimates, not results independently reproduced by a third party on commercially available servers. Since the EPYC 9996 has twice the core count of the Intel comparison target, workloads using per-core software licensing, or tasks that don't parallelize well, won't achieve the same economics. One needs to weigh the number of concurrent tasks that can fill 256 cores together with overall server power consumption and pricing.

A Different Approach to Bandwidth Allocation Than Vera

In its comparison against NVIDIA Vera, AMD used two different Venice configurations depending on the metric. For SPEC CPU 2026 integer throughput, AMD compared a 256-core, 600W Venice configuration against its own estimate for Vera, claiming a 2.2x advantage. For per-core performance, meanwhile, AMD used the 96-core high-clock variant, claiming a 1.2x advantage over an 88-core, 450W Vera. Both figures are explicitly labeled as "estimates" in AMD's materials, and the compiler was standardized to GCC 15.2.

Memory tells an opposite story. The Venice SP7 reaches up to 1.6TB/s in total bandwidth, exceeding Vera's LPDDR5X-based maximum of 1.2TB/s. However, if one simply divides the stated peak bandwidth by the maximum core count, Venice comes out to 6.25GB/s per core, while Vera comes out to roughly 13.6GB/s per core. NVIDIA emphasizes single-core speed and memory efficiency with Vera, while AMD foregrounds total bandwidth and density across 512 threads.

This ratio does not predict sustained bandwidth or application-level performance. Real-world differences will depend on actual memory efficiency, on-chip data movement, and the degree of parallelism software can exploit. Even so, one can see why AMD chose the 256-core variant for its throughput comparison and the 96-core variant for its per-core performance comparison. Venice allows users to choose different design points depending on the use case.

AD

A Use-Case-Specific Roadmap Extending to 2027

The four product families won't all reach the market at the same time. Laying out the delivery timelines reported by Tom's Hardware based on AMD's briefing, alongside the key specifications AMD has disclosed, produces the following picture.

Product Family Key Specifications Use Case Availability
EPYC 9006 SP7 Up to 256 cores, 16 channels, PCIe 6.0 High-density workloads, GPU hosts, general-purpose Q4 2026
EPYC 9006 SP8 8–128 cores, 8 channels, 128 PCIe 6.0 lanes Enterprise, edge, power-efficient configurations First half of 2027
EPYC 9006X SP7 Up to 96 cores, 1,152MB L3, up to approx. 5.15GHz HPC, EDA, analytics Second half of 2027
EPYC 9006 LP Up to 72 cores, LPDDR5X, 112Gbps xGMI Next-gen rack AI hosts Second half of 2027

SP8 supports both single-socket and dual-socket configurations, emphasizing a balance between memory capacity and I/O. The 9006X expands the working set that fits within cache, reducing latency for simulation and similar workloads. The 9006 LP, formerly codenamed Verano, pairs LPDDR5X with CPU-to-GPU interconnects and will serve as the host CPU for the next-generation Helios platform.

AMD has not yet disclosed pricing for all SKUs, the base clock speed of the EPYC 9996, or power consumption figures for complete servers. Turning the claimed 3.45x figure into an actionable procurement decision will require independent measurements—using identical software and power caps—on SP7-equipped systems once they ship in Q4 2026. Whether 256 cores can actually reduce the number of servers and the amount of power needed for a given workload will only become measurable once real-world performance and server pricing are both in hand.