Apple has released the Mac Studio with M5 Ultra, bringing up to an 80-core GPU and 1.2TB/s of memory bandwidth to the desktop. Measurements from production units have poured in since launch day, giving us enough data to place the M5 Ultra alongside desktop GeForce cards in the same table. But the picture is far from uniform: it approaches the RTX 4080 in Geekbench's OpenCL test, tracks closer to an RTX 5070 machine in gaming, and in some cases beats NVIDIA's compact AI systems at local LLM inference. Let's check the numbers to see where the shorthand "RTX 5080-class" holds up—and where it breaks down.

AD

RTX 4080-class in OpenCL, but not quite 5080-class

Geekbench 7's public OpenCL chart lets you line up Apple, NVIDIA, and AMD GPUs using the same API and the same benchmark generation. As of September 22, 2026, the average score shows the M5 Ultra at 214,954 points—within striking distance of the GeForce RTX 4080's 219,128 and above the Radeon RX 7900 XTX's 212,264. It's clear that Apple's GPU has muscled its way into high-end discrete GPU territory.

According to Geekbench 7 OpenCL's published averages, the M5 Ultra's 214,954 points come within 2% of the RTX 4080's 219,128, but fall roughly 14% short of the RTX 5080's 250,711.

Geekbench 7 OpenCL GPU Score Comparison横棒グラフ。カテゴリ 9 件、系列: OpenCL Score(単位: score)GeForce RTX 5090GeForce RTX 5090GeForce RTX 5090 — OpenCL Score: 355,083score355,083GeForce RTX 4090GeForce RTX 4090GeForce RTX 4090 — OpenCL Score: 274,997score274,997GeForce RTX 5080GeForce RTX 5080GeForce RTX 5080 — OpenCL Score: 250,711score250,711GeForce RTX 4080 SUPERGeForce RTX 4080 …GeForce RTX 4080 SUPER — OpenCL Score: 220,751score220,751GeForce RTX 4080GeForce RTX 4080GeForce RTX 4080 — OpenCL Score: 219,128score219,128Apple M5 UltraApple M5 UltraApple M5 Ultra — OpenCL Score: 214,954score214,954Radeon RX 7900 XTXRadeon RX 7900 XT…Radeon RX 7900 XTX — OpenCL Score: 212,264score212,264GeForce RTX 4070 Ti SUPERGeForce RTX 4070 …GeForce RTX 4070 Ti SUPER — OpenCL Score: 204,801score204,801Apple M3 UltraApple M3 UltraApple M3 Ultra — OpenCL Score: 131,445score131,445単位: score
データを表で見る
OpenCL Score (score)
GeForce RTX 5090355,083
GeForce RTX 4090274,997
GeForce RTX 5080250,711
GeForce RTX 4080 SUPER220,751
GeForce RTX 4080219,128
Apple M5 Ultra214,954
Radeon RX 7900 XTX212,264
GeForce RTX 4070 Ti SUPER204,801
Apple M3 Ultra131,445
Geekbench 7 OpenCL GPU Score ComparisonPublished averages as of September 22, 2026. Higher is better. Only GPUs with at least 5 distinct results are included.出典: Geekbench 7 OpenCL Benchmark Chart

So if we limit ourselves to OpenCL, "RTX 4080-class" fits the numbers, but "RTX 5080-class" stretches the label too far. The gap to the RTX 4080 is within 2%, but the gap to the RTX 5080 widens to roughly 14%. The RTX 5090 sits at 355,083 points, pulling even further ahead. The difference between top-tier GeForce cards isn't small enough to lump them together under one performance tier.

That said, this chart isn't a controlled ranking from manufacturer-run test rigs—it's an average of at least 5 user-submitted results, subject to driver variance and machine condition. Geekbench uses the RTX 4060's 100,000-point score as a baseline, describing double the score as double the OpenCL performance, but that ratio doesn't translate directly into gaming frame rates or AI token generation speed. What this tells us is where the M5 Ultra stands in shared-API compute workloads—nothing more.

Why topping Metal doesn't guarantee gaming supremacy

There's a single published Geekbench 7 Metal result for the M5 Ultra: 354,301 points. On its face, that number is close to the RTX 5090's OpenCL score and falls within the 350,000–360,000 range that AppleInsider reported. But Metal is Apple's own API, while OpenCL serves multiple vendors—the sub-tests and aggregation methods differ. You can't subtract the RTX 5080's OpenCL score of 250,711 from the M5 Ultra's Metal score of 350,000 and conclude the Mac is faster.

The difference shows up in gaming. Tom's Hardware tested a review unit with a 36-core CPU, 80-core GPU, and 256GB of memory, measuring 66fps in Cyberpunk 2077 at 1080p. The outlet judged this comparable to an Acer machine with the RTX 5070, and noted that 4K performance fell short of playable levels. By contrast, a machine with the RTX 5090 hit 59fps at 4K. In 3DMark Steel Nomad, the M5 Ultra outperformed the RTX 5070 machine but was left far behind by the RTX 5090 system.

This isn't a contradiction. OpenCL measures general-purpose compute, Metal is Apple's own graphics and compute framework, and gaming performance depends on rendering APIs, drivers, ray tracing, upscaling, and how well a title is ported. Even with the same 80-core scale, rankings shift depending on which circuits the software can actually engage. Even as more games arrive on Mac, predicting GeForce-matching frame rates from OpenCL rankings alone is risky.

AD

How 1.2TB/s bandwidth and massive memory reshape local LLM rankings

In local LLM inference, what matters isn't just the peak speed of the GPU's compute units—it's how fast you can pull model weights from memory. The M5 Ultra offers 1.2TB/s of bandwidth, with CPU and GPU sharing the same unified memory pool. Token-by-token generation leans heavily on repeatedly reading weights, and this architecture produces a different ranking than gaming does.

Tom's Hardware tested Qwen 3.8-27B-Q4_K_M and reported that the M5 Ultra's generation speed was roughly double that of the M4 Max and about 4x that of the DGX Spark. It reportedly also outpaced the DGX Spark in prompt processing. That said, this is a single test using a dense model with strong memory-bandwidth dependency—the same multiplier won't necessarily apply to image generation, training, or other quantization schemes. Differences between CUDA and Metal implementations also factor into the results.

ITmedia PC USER's generation-to-generation comparison shows similarly large gains. Feeding LLaMA 7B Q4_0 a 512-token prompt and generating 128 tokens, the M5 Ultra hit 204.5 tokens/s, compared to 119.9 for the M5 Max and 92.1 for the M3 Ultra. Within this specific test, the M5 Ultra more than doubled the M3 Ultra's speed.

Local LLM Generation Speed on Apple Silicon横棒グラフ。カテゴリ 3 件、系列: Generation Speed(単位: tokens/s)M5 UltraM5 UltraM5 Ultra — Generation Speed: 204.5tokens/s204.5M5 MaxM5 MaxM5 Max — Generation Speed: 119.9tokens/s119.9M3 UltraM3 UltraM3 Ultra — Generation Speed: 92.1tokens/s92.1単位: tokens/s
データを表で見る
Generation Speed (tokens/s)
M5 Ultra204.5
M5 Max119.9
M3 Ultra92.1
Local LLM Generation Speed on Apple SiliconLLaMA 7B Q4_0, 512-token input, 128-token generation. Not directly comparable to figures from other outlets or models.出典: ITmedia PC USER

In the same test, prompt processing speeds were 6,298 tokens/s for the M5 Ultra, 3,220 for the M5 Max, and 1,471 for the M3 Ultra. Generation and prompt processing are distinct stages, but the M5 Ultra widened the generational gap in both. Even though Apple touts up to 4.3x peak AI compute and up to 1.8x graphics performance over the M3 Ultra, that "up to" figure shouldn't be mistaken for real-world app latency. Actual speed depends on the model, quantization, inference engine, and processing stage involved.

512GB of unified memory isn't the same as VRAM

The up-to-512GB capacity is what most sets the M5 Ultra apart from GeForce cards. According to NVIDIA's official specs, the RTX 5080 has 16GB of GDDR7, the RTX 5090 has 32GB of GDDR7, and even the workstation-grade RTX PRO 6000 Blackwell Workstation Edition tops out at 96GB of GDDR7 ECC. Try to fit a model exceeding 100GB plus its KV cache onto a single GeForce GPU's dedicated memory, and capacity becomes the first wall you hit. The Mac Studio can offload data to the CPU side and cut down on round-trips over PCIe, widening the range of models a single machine can handle.

That said, it's not quite accurate to call 512GB "512GB of VRAM" outright. macOS, the CPU, the GPU, and individual apps all share the same memory pool—the GPU can't claim the entire capacity for itself. Unlike dedicated VRAM, this shared capacity means competing demands on the same pool. On the other hand, because the CPU and GPU can reference the same memory without copying data, that shared architecture becomes an advantage precisely when loading massive models.

The M5 Ultra's starting price of $5,499 covers a configuration with a 30-core CPU, 64-core GPU, and 96GB of memory—not the maximum 80-core GPU with 512GB of memory. The configuration Tom's Hardware tested—36-core CPU, 80-core GPU, 256GB memory, 4TB SSD—cost $12,299. The 512GB configuration was slated to ship in late October at the time of the article's publication, meaning the actual usable GPU limit, sustained performance under prolonged load, and thermal behavior haven't yet been confirmed through independent testing.

AD

Power draw and pricing don't fit on the same scale

Apple lists the Mac Studio's maximum sustained power draw at 480W. NVIDIA specifies the RTX 5080's Total Graphics Power at 360W, the RTX 5090's at 575W, and the RTX PRO 6000 Blackwell Workstation Edition's maximum power consumption at 600W. But that 480W figure covers the entire computer, while NVIDIA's numbers apply specifically to the GPU board. You can't use this to claim the Mac is more power-efficient than the RTX 5090, nor to claim the M5 Ultra's GPU alone draws 480W. What's needed is a real-world measurement of wall-outlet power draw running the same workload.

Which comparison matters also depends on your use case. If you're after gaming performance per dollar, the Cyberpunk 2077 results tilt heavily toward GeForce. Even for CUDA-based development work, memory capacity alone can't offset the cost of migrating existing code and libraries. On the flip side, if you want to run a quantized model exceeding 100GB on a single machine—avoiding network-based distribution or CPU memory swapping—the Mac Studio's up-to-512GB capacity carves out a niche that's hard to replicate with discrete GPUs.

Calling the M5 Ultra "RTX 5080-class" only holds up for specific compute tests where the results happen to land close together. The next numbers that could shift this judgment are the actual usable GPU memory and power draw from production 512GB units, along with Metal-versus-CUDA speed comparisons run on identical models, quantization, batch sizes, and inference engines. Once those numbers are in, the question won't be whether the Mac Studio is fast—it'll be which jobs you can finish on a single machine.