On October 7 (US time), Microsoft opened pre-orders for the Surface Laptop Ultra, powered by NVIDIA's Arm-based "RTX Spark" chip. It is scheduled to ship on October 16. US pricing starts at $2,599.99, and the Japanese Microsoft Store is taking pre-orders from ¥513,480 including tax. According to the official announcement, the creator- and AI-development-focused laptop previewed in May has now become a product you can actually buy.

However, the starting price and the headline claims of "up to 128GB" and "up to 1 petaflop" do not refer to the same configuration. Microsoft's AI comparison against the MacBook Pro also comes with conditions on memory capacity, model, and numerical precision. The value of the Surface Laptop Ultra will vary widely depending on which configuration you choose and which software you run.

AD

The ¥510,000 entry point and the 128GB model are different configurations

On the Japanese sales page, the cheapest model pairs the RTX Spark N1X, with an 18-core CPU and a GPU with 5,120 CUDA cores, with 24GB of shared memory and a 512GB SSD. The higher-end N1X, with a 20-core CPU and 6,144 CUDA cores, starts at ¥707,080 in a 32GB/1TB configuration.

Representative Japan configurations CPU / GPU Memory / SSD Price (tax incl.)
Lowest-priced model 18 cores / 5,120 CUDA cores 24GB / 512GB ¥513,480
Entry to the higher-end chip 20 cores / 6,144 CUDA cores 32GB / 1TB ¥707,080
64GB model 20 cores / 6,144 CUDA cores 64GB / 1TB ¥812,680
Maximum-capacity model 20 cores / 6,144 CUDA cores 128GB / 1TB ¥1,094,280

Prices and configurations reflect what the Japanese Microsoft Store displayed on October 8. The 128GB model was listed as "out of stock" at the time. Available configurations also vary by color and keyboard layout. The cheapest 24GB model ships with Windows 11 Home, while the 32GB and larger models in the table ship with Windows 11 Pro.

The cheapest configuration in Japan has 24GB, a different memory capacity from the 64GB configuration used in the official AI comparison. The comparison targets in Microsoft's October 7 announcement, footnotes 5–8, are a prototype PC with RTX Spark and 64GB of memory and a 64GB MacBook Pro. Checking the table above against that memory capacity, the listed price for a 64GB configuration in Japan is ¥812,680.

Of course, there is no guarantee that a 64GB production Surface will match the prototype's speed. The CPU, GPU, SSD and OS also differ from the cheapest model, so the price gap can't be treated as simply an extra charge for memory. If you're buying for local AI, decide which models you'll run and how much capacity your other apps need before looking at the starting price.

What RTX Spark combines in one machine

RTX Spark pairs a Grace CPU with a Blackwell RTX GPU, and the CPU and GPU share the same memory. Unlike designs with separate GPU memory, the system's large memory pool can also be used for AI and graphics workloads. The design aims to bring CUDA-compatible development environments and RTX-accelerated creative software together in a portable Windows machine.

According to NVIDIA's official specifications, the top laptop N1X supports up to 128GB of memory, and the 18-core lower-tier N1X up to 64GB. Both have a TDP of 45–80W, and sustaining real-world performance depends on the chassis's cooling capacity. TDP is a design power envelope, not a figure indicating that the chip draws that power in every task.

It would also be premature to treat the 128GB of shared memory as 128GB available for models. Microsoft states that the maximum amount the GPU can use varies with configuration and workload and is less than the total system capacity. The OS and other running apps also consume memory. Whether a model fits has to be judged by including the quantization setting, which represents weights at lower precision, and how long an input context you keep.

The "up to 1 petaflop" is likewise a theoretical FP4 AI performance figure that uses sparsity, a feature that exploits sparse data. The computational conditions differ from standard FP32 operations and from competitors' NPU INT8 TOPS figures. The magnitude of the number can't be translated directly into app speed.

The Surface hardware is tailored to creative work. The 15-inch mini-LED touch display supports 120Hz and has a peak HDR brightness of 2,000 nits. That is a peak value measured over 10% of the screen, though, and differs from the brightness sustained across the full panel. There are three USB-C/USB4 ports, and the Magnetic Connect-compatible port on the right side attaches a dedicated cable magnetically while also working as a regular USB-C port for data and video.

The Japanese sales page says a 140W USB-C power adapter and magnetic cable are included. With HDMI, USB-A and an SD card reader also on board, you should spend less time hunting for adapters every time you import footage. The SSD is removable, using PCIe Gen 4 for 512GB and Gen 5 for 1TB and 2TB. There is room to evaluate it as a creative machine, including its display, ports and serviceability.

AD

How to read the "up to 6.2x" claim against the M5 Pro

Microsoft announced that a PC with RTX Spark is up to 2.1x faster than the M5 Pro 16-inch MacBook Pro in time to first token, up to 4.3x faster in image generation, and up to 6.2x faster in video generation. Reading the measurement conditions shows that these figures come from separate tests for each use.

Workload compared Maximum multiple announced Model and main settings
Time to first token 2.1x llama.cpp, Qwen3.5 27B, Q4K-Medium, fixed 8,192-token input
Image generation 4.3x ComfyUI, FLUX.2 Klein 4B, NVFP4, 4 steps, 1024×1024
Video generation 6.2x ComfyUI, LTX 2.3 22B, NVFP4, 121 frames, 25fps, 8 steps, 1280×720

All were measured in September 2026, and both comparison machines had 64GB of memory. The time-to-first-token test was commissioned by Microsoft, while the image and video tests were measured by NVIDIA. The RTX Spark side included multiple prototype Windows PCs, and the text test included a prototype Surface Laptop Ultra. It is not an announcement that the Surface alone achieved each multiple.

The 2.1x for text compares how long it takes for a response to begin after a long input is given to a 27-billion-parameter Qwen model. That is a different metric from text generation speed once the response has started. The image and video tests also run specific models at low precision with fixed output settings. The same multiples won't necessarily appear with other models or higher-precision settings.

The footnotes also don't detail the Mac side's numerical precision, the mechanism used to execute the processing, or raw processing times in seconds. The NVFP4 label doesn't allow us to read this as a comparison in which the Mac side used the same hardware feature. The published numbers are best treated as the manufacturer's measurements showing that NVIDIA-side AI processing had the advantage under these conditions.

The comparison offers concrete guidance for readers who use CUDA and low-precision AI processing. It doesn't, however, show general CPU performance, performance in all creative software, or a ranking that includes the M5 Max. From here, the judgment changes depending on which competing chip you are weighing it against.

Apple, AMD, Intel and Qualcomm are not competing on the same ground

If you're choosing a Surface for its large memory, comparing it with the M5 Pro alone isn't enough. Apple's M5 Max can also handle 128GB, and AMD's Ryzen AI Max+ PRO 495 supports up to 192GB. Beyond capacity, NVIDIA's bid is whether work that relies on CUDA and RTX can be consolidated on the same machine.

Chip Published CPU configuration Published integrated GPU configuration Maximum memory Main condition for choosing it
RTX Spark N1X (higher tier) Arm-based, 20 cores Blackwell RTX, 6,144 CUDA cores 128GB Use CUDA/RTX on Windows on Arm
Apple M5 Pro Up to 18 cores Up to 20 GPU cores 64GB Use macOS creative environment
Apple M5 Max 18 cores Up to 40 GPU cores 128GB Use large capacity and high bandwidth on macOS
Ryzen AI Max+ PRO 495 x86, 16 cores/32 threads Radeon 8065S, 40 CUs 192GB Check x86 environment and Radeon-compatible software
Core Ultra X9 388H x86, 16 cores/16 threads Arc B390, 12 Xe cores 96GB Check existing x86 environment and Intel-compatible software
Snapdragon X2 Elite Extreme X2E-94-100 Arm64, 18 cores Adreno X2-90 128+GB; actual installed amount depends on device Use NPU-compatible workloads on Windows on Arm

The table compares each company's published specifications as checked on October 8. For Intel it draws on the product specifications, and for Qualcomm on the spec table in the product brief. It lines up chip ceilings, which differ from the configurations sold in actual PCs and from the amount the GPU can actually use. CUDA cores, CUs, Xe cores and Apple's GPU cores are counted differently, so dividing the counts to get a performance ratio isn't valid.

On Apple's side, the M5 Pro has memory bandwidth of up to 307GB/s and the M5 Max up to 614GB/s. Even at the same 128GB capacity, the speed at which data moves is a separate condition. Meanwhile, we haven't been able to confirm the Surface's DRAM bandwidth in the same units, so we can't rank by bandwidth-based speed here. The first major dividing line is whether you use software that runs on CUDA or whether your work is fully served by an existing macOS environment.

AMD goes further on capacity. Of the Ryzen AI Max+ PRO 495's 192GB, graphics memory is described as up to 160GB. That does not mean 192GB and 160GB add up. Its configurable TDP is 45–120W, with a different upper limit from the RTX Spark laptop's 45–80W. Headroom for handling large models and how fast a thin chassis can run for long periods should be considered separately.

Intel's 388H has a base power of 25W and a maximum turbo power of 80W, with an integrated Arc GPU and NPU. For people who rely heavily on existing x86 apps and peripherals, it is a starting point, alongside AMD, for checking migration conditions. However, base power, AMD's cTDP and NVIDIA's TDP are not the same kind of measurement, and the numbers alone can't determine battery life.

Qualcomm's X2E-94-100 touts 228GB/s of memory bandwidth and a Hexagon NPU rated at 80 TOPS in INT8. A design that uses NPU-compatible AI processing and one that loads large models onto the GPU via CUDA suit different software. Not putting NVIDIA's FP4 performance and NPU TOPS in the same speed ranking table is the first step in a fair comparison.

AD

Software support and sustained performance will decide the purchase

On Windows on Arm, even a fast GPU requires you to check migration issues. According to Microsoft's developer documentation, x86 and x64 apps can run via emulation, but kernel components such as drivers must be built for Arm64. Even if CUDA workloads run, you need to verify separately whether you can finish your work including surrounding plug-ins and peripherals.

Games also require title-by-title support. Microsoft showcased Gears of War: E-Day as an example of RTX Spark's graphics features and said Call of Duty support would come in 2027. Support for DLSS and ray tracing is appealing, but this is not information that guarantees, across the board, that the games you usually play will launch or work with anti-cheat.

Battery life needs the same careful reading. The Surface Laptop Ultra's stated figures are up to 15 hours of video playback and up to 12 hours of web use. However, those come from tests of a prototype with the higher-tier N1X, 64GB and 1TB, running at 150 nits of screen brightness, not from continuous AI or gaming workloads. Whether it can sustain speed under heavy load, and how that changes when unplugged, should be verified on production units.

Microsoft itself promotes an AI environment that splits work between the PC and the cloud. The stationary Surface RTX Spark Dev Box, whose pre-orders also opened at the same time, is US-only at $5,999, with shipping planned for November, a different schedule from the laptop. Where you put your investment depends on whether you want to carry local AI with you or leave it to a desktop machine.

With pre-orders open, the conditions for choosing the Surface Laptop Ultra have become more concrete. If your models fit, the CUDA/RTX software and peripherals you need run, and it can hold its speed during long jobs, it becomes an option for carrying creative and AI development work with you while cutting cloud spending. Making that judgment requires separating the 24GB starting price from the 128GB maximum spec and testing real workloads on the configuration you plan to use.