On October 2, 2026, NVIDIA announced a new configuration of its desktop AI computer, DGX Spark, with unified memory cut to 64GB. Partners will begin selling it on October 23, with prices starting at $4,999. It keeps the same GB10 Grace Blackwell Superchip and AI software environment as the 128GB model, giving buyers a choice of memory capacity to match their use.

Still, it is not simply a "cheaper DGX Spark." The starting price of the new 64GB model is higher than the price once published for the 128GB model. And the memory needed to run AI locally is not determined by model size alone. Handling long contexts or running several tasks at once increases the requirement. The 64GB model trades a lower price for less memory, so buyers need to judge, use case by use case, how much they can cut.

AD

Is $4,999 cheaper than the original 128GB model?

Six companies will sell the 64GB model: Acer, ASUS, Dell, Gigabyte, HP and MSI. It will be offered as partner products, not as a NVIDIA-sold Founders Edition. According to NVIDIA's announcement, the starting price is $4,999.

The 128GB Founders Edition, meanwhile, has already seen price increases. In an official forum post on February 25, NVIDIA said it was raising the suggested retail price from $3,999 to $4,699, citing global memory supply constraints. The hardware configuration was unchanged at that point, and NVIDIA directed buyers to check with each manufacturer for OEM pricing. The Register now reports that the 128GB model's price was raised to $6,950 on October 2.

Model and point in time Memory Published/reported price (USD) What the price represents
Founders Edition at launch 128GB $3,999 Pre-revision price cited in the February official notice
Founders Edition after the February 2026 revision 128GB $4,699 NVIDIA's suggested retail price
OEM 64GB model, at planned October 23 launch 64GB From $4,999 Starting price announced by NVIDIA
128GB model after the October 2 revision 128GB $6,950 New price reported by The Register

The sources are NVIDIA's February price-change notice, the current announcement, and The Register's report. Because OEM models and the Founders Edition differ in configuration and sales terms, these are not prices compared under identical conditions.

The 64GB model's $4,999 starting price is $300 higher than the 128GB Founders Edition's previous price of $4,699, and $1,000 higher than its launch price of $3,999. The differences are calculated as 4,999−4,699 and 4,999−3,999. Buyers can spend less than on the current 128GB model, but compared with the earlier 128GB model, they would pay more for a product with half the memory.

Halving the memory does not halve the price. The GB10, the networking features and the CUDA environment are all retained. Whether the 64GB model is worthwhile depends on whether that capacity can handle your actual workloads.

The 64GB limit isn't set by model weights alone

The 64GB in DGX Spark is unified memory shared by the CPU and GPU. There is no separate system memory on top of dedicated GPU memory. Model weights, the OS, applications and the working space needed during inference all draw on the same 64GB.

NVIDIA says the 64GB model can handle models of up to 100 billion parameters. But the memory required is not determined by parameter count alone. It varies with the precision at which weights are stored and with model architecture, so this does not mean every model runs comfortably at every context length.

Keeping a long conversation going requires memory beyond the model itself. A typical example is the KV cache, which stores the results of computations on previously processed text and reuses them when generating the next token, reducing the burden of repeating the same calculations.

According to Hugging Face's Transformers documentation, a dynamic KV cache grows as generation proceeds and can become a major memory constraint with long contexts. However, with methods that reference only a fixed recent window, cache growth for the affected layers may be capped. The actual burden needs to be checked for each model and inference method.

For example, a model that fits in memory for short questions may not be able to handle several requests at once with a long document loaded. Quantizing model weights to shrink them and quantizing the KV cache to compress it are also separate settings. Beyond whether a model loads at all, you need to confirm that the required context length and number of concurrent users can be sustained.

Halving capacity also doesn't uniformly halve processing speed. The Register relays NVIDIA's explanation that the 64GB model's memory bandwidth is unchanged at 273GB/s. The earlier model's specification page likewise lists 273GB/s of memory bandwidth and a 20-core CPU. Capacity indicates how much data can be held; bandwidth indicates how much data can be transferred in a given time. They are separate measures.

The up-to-1 PFLOP compute performance DGX Spark advertises is also a nominal figure that assumes FP4 precision and sparsity. It reflects performance under low-precision arithmetic and data sparsity, and does not directly indicate how fast text is generated. Even with the same GB10, real-world response speed has to be compared with the model, inference software and settings held constant.

AD

Do two units add up to 128GB?

The 64GB model also includes ConnectX-7, allowing two units to be connected over a 200Gbps network. Physically, two 64GB units give a combined 128GB of memory. But this assumes inference software that supports distributed processing; connecting a cable does not make the pair behave like a single 128GB machine.

In NVIDIA's own tests using "Qwen 3.8 27B," it says a two-unit cluster of 64GB models delivered up to 1.7 times the performance of a single unit. However, the announcement blog does not give input and output lengths, the quantization method, detailed inference engine settings, or tokens generated per second. The 1.7x figure cannot be applied directly to other models, nor interpreted as the performance gap against a single 128GB unit.

Price also needs attention. Simply doubling the starting price gives $4,999 × 2 = $9,998, which does not include costs such as connecting cables. Ways to increase memory capacity and ways to reduce purchase cost need to be considered separately.

The tool that helps set up multiple units is the Cluster Assistant in NVIDIA's management app, NVIDIA Sync. According to the official manual, automatic setup covers 2 to 4 units, and the required connections differ as follows.

Units Direct connection Via switch
2 Possible with one QSFP cable Each unit connects to the switch
3 Each unit interconnected with three cables Each unit connects to the switch
4 The Assistant cannot automatically configure direct connections Switch required

Two units can be connected directly, but a four-unit setup requires a network switch. This indicates the range of configurations Cluster Assistant can set up automatically, not an absolute limit on how many DGX Spark units can be connected. Before use, software updates, account setup and the corresponding physical cabling are also required.

What Cluster Assistant automates is the network configuration between machines; distributed setup for inference or fine-tuning itself is a separate task. The official manual states this scope explicitly, and actually running workloads requires following NVIDIA's procedures or similar. Completing the network connection is not the same as having a model properly distributed across multiple units.

Cluster Assistant itself was already introduced in NVIDIA's technical blog on June 1. The newly announced Model Launcher is due at the end of October and will help download Qwen 3.8 27B and launch it on a single unit or a cluster, with support for configuring OpenCode. The network setup assistance available now and the simplified model launching planned for later are separate features.

If you choose 64GB, test the headroom on real work

MSI's announcement the same day shows how the two models might be divided. The company positions the 128GB model for large-model development and fine-tuning, and the 64GB model for running AI agents at branches, factories and similar sites. The idea is that the development environment and the operating environment at each site need not have the same memory capacity.

If the model is fixed and the required context length and concurrency are known, it is relatively easy to judge whether 64GB is enough. For development work where models are swapped frequently, or for work involving long documents, judging by free memory right after loading a model may lead to running out of capacity during actual processing.

To compare before buying, first decide on the model and quantization format, then input text of a length comparable to what you normally handle. Next, check memory usage and response speed while running the number of concurrent tasks you need. If fitting into 64GB means shortening the context or reducing concurrency, compare not only the price difference but also the constraints this places on your work.

At the October 23 launch, it will be worth checking each manufacturer's final configuration and actual selling price. Pricing and availability for Japan cannot be determined from the announcements reviewed here. If Model Launcher arrives as planned and 64GB can handle real work while sustaining the necessary context length and concurrency, it would widen the option of placing local AI environments at individual sites, separate from a high-capacity development machine.