On October 7, 2026 (U.S. time), NVIDIA showed off its enterprise AI workstation, the "DGX Station for Windows," once again at Microsoft's Windows and Surface event.
It is built around the "GB300 Grace Blackwell Ultra Desktop Superchip" and offers up to 748GB of combined CPU and GPU memory. FP4 compute performance reaches up to 20 PFLOPS, and the system is scheduled to launch in the fourth quarter of 2026.
For developers who have had to move back and forth between a Linux environment for AI development and a Windows environment for business applications, the appeal is that a single machine can handle both.
However, being able to fit a trillion-parameter AI model in memory is a separate question from being able to run that model at a sufficient speed.
Business Windows apps and large-scale AI on one machine
DGX Station for Windows is aimed at companies that want to carry out design, research, and other work on Windows while also running large AI models and always-on AI agents locally.
The previous DGX Station was a Linux-centered system. With the Windows version, Windows business applications and the powerful compute resources of the GB300 can be used on the same machine. NVIDIA also says Linux-based AI development tools can run through the Windows Subsystem for Linux (WSL).
The Windows version was first announced in NVIDIA's announcement of May 31, 2026.
The October 7 event was a renewed showing ahead of launch. It does not mean a new GPU was announced or that sales began immediately. The fourth-quarter 2026 launch target had also been indicated from the start.
Still, the showing is significant.
At the same event, Microsoft laid out a new direction for Windows that uses on-device AI and cloud AI according to the task.
Within that, DGX Station is positioned not only as a high-performance computer for individual researchers and developers, but also as an AI inference platform shared by teams.
For the latter use, Microsoft cited an example of running 32 or more AI agents at the same time.
What matters to enterprises is how far this can reduce the hassle of leaving their usual work environment each time they use AI, and the burden of managing dedicated AI machines through a separate system.
An always-on AI agent does not finish its job simply by generating text. It may operate business applications, process files, and keep working for long periods.
The aim of this product is to make it easier to fold the compute resources and access-permission management these tasks require into existing Windows environments.
748GB of memory, but not all of it is fast GPU memory
At the core of DGX Station for Windows is the GB300, which pairs a 72-core Grace CPU with a Blackwell Ultra GPU over NVLink-C2C.
The maximum capacity of 748GB is the sum of 252GB on the GPU side and 496GB on the CPU side, so not all of it is high-speed GPU memory.
Organizing the Windows version specifications published by NVIDIA gives the following picture.
| Component | Memory capacity | Nominal bandwidth | Main role in large-scale AI |
|---|---|---|---|
| GPU-side HBM3e | 252GB | 7.1TB/s | Holds model weights and computation data the GPU accesses at high speed |
| CPU-side LPDDR5X | 496GB | 396GB/s | Stores model weights and data that do not fit in GPU memory |
| NVLink-C2C connecting CPU and GPU | Not included in memory capacity | 900GB/s | Link for exchanging data between the CPU and GPU |
Note: Figures are nominal values based on NVIDIA's official specifications as of October 8, 2026. Memory bandwidth and CPU-to-GPU interconnect bandwidth are different performance metrics and cannot simply be added together.
What NVIDIA calls "coherent memory" is a mechanism that lets the CPU and GPU use memory while maintaining data consistency.
Because the large CPU-side memory can be used in addition to the GPU side, it becomes possible to handle large AI models that do not fit in GPU memory alone.
However, memory bandwidth differs greatly between the GPU side and the CPU side, and NVLink-C2C, which connects them, also has an upper limit on transfer speed.
Being able to use 748GB of memory does not mean all of it can be accessed at 7.1TB/s.
The time needed for computation and data transfer changes depending on which parts of the model are placed in GPU-side HBM3e and which in CPU-side LPDDR5X.
The 748GB capacity widens the scale of AI models that can be used, but judging actual inference speed requires considering memory placement and bandwidth as well.
The 20 PFLOPS maximum is peak FP4 performance
The headline figure of up to 20 PFLOPS also comes with conditions.
This number is the theoretical peak when running Tensor Core operations at FP4 precision with sparsity.
Sparsity refers to a mechanism that uses a specific data structure to skip some operations and improve computational efficiency.
NVIDIA's specification table lists FP4 performance without sparsity at 15 PFLOPS, FP16/BF16 Tensor Core performance at 5 PFLOPS, and FP32 performance at 80 TFLOPS.
Unless otherwise noted, Tensor Core figures include sparsity, and peak performance assumes the GPU's boost clock.
FP4 is a low-precision data format that represents numbers in 4 bits. It is used to reduce the memory capacity and computation needed when running large AI models.
However, the 20 PFLOPS at FP4 cannot be evaluated as-is as performance for high-precision scientific and technical computing.
Nor is it a number that can be simply converted into processing speed for AI inference such as text generation.
Comparing real-world performance requires aligning the numerical precision and computation methods the model uses.
Running a trillion-parameter model requires lower precision and careful use of memory
NVIDIA says DGX Station for Windows can handle AI models of up to one trillion parameters.
So how much memory does a trillion-parameter model actually need?
A model's parameters are the numerical values adjusted through training. In large models, just storing these parameters consumes a vast amount of memory.
For example, if one trillion parameters are each stored in 4 bits, the required capacity is 500GB.
Including the block-wise scale factors used in NVIDIA's NVFP4 format, the calculated capacity becomes 562.5GB.
According to NVIDIA's technical explanation of NVFP4, NVFP4 shares one 8-bit scale factor across every 16 4-bit values.
The average storage per parameter is therefore:
4 bits + 8 bits ÷ 16 = 4.5 bits
Converted to one trillion parameters, the required capacity is as follows.
| Storage format for 1 trillion parameters | Calculated capacity | Data included |
|---|---|---|
| 4-bit values only | 500GB | Parameter values |
| With NVFP4 block scaling | 562.5GB | Parameter values plus an 8-bit scale factor for every 16 values |
These are theoretical figures based on the maximum model size NVIDIA has published and the NVFP4 specification, assuming all parameters are stored in the same format.
They do not include additional per-tensor scale factors, layers that are not quantized, or working space needed at runtime.
Memory used by the OS and the KV cache required during inference must also be reserved separately. Actual model file sizes and memory use while running therefore will not match this calculation exactly.
What stands out is that even when stored in a 4-bit format, the weights of a trillion-parameter model far exceed the 252GB on the GPU side.
In other words, keeping a trillion-parameter model entirely in memory requires using CPU-side memory as well.
DGX Station for Windows is designed to handle trillion-parameter-class models by combining large CPU memory with low-precision techniques.
The entire trillion-parameter model cannot fit in high-speed HBM3e alone.
The KV cache affects inference speed and the number of concurrent users
When running a large AI model, the memory needed for the KV cache must be considered in addition to the storage capacity for parameters.
The KV cache stores intermediate computation results for previously processed tokens and reuses them when generating the next token.
This removes the need to repeat the same computation and makes inference more efficient.
On the other hand, the KV cache itself consumes memory and uses memory bandwidth when accessed.
Loading long documents or handling requests from multiple users at once increases the capacity the KV cache requires.
Therefore, even if the model weights fit in 748GB of memory, there is no guarantee that enough headroom will remain for long contexts or many simultaneous requests.
It is also necessary to distinguish between being able to use a trillion-parameter model for inference and being able to train a model of that size from scratch.
In AI training, in addition to parameters, gradients and the state needed for parameter updates must also be kept.
Inference, fine-tuning, and pre-training each require different amounts of memory and computation.
It would not be appropriate to interpret the "trillion-parameter support" claimed for DGX Station for Windows as the ability to train a model of that scale from the beginning.
For enterprise adoption, agent management and Windows app compatibility are also issues
To use DGX Station for Windows in a company, it matters not only how much AI compute it offers but also whether the agent execution environment and access permissions can be managed securely.
On October 7, Microsoft announced general availability of "Microsoft Execution Containers (MXC)," which run AI agents in isolated environments.
MXC is an OS-level mechanism that supports organization-wide management by isolating the execution environment of each agent and identifying which agent performed which operation.
Microsoft is also promoting integration with Agent 365 and Intune. It notes, however, that using management features in large-scale environments may require additional services.
NVIDIA, for its part, plans to support "OpenShell" on DGX Station for Windows.
OpenShell is also a mechanism that gives each AI agent an isolated execution environment, and its distinguishing design separates application operations from policy enforcement on the infrastructure side.
This makes it possible to restrict agent actions in the execution environment itself, rather than relying only on instructions given to the AI model.
However, the general availability of MXC announced by Microsoft and the OpenShell support NVIDIA plans for DGX Station for Windows are at different stages of availability.
The Grace CPU is Arm-based, so compatibility with existing Windows apps needs attention
Another point to check is compatibility with Windows apps and peripherals.
The Grace CPU in DGX Station for Windows uses the Arm architecture.
So even though Windows runs on it, software cannot necessarily be used under exactly the same conditions as on a workstation with a conventional x86/x64 processor.
Microsoft provides emulation for running x86/x64 apps on Windows on Arm.
However, emulation carries additional processing overhead, so some apps may see a performance impact.
Kernel drivers and user-mode printer drivers also need native Arm64 versions.
Companies adopting DGX Station for Windows will need to verify operation in an Arm environment, including the design software and research applications they use, as well as peripheral drivers.
Up to 1,600W of power, so the installation environment matters too
DGX Station for Windows cannot necessarily be installed the way an ordinary desktop PC can.
NVIDIA has published a total system power of 1,600W and states that a 20A power circuit is required.
This is not a measurement showing that the system always draws 1,600W for every task.
Still, installing it near an office desk calls for attention to power supply and heat dissipation.
Configurable options, such as visualization features using additional RTX PRO GPUs, also need to be confirmed with each OEM that offers the product.
How practical is "running 32 or more AI agents at once"?
In this announcement, Microsoft presented DGX Station as a platform for running 32 or more AI agents simultaneously.
However, being able to run multiple AI agents at the same time is a separate matter from each one completing its work at a sufficient speed.
For example, when 32 agents send requests to a model at once, response wait times change depending on the model's size, the context length, and the processing involved.
To judge whether performance is adequate for business use, you need to check at least the following metrics:
- Time to first response: The time from sending a request until the first token is generated
- Sustained generation speed: The number of tokens the model can generate per unit of time
- Performance under concurrency: How much response time and throughput change when multiple agents are running
Although Microsoft's announcement presents use cases running 32 or more agents at once, it has not published specific test conditions for DGX Station or measurement results combining these metrics.
For actual purchasing decisions, you need to align the AI model and context length and measure performance as the number of concurrent agents increases.
Simply being able to launch many agents is not enough to evaluate practicality.
What matters is whether each agent can finish its work within an acceptable wait time even when several are running at once.
Cost-effectiveness compared with cloud APIs also matters
When a company runs large-scale AI locally, a cost comparison with cloud APIs is essential.
Deploying DGX Station makes it possible to run AI models on local compute resources without depending on external cloud APIs.
However, buying the machine does not make AI inference free.
While usage-based cloud API charges decrease, hardware purchase costs, electricity, and maintenance and operating costs are incurred instead.
Cost-effectiveness therefore depends on how much cloud API usage can be replaced with local inference.
The comparison needs to include not only current cloud API fees but also the performance of the required models on DGX Station and ongoing operating costs.
NVIDIA's announcement and product page this time have not disclosed a unified selling price for the Windows version or a confirmed shipping date for Japan.
It is therefore difficult at this point to assume a product price and calculate in concrete terms how many years it would take to recoup the investment.
The real value of DGX Station for Windows will be decided by actual AI inference performance
DGX Station for Windows is a product that offers up to 748GB of memory and high AI compute performance, enabling large AI models to be used in a Windows environment.
Being able to handle Linux-based AI development tools and Windows business applications on one machine could be a major advantage for companies bringing AI into research and development and design work.
It is also expected to serve as a shared platform for keeping multiple AI agents running continuously.
On the other hand, the maximum model scale of one trillion parameters and the FP4 peak of 20 PFLOPS alone cannot tell you how usable the machine will be in practice.
The 748GB of memory is split between the CPU and GPU sides, with large differences in access speed. The larger the model, the more the placement of weights and the KV cache affects inference performance.
There are also many points to check before enterprise deployment, including application compatibility on Windows on Arm, power and cooling facilities, and AI agent management features.
Once the actual hardware launches, planned for the fourth quarter of 2026, it will be important to run the AI models and business applications companies actually use and to verify response speed when multiple agents run at once.
The value of DGX Station for Windows will be determined not only by the specification of being able to run a trillion-parameter model, but by whether that large-scale AI can be used in daily work with sufficient speed and stability.
Whether the large memory and high compute performance can be turned into practical AI processing on Windows will be the focus going forward.
