On October 7, 2026 (US time), Microsoft added an experimental feature to Windows ML, its local AI runtime for Windows, that lets apps run GGUF-format AI models through llama.cpp. The official announcement also introduces new APIs for text generation and speech recognition, along with the Windows-native "Runtime API" that underpins them. Windows ML had centered on the ONNX format; now publicly available GGUF models can also be built into apps. However, being able to call a model through a common API is not the same as being able to use every feature on the same hardware. Understanding this change requires looking at how each model format is executed and which hardware it supports.

AD

Windows ML Adds GGUF Support, Letting Apps Embed llama.cpp Models

The newly added Windows ML "Text Generation API" supports language models in both ONNX and GGUF formats, and it selects the appropriate inference engine for each format.

For GGUF, it uses llama.cpp. GGUF models are not converted to ONNX before running.

Windows ML has traditionally combined ONNX, a common model format, with "Execution Providers" tailored to each vendor's hardware.

An execution provider is software that runs an AI model's computations on hardware such as a CPU, GPU, or NPU.

Because Windows manages the acquisition and updating of execution providers, developers are spared from bundling all the software needed for each type of hardware into their apps.

The existing ONNX Runtime API remains supported and is offered alongside the new Runtime API.

With GGUF support, developers can more easily bring language models they have been trying out in llama.cpp into apps that use Windows ML.

The new Text Generation API handles tasks such as converting input text into tokens the model can process, streaming output that returns results as they are generated, and canceling generation.

Choosing and distributing the model, however, remains the developer's job.

Tokenizer handling also differs by format. With ONNX, a tokenizer matching the model must be prepared separately, whereas with GGUF the tokenizer is included in the model file, so no separate preparation is needed.

To make prototyping easier, a local endpoint compatible with the OpenAI API has also been added.

In Microsoft's example, a GGUF model is launched with "WinMLServer," and the OpenAI SDK's connection target is changed to a local address on the PC.

This lets developers try local AI models using calling conventions familiar from the cloud-based OpenAI API.

However, API compatibility in how calls are made does not mean the performance or features of OpenAI's cloud models are available.

The other new feature is a speech recognition API that can use Whisper models supplied by the developer.

This API converts speech to text by combining an ONNX-format Whisper encoder, decoder, and tokenizer.

Passing the recognized text to a GGUF language model makes it possible to build apps that combine voice input with text generation.

Unlike the existing Windows AI APIs, where Microsoft manages the models, developers can freely choose which models to use.

In exchange, they are responsible for distributing and updating those models themselves.

GGUF Support Does Not Mean NPU Execution

The newly added Runtime API and the text generation and speech recognition APIs are all experimental.

The Runtime API documentation from Microsoft states that production use is not supported and that apps using these APIs should not be published to the Microsoft Store.

This does not mean Windows ML as a whole is unsuitable for production, however. The existing mechanism that uses ONNX Runtime remains officially supported.

The key caveat for GGUF support is the hardware difference.

The execution environments currently validated for GGUF in Windows ML are CPU and NVIDIA CUDA; NPUs are not supported.

While Windows ML as a whole can use NPUs, the newly added GGUF execution does not run on the NPUs found in devices such as Copilot+ PCs.

Based on Microsoft's Runtime API documentation and the description of the llama.cpp package for Windows ML, the differences by model format are as follows.

モデル形式 実行エンジン 対応するハードウェア 実行に必要なソフトウェア
ONNX/ORT ONNX Runtime CPU、GPU、NPU。ただし、モデルや実行プロバイダーが対象機器に対応している必要がある Windows MLが実行プロバイダーの取得・更新を管理
GGUF llama.cpp 今回検証されているのはCPUとNVIDIA CUDA。NPUには非対応 CPU用モジュールと、必要に応じてCUDA用モジュールをアプリ側で用意

Note: This table is compiled from official documentation on the experimental Runtime API as of October 8, 2026. It is not a speed comparison, nor does it list all hardware supported by llama.cpp itself.

Even for ONNX, NPU support does not mean every model runs as-is.

In practice, the operations a model uses and the level of execution provider support determine which hardware can be used.

To use CUDA with GGUF, a compatible NVIDIA GPU and driver are required.

The llama.cpp package for Windows ML is described as running on the CPU without CUDA when no compatible GPU is found.

In other words, even when models can be handled through a common API, the way inference engines are distributed is not unified.

Developers embedding GGUF models in apps must decide how to distribute the CPU and CUDA modules and what hardware to require of target PCs.

This also affects an app's distribution size and system requirements.

Hardware constraints remain for ONNX as well.

According to the Windows ML overview, execution providers optimized for NPUs and certain GPUs require Windows 11 24H2 or later.

The arrival of new APIs does not mean everything works the same way on every Windows PC.

When developing local AI apps, it is important to check not only the model format but also the combination of inference engine and target hardware.

AD

Not All llama.cpp Acceleration Features Are Available Through the New API

In the announcement, Microsoft said it has worked with NVIDIA and the open-source community to improve llama.cpp performance.

Specific examples included kernel fusion, which combines multiple GPU operations into one, and optimizing the ordering of CPU and GPU work.

Support is also in progress for techniques that speed up text generation, such as Eagle-3, MTP, and D-Flash2.

These are primarily improvements to llama.cpp itself, however, and do not mean all of them can be used from Windows ML's Text Generation API.

The current Text Generation API specification shows that many constraints remain in controlling text generation.

For example, the only generation method currently supported is greedy decoding.

This is a method that, when generating text, always selects the most probable candidate as the next token.

Features that adjust the diversity of generated text, such as temperature, top-p, and top-k, which are common in language model APIs, are not yet supported.

Speculative decoding and MTP (Multi-Token Prediction) are also unavailable in the Text Generation API.

Speculative decoding speeds up text generation by looking ahead at the tokens to be generated. MTP likewise aims to make generation more efficient by predicting multiple tokens.

Even if llama.cpp supports such features, apps cannot specify them directly unless they are exposed through the Text Generation API.

In other words, the features an inference engine has and the features available from the API built on top of it do not necessarily match.

Similar limits apply to chat templates and structured output.

Chat templates format user and assistant messages in the form the language model expects.

Structured output is a feature that controls results so they are returned in a fixed format such as JSON.

The current Text Generation API does not support these either.

As a result, while simply loading a model and generating text may work, embedding it in business apps that need complex dialogue features or strict JSON output may require additional implementation.

There are also limits on supported programming languages.

The current Runtime API and Text Generation API support C++ and Python, but not C#.

ONNX text generation models must also have the input and output structure the API expects.

Specifically, in addition to input token IDs and output logits, they need a configuration for passing a KV cache that holds past computation results and state related to the current generation position.

Simply preparing an ONNX model file is therefore not necessarily enough.

You need to check in advance whether the model you want to use matches the structure the new Text Generation API requires.

Develop with PyTorch and Triton, Distribute to Apps via ONNX

Beyond GGUF and llama.cpp support, the announcement also covered progress in PyTorch and Triton, which support AI model development on Windows.

However, native CPU builds of PyTorch for Windows Arm64 were not offered for the first time here.

Microsoft already announced that native builds for Windows Arm64 became available with PyTorch 2.7 in April 2025.

In addition to these CPU builds, the latest explanation also covered CUDA-enabled Windows Arm64 packages that NVIDIA provides for supported hardware.

That does not mean CUDA is available on every Arm-based Windows PC.

Using CUDA requires the necessary hardware and software, such as a compatible NVIDIA GPU.

In AI model development, PyTorch handles training and additional training, while Triton is used to generate code for running computations efficiently on GPUs.

In Microsoft's example, PyTorch's torch.compile is used to consolidate multiple operations into GPU kernels generated by Triton.

This can reduce the overhead of running many small operations individually and may speed up the model's computation.

Speeding up a model in the development environment, however, is a separate matter from achieving the same performance once the model is embedded in an app.

In Microsoft's example, the developed model is exported to ONNX format so it can run on Windows ML.

What gets exported to ONNX is the graph representing the model's computation, not the GPU kernels that Triton generated.

When the app runs, Windows ML's execution provider processes the ONNX computation graph on the target hardware.

That means even if the PyTorch and Triton development environment delivers high performance, simply converting to ONNX does not guarantee the same speedup on other GPUs or NPUs.

Optimization for the target hardware and actual testing are needed.

The "Windows ML CLI" helps with that preparation.

The Windows ML CLI provides features for analyzing a model and performing optimization, quantization, and compilation.

It aims for efficient execution in the app by reshaping the model's computation graph to suit the target hardware and compiling it for the corresponding execution provider.

Finally, you can also run benchmarks to check actual performance.

In Microsoft's development example, an x86_64 Python environment for using the Windows ML CLI is set up separately from the Arm64 PyTorch environment.

Even when everything from model development to distribution preparation can be done on Windows, not every tool can necessarily run in the same environment.

Triton also has constraints on supported GPUs and software versions.

In the compatibility table for Triton on Windows, the older NVIDIA GPU architectures Turing and Volta are supported in Triton 3.2, but support ends with 3.3 and later.

There are also compatibility conditions between PyTorch and Triton versions.

Smoothly moving from AI model development to embedding in a Windows app therefore requires checking not only supported hardware but also the combination of frameworks and development tools.

AD

Choosing Models and Runtimes Becomes Key to Practical Local AI

The newly added Runtime API lets developers combine multiple AI models and specify execution targets such as CPU, GPU, or NPU for each stage of processing.

Microsoft also introduced mechanisms for efficiently passing Windows image and audio buffers to models, and a feature that shortens startup time by compiling models in advance.

In an app that converts input audio to text with a speech recognition model and passes the result to a language model to generate a response, for example, being able to specify which processing runs on which hardware is important.

Choosing execution targets according to the workload may improve performance and power efficiency.

However, as noted above, the newly added GGUF execution is limited to CPU and NVIDIA CUDA.

Being able to select an NPU in the Runtime API as a whole is not the same as being able to run GGUF models on an NPU.

Developers already building and running production apps on ONNX Runtime could keep their existing environment and try the new Runtime API only where it is needed.

For developers who want to prototype Windows-only apps that use publicly available GGUF models, the new features offer a new option.

On the other hand, those who need C# development, NPU execution of GGUF models, or fine-grained generation control such as temperature may find the current experimental API insufficient.

Local AI has advantages such as running processing without sending data to the cloud and avoiding the per-token usage fees incurred by cloud inference.

At the same time, it adds items developers must manage, including model distribution and updates, securing the necessary memory, and processing speed on target PCs.

Because CPU, GPU, and NPU configurations vary widely among Windows PCs, the same AI model can differ in which machines can run it and in the performance it achieves.

This Windows ML expansion increases the options for bringing publicly available GGUF models into Windows apps. But a common API does not eliminate the constraints imposed by model format and hardware.

What matters most is that llama.cpp and GGUF have been added as a new option while the existing ONNX-centered runtime remains in place.

If the new Runtime API moves to general availability and its text generation features and supported hardware expand, the burden on developers bringing various public models into Windows apps could ease further.

In actual use, though, startup time, memory usage, and inference speed on target PCs still need to be checked.

Being able to try GGUF models easily is a step forward, but turning that into an app that runs reliably on a wide range of Windows PCs will continue to demand careful validation, from model selection to runtime optimization.