On October 7, 2026, Microsoft and GitHub announced a GitHub Copilot feature that automatically switches between AI models running on a PC and cloud-based models depending on the task. An experimental preview for Windows is scheduled to begin in late October, covering GitHub Copilot CLI, the GitHub Copilot app, and Visual Studio Code (VS Code). Developers will also be able to specify a local model directly, so they can either leave the choice of where to process work to Copilot or make it themselves. Running AI models on a PC was possible before, but what sets this update apart is that Copilot can decide where to process each task. How useful it turns out to be will depend on how quality, memory use, and cloud communication change during long coding sessions.

AD

From manual model selection to automatic local/cloud routing

In VS Code, the BYOK (Bring Your Own Key) feature lets you register your own API keys and endpoints, and then select models running on Ollama or Foundry Local from the chat view. Microsoft already covered development work with local models in its official June 18 post. On top of this existing mechanism, the new update adds a way to easily register models running on Ollama, as well as automatic routing between local and cloud.

Feature What the developer chooses Availability confirmed in official materials
Local connection via VS Code BYOK The endpoint and the model used in chat Available as of the June 18 post
Ollama detection in Copilot CLI Whether to add detected models Announced October 7; available in CLI 1.0.94-0 and later
Local/cloud Auto Whether to leave the choice of where to process to Copilot Experimental preview planned for late October

The table summarizes how each feature works and its availability, based on official Microsoft and GitHub information as of October 9, 2026. June 18 is not the date BYOK was launched. Also, the already available Ollama detection feature and the upcoming Auto are separate features.

In Copilot CLI, the /model command detects compatible models available on a running Ollama instance. Developers check the model's provider and endpoint, then add the model they want to use. However, this does not automatically install Ollama or the models themselves, so some preparation is required. Models must also support tool calling and streaming output, which returns the response in sequence as it is generated.

Auto, by contrast, takes the conversation context and cache state into account and automatically decides whether to process a request with a local or cloud model. HydraFusion, which GitHub announced as a research preview on September 4, went beyond single-model processing: it could switch to a higher-performance model when needed, or have another model verify and correct an answer. With this update, local models running on the PC are added to those options.

For developers who want to specify the processing location themselves, there are two options: using Microsoft's "MAI Code 1.1 Flash" through Windows ML, or connecting to a local OpenAI-compatible API. OpenAI-compatible here means the API specification is compatible; it does not mean inference is sent to OpenAI's cloud. The actual destination is determined by the endpoint the developer configures.

A 53GB model that uses 75.5GB of memory at runtime

The local version of MAI Code 1.1 Flash is a Mixture of Experts (MoE) model with 137 billion total parameters and 6.8 billion active parameters used during inference. MoE activates only the parts of the network, among multiple specialized sub-networks, that are needed for a given input. This reduces computation, but memory is still required to hold the weights of the entire model. So having 6.8 billion active parameters does not necessarily mean it can run in a small amount of memory.

Microsoft explains that it reduced the model size to 53GB through quantization, which represents weights and other data with fewer bits. That is roughly an 80% reduction compared with the cloud version in Bfloat16 format. According to the technical write-up, it uses mixed-precision quantization of about 3.3 bits per weight.

However, a 53GB model does not mean runtime memory use of 53GB. In Microsoft's measurements, peak memory usage reached 75.5GB when processing a 256k-token context on a Surface Laptop Ultra. The size of the model itself needs to be distinguished from the memory actually required during work.

AI agents work by reading source code and receiving the results of tool executions. In the process, the "KV cache," which holds information about tokens already processed, also grows. The OS, editor, and inference software use memory as well. Being able to answer a short question is not the same as being able to see a large repository change through to the end.

The Surface Laptop Ultra used in the measurements has NVIDIA RTX Spark and up to 128GB of unified memory. The CPU and GPU can share the same memory space, but not all of the capacity can be allocated to the AI model. Nor does the 128GB maximum specification indicate the minimum requirement for running this model.

Keeping the model loaded in memory saves the time of loading it for each request. But if conversations grow long and memory contention arises with other apps, response speed may drop. When evaluating local AI, it is necessary to check not only whether it can start, but whether it maintains sufficient performance during extended use in your everyday development environment.

AD

Can coding quality hold up when the model is made smaller?

Even if a model's weights are represented at lower precision through quantization, the code it generates must be accurate. A single wrong identifier or tool call can lead to syntax errors or unintended changes. For that reason, Microsoft evaluates not just model size and processing speed but also whether the model can actually complete coding tasks.

In benchmark results published in its October 7 technical write-up, the impact of quantization on performance differed depending on the metric.

Metric Tasks MAI Code 1.1 Flash Local quantized version
SWE-Bench Verified 500 72.6% 70.80%
Terminal-Bench 2.1 89 62.9% 66.29%

The source is the evaluation results published by Microsoft, not measurements by an independent third party. On SWE-Bench Verified, the quantized version scored slightly lower, while on Terminal-Bench 2.1 it scored higher.

These results suggest that a certain level of coding ability was maintained despite the reduction in size. They do not show, however, that quality equal to the original model will be achieved on every development task, nor that quantization always improves performance.

Another technique used to increase processing speed is speculative decoding. A small model generates candidates first, and the original model verifies them. If candidates are accepted efficiently, generation speed can increase. However, generating and verifying candidates requires additional working memory. In other words, part of the memory saved through quantization is redirected to speeding up processing.

The speed figures also require care about what was actually measured. Microsoft reported 923.5 tokens per second at a 64k-token context and 769.8 tokens per second at 128k. The body of the technical write-up describes these as prompt input processing speeds.

However, the figure captions and footnotes in the same material describe them as decode speeds, so the labeling of the metric is inconsistent. It is therefore not appropriate to interpret these numbers directly as response generation speed in typical usage environments.

The measurements were taken on October 5 in a llama.cpp CUDA environment for Windows ARM64, using speculative decoding with DFlash2. A footnote also explains that they used an artificial code-generation workload, so results will vary with the actual device, settings, and tasks.

The aim of this automatic routing is not to move all processing to the PC, but to choose the model and execution environment best suited to each task. Even if dependence on the cloud decreases, development time will not be shortened if code fixes and rework increase. Evaluation needs to include the quality of the final output, not just processing speed.

Where inference happens is a separate question from communication and tool-execution permissions

Choosing a local model does not make the entire Copilot session offline. In its announcement of the Ollama detection feature, GitHub states explicitly that selecting a local model does not enable offline mode or automatically stop the sending of usage data to GitHub.

In Copilot CLI, setting COPILOT_OFFLINE=true disables connections to GitHub's servers. However, if an external service is specified as the model endpoint, prompt and code information may still be sent to that service even in offline mode.

How far communication can be restricted depends on the settings described in GitHub's official documentation and on the endpoint of the model actually used. This is also a Copilot CLI setting and does not apply to all VS Code features.

The question of where an AI model runs inference should also be kept separate from the question of what operations an AI agent can perform.

On the same day, October 7, GitHub also announced the general availability of local sandboxing. It covers Copilot CLI, the Copilot app, and VS Code sessions that use Agent Host, and is available at no additional cost.

This feature uses Microsoft Execution Containers (MXC) to restrict file access, network connections, and access to credentials for commands the agent runs. Whether the model is local or cloud-based, the need to properly control the agent's execution permissions remains the same.

However, not all tools are isolated in the same way.

When the sandbox is enabled, in addition to shell commands, local MCP servers and language servers are in principle also run inside processes whose access is restricted by the OS.

For file-operation tools built into Copilot, on the other hand, Copilot itself checks operation requests against policy. This differs from a mechanism in which the OS restricts the access of a separate process. Also, remote MCP servers are not within the scope of local sandbox isolation, and restrictions on connections to them are checked on the Copilot side.

For development teams handling highly confidential source code, what matters is not simply seeing a "using a local model" label. What matters is being able to understand what information is sent where and which operations the AI agent is permitted to perform.

Regarding Auto, Microsoft and GitHub explain that it decides where to process based on conversation context and cache state. However, they have not disclosed the specific scope of information sent to the cloud or the detailed conditions that govern routing decisions. When adopting it, teams need to check both where the model runs and the conditions under which data may be sent externally.

AD

The practicality of automatic routing should be judged in real development work

Being able to run AI models on a PC widens the choice of computing resources compared with relying on the cloud alone. However, this announcement does not give a specific cost-reduction rate for the new Auto feature, nor the conditions for all supported devices and plans. The cost savings obtained in the existing HydraFusion experiment also cannot simply be applied as the effect of local-model support.

To test practicality in the late-October preview, one approach would be to prepare equivalent fix tasks in a repository you use regularly and run them with Auto, a cloud model, and a local model, then compare the results.

What to check is not just how quickly the first response comes back. The quality of the final generated code, the total time for the work including fixes and redos, and memory usage over long sessions are what matter. Recording the number of tokens consumed on the cloud side would also make it easier to judge how much burden local models actually relieve.

The advantage of Auto is that Copilot can choose an appropriate execution environment without the developer having to reselect the model or processing location for every task. Where processing locations must be controlled strictly, the ability to specify a model explicitly remains useful.

Ultimately, what matters is whether automatic routing can reduce working time and cloud dependence without lowering development quality. If data destinations and agent execution permissions can also be managed properly, the value of using a PC's computing resources for everyday AI coding grows.