Copilot now supports local inference

Microsoft announced that GitHub Copilot will support local model inference by the end of this month, allowing developers to automatically orchestrate or manually switch between the cloud and their devices. The real change is not merely the addition of another model option, but that Copilot is beginning to incorporate the location of inference into its scheduling.
Copilot Reasoning No Longer Has to Happen in the Cloud
Microsoft announced on October 7, San Francisco local time, that GitHub Copilot will support local AI model inference by the end of this month. After the update, developers will be able to let Copilot switch automatically between cloud-based and on-device models, or explicitly require the use of a local model. The relevant settings will be available in GitHub Copilot CLI, the Copilot app, and Visual Studio Code.
This is not simply an expansion of the model catalog. Previously, Copilot requests were handled primarily by cloud-hosted models, with a backend orchestrator routing them based on factors such as performance, cost, and accuracy. What Microsoft is expanding this time is the scope of the orchestrator's decisions: it will determine not only which model to call, but also where inference should take place.

Microsoft offers two basic usage modes: automatic orchestration or forced use of an on-device model. Developers can also specify a model provider, model name, or endpoint. Local models can be accessed through Windows ML or connected to local endpoints that expose OpenAI-compatible interfaces. One example provided by Microsoft is running MAI Code 1.1 Flash through Windows ML.
The "automatic switching" feature is worth watching, but it should not be overinterpreted. Microsoft has announced the direction in which cloud and local models can be coordinated by the backend, but the information disclosed so far does not fully explain the scheduling strategy. For example, it does not specify which requests will be sent locally, how task complexity will be assessed, whether network conditions will factor into the decision, or whether users will be able to audit or override each routing decision. For developers, the eventual experience will depend on these implementation details, not merely on the addition of a switch to the settings page.
The Value of Local Models Goes Beyond Saving Network Traffic
Moving inference to the device most directly reduces dependence on network round trips. Short requests such as code completions are highly latency-sensitive: when the user stops typing, the interaction feels more fluid if the model can begin generating output more quickly. For scenarios involving unstable network conditions or offline work, local inference also provides an additional option.
More importantly, it changes the data boundary. When sending code to a cloud model, developers need to consider how repository contents, prompts, and context are handled, as well as whether team or customer data policies permit it. If a request is genuinely completed locally, sensitive code has a chance to remain on the device. However, "supporting local models" does not automatically mean that the entire Copilot workflow stays off the device: the extension may still need a network connection for login, synchronization, model retrieval, or calls to other services. Before incorporating local mode into a compliance plan, teams should still verify what data is sent where, how logs are stored, and how the specific models and endpoints behave.
Local inference is also not an unconditional replacement for cloud inference. On-device models are constrained by video memory or unified memory, memory bandwidth, cooling, and power consumption, while cloud infrastructure can provide larger models and more computing power. A lightweight request may be faster and more suitable locally, but for tasks such as cross-file refactoring, complex troubleshooting, or work requiring stronger reasoning capabilities, a local model may not offer an advantage in either quality or speed. The purpose of automatic orchestration is precisely to try to balance these two types of resources. Whether it can do so effectively will depend on routing performance in real tasks, rather than peak figures measured on a demonstration machine.
MAI Code 1.1 Flash: Parameter Count Is Not the Same as Device Cost
Microsoft is also launching MAI Code 1.1 Flash. According to Microsoft's published information, it is a mixture-of-experts (MoE) model with 137 billion total parameters and 6.8 billion active parameters. It uses quantization and speculative decoding to optimize response speed and memory usage.
An MoE model can be understood as a group of experts sharing the same system: when processing each input, it does not need to involve all parameters in computation, activating only a subset instead. Therefore, the total parameter count does not directly represent the amount of computation required for every token. However, this does not mean that a device only needs to provision resources for 6.8 billion parameters. The model weights still occupy storage, while the inference process also needs memory for context, caches, and the runtime. Quantization compresses the weights, while speculative decoding attempts to quickly generate candidate tokens first and then have the target model verify them, reducing wait time. The two techniques optimize different stages, and their actual benefits depend on the combination of model, hardware, and task.
Microsoft's published tests were conducted on a Surface Laptop Ultra. As the prompt length increased from 2K to 256K tokens, decoding throughput ranged from 40 to 63 tokens per second. This figure provides a performance reference for local execution, but it cannot be treated as a speed guarantee for all Copilot users. Throughput is affected by device configuration, prompt length, output content, and testing methodology. Moreover, generation speed is only one part of the experience; time to first token, response quality, power consumption during sustained operation, and memory usage are equally important.
In particular, the 256K-token long-context figure can easily lead people to equate "the model supports long prompts" with "it can smoothly handle long contexts in real development tasks." Long contexts increase cache and memory pressure and may also slow generation. Regardless of how large the codebase is, what matters is whether the system can select the relevant files, retain the necessary context, and provide reliable suggestions within limited resources. For IDE assistants, context management is often just as important as model parameters.
Automatic Orchestration Addresses the Cost of Choosing
Letting users select models manually is flexible, but it also adds to their burden. Developers must determine which tasks are suitable for local execution and which are worth sending to the cloud, while also understanding the speed, capabilities, and resource requirements of different models. Automatic orchestration aims to delegate this judgment to the product: simple, latency-sensitive, or data-sensitive tasks should be considered for on-device execution first, while requests requiring stronger capabilities can then be considered for cloud execution.
This logic appears reasonable, but the difficulty lies in what "automatic" is actually based on. If the scheduler is too conservative, most requests will still go to the cloud, leaving local mode as merely a backup entry point. If it is too aggressive, complex tasks may be assigned to an underpowered on-device model, resulting in lower quality or longer wait times. For teams, model selection also involves policy: some repositories may be allowed to use the cloud, while others must remain local; some developers may choose their own endpoints, while others need centralized management by administrators. Microsoft has not provided the full policy details in this announcement, so the follow-up documentation released when the feature launches will be worth watching.
Accordingly, when evaluating this update, development teams should observe at least four separate aspects: whether routing is predictable, whether users can override automatic selection, whether model and endpoint configuration is suitable for centralized team management, and where requests and code context actually flow. The mere availability of "local inference" is not enough to determine whether it meets an organization's security and compliance requirements.
What This Means for Developers
For individual developers, this update provides a practical workflow option: reduce cloud dependence for tasks that local models can handle adequately while retaining access to cloud capabilities when stronger models are needed. The local option also has real value for people who frequently work in environments with restricted network access. As for whether it can reduce costs, that depends on the purchase and operating costs of local hardware, Copilot's specific billing rules, and which requests ultimately still call cloud services. The information currently available does not support the conclusion that "local is cheaper."
For enterprises and open-source maintainers, the key question is not whether they can run a model, but whether this capability can integrate with their existing permissions, endpoint, and data-governance processes. Support for open, OpenAI-compatible local endpoints makes it easier to connect different providers, but interface compatibility only addresses the calling convention. It does not guarantee that model quality, tool-calling capabilities, or security policies will be fully consistent. Adding an endpoint to the configuration does not mean that all models can be substituted seamlessly.
Overall, Copilot's addition of local inference is a product update with a clear direction, but its success will depend on implementation. It advances the question of "which model should be used" by one step and begins to answer "where should the model run?" This has practical significance for privacy, latency, and offline workflows. However, the differences in capabilities and resource requirements between local and cloud models will not disappear simply because they share a unified entry point. What is truly worth watching is whether, after the feature launches at the end of this month, automatic orchestration can make explainable choices that balance speed, quality, data boundaries, and controllability.



