Kimi K3.1 spotted: another upgrade to its million-token context window?

The Moonshot AI API platform has revealed the `kimi-k3-1` model identifier. The official Open Platform subsequently teased K3.1 or K3 11111 and indicated support for up to a 1-million-token context window. Three levels of reasoning intensity, along with Agent and Swarm multi-agent modes, may also launch with the new model.
Kimi K3.1 Identifier Surfaces, Moonshot AI May Be Preparing Another Flagship Model
Moonshot AI has not officially announced Kimi K3.1 yet, but developers have already spotted it on the API platform.
On September 28, multiple developers discovered that the identifier kimi-k3-1 had appeared in the model registry of Moonshot AI’s API platform. More importantly, the identifier had already passed API probing and was available for invocation. Later that day, the official Kimi Open Platform also published related teasers for K3.1 or K3 11111, with the page indicating support for context windows of up to 1M tokens.
This suggests that K3.1 has most likely entered the final preparation stage before release. Based on the information currently available, its official launch may be imminent, possibly as early as October. However, until the company officially announces the model name, weights, API documentation, and pricing details, developers should not treat the backend identifier as an official version ready for use.

Three Reasoning Levels: The Key May Not Be the Model Name
Among the information revealed so far, the most noteworthy detail is not the K3.1 minor version number, but the three reasoning levels: Low, High, and Max.
This type of design is usually more than a simple switch between response styles. It more likely corresponds to different reasoning budgets, computing resources, and service prices. The Low level would be suitable for classification, extraction, rewriting, and simple code generation; High might be intended for multi-step analysis, complex debugging, and longer tool-calling chains; Max would target tasks requiring continuous planning and repeated verification.
For developers, this tiered approach is more practical than a single thinking toggle. In real-world applications, using the highest reasoning level for every request would rapidly increase costs and latency, while using the lowest level for everything would sacrifice completion rates on complex tasks. If Kimi K3.1 can differentiate the capabilities and prices of its various tiers, applications could route requests according to task requirements instead of having one expensive model handle everything.
But this also raises an important question that needs to be clarified: Are the three levels the same model operating with different compute budgets, or do they correspond to three different service configurations? If they are merely parameter switches, the differences may mainly be reflected in response time and reasoning depth. If they rely on separate computing pools, pricing, concurrency, and rate-limiting policies will become the aspects developers truly care about.
One Million Tokens Is Not Kimi’s First Attempt
The leaked information also suggests that K3.1 may offer an ultra-long context option called Extra Long, supporting up to 1 million tokens.
This capability is not, by itself, an exclusive selling point of K3.1. Kimi K3, released by Moonshot AI in July, already supports a 1-million-token context window and positions long-horizon programming, knowledge work, and deep reasoning as its primary application areas. Therefore, if K3.1 merely retains a 1M context window, its true upgrade will not be about “how much content it can fit,” but about “whether it can continue to use that content correctly after it has been loaded.”
The aspect of long-context models that is most easily misunderstood is that window size does not equal effective memory. Placing a large codebase, hundreds of contracts, or tens of thousands of pages of research materials into the context only solves the visibility problem. Whether the model can accurately locate relevant information, maintain consistent cross-section references, and avoid losing constraints over multiple rounds of operation determines whether it is truly suitable for production environments.
The API documentation for Kimi K3 already provides an automatic caching mechanism: when the request prefix is sufficiently long and remains stable, subsequent requests can reuse the cache and reduce the cost of repeated inputs. If K3.1 continues this design and optimizes long-context retrieval, cache hits, and tool calls, its value for codebase analysis, legal contract review, long-term research, and complex project planning would be fairly clear.
Conversely, if the 1M context is merely a marketing figure, while the model remains unstable at long-distance information retrieval, recalling information from the middle of the context, or performing consecutive tool calls, developers will ultimately continue breaking tasks into smaller 128K or 256K chunks. A larger window does not mean that application architectures can simply eliminate retrieval, summarization, and state management.
Agents and Swarms: Competition Shifts from Answers to Execution
The internal configurations currently circulating also include task modes such as Agent, Swarm, search, and batch processing.
Agent mode is nothing new. Kimi K3 and Kimi Code have already established a product foundation in long-horizon programming, tool calls, and sub-agent tasks. What is truly worth watching is multi-agent collaboration through Swarm. If it eventually becomes available through the API rather than existing only in official products, Moonshot AI would be offering more than a text-generation interface: it would be providing execution orchestration capabilities for complex tasks.
A typical workflow might look like this: The primary agent first breaks down the requirements, then assigns different sub-agents to handle code localization, test generation, documentation retrieval, and result review, before the primary agent consolidates the results and submits the modifications. For large codebases, deep research, and analysis of multi-source materials, this kind of parallel execution has a better chance of reducing completion time than having one model perform every step sequentially.
However, multi-agent systems can also easily turn “one call into ten calls.” Each sub-agent consumes input and output tokens, and intermediate results must be passed between tasks. Without clear stopping conditions, shared state, and error rollback mechanisms, Swarm could become merely a more complicated and expensive workflow. What developers most need to examine is not a demo video, but the following API details:
- Are sub-agents hosted by the platform, or do developers need to maintain their own orchestrator?
- Can multiple agents run concurrently, and do they have independent contexts and tool permissions?
- Are intermediate results included in input and output charges?
- After a task fails, can it be retried, paused, or resumed from a particular step?
- Do tools such as search, batch processing, and code execution provide a unified calling protocol?
These factors will directly determine whether K3.1 is “a model that is better at getting things done” or “a collection of model capabilities that developers must assemble themselves.”
Pricing Is Temporarily the Same as K3—Don’t Treat It as the Final Price Yet
The Kimi Open Platform currently shows that K3.1’s API pricing is the same as that of the existing K3. However, this price may simply be placeholder configuration before the new model goes live.
According to Kimi K3’s currently published prices, input costs ¥20 per million tokens, output costs ¥100 per million tokens, cache writes cost ¥20 per million tokens, and cache hits cost ¥2 per million tokens. For long-context applications, the true cost usually depends on more than the output price. Developers also need to consider caching strategies, context reuse rates, the number of tool calls, and how many agents a single task needs to pass through.
If K3.1’s Max level uses more reasoning tokens while retaining K3’s uniform pricing, the platform may later adjust the structure through different tiers, concurrency limits, or separate billing. When migrating, developers should not compare only the per-million-token unit price. They should also calculate the total cost and end-to-end latency of a complete task.
At present, connecting to different models through OpenAI-compatible APIs has become a common approach for domestic applications. If Kimi K3.1 eventually offers a standardized API, developers could first test model switching and task routing on an aggregation platform such as OpenAI Hub, then decide whether to bind directly to Moonshot AI’s standalone service. For teams that need to test GPT, Claude, Gemini, DeepSeek, and Kimi simultaneously, the value of a unified interface lies primarily in reducing SDK modification and vendor-switching costs. However, the model’s agent orchestration capabilities, vision support, and caching rules will still require separate adaptation.
The Real Significance of This Update
Based on the information currently available, K3.1 appears more like a product upgrade centered on “long-horizon execution” than a simple expansion of model scale.
Kimi K3 already features flagship-level specifications, including 2.8 trillion parameters, native vision, and a 1-million-token context window. If K3.1 merely continues to improve benchmark scores, it may not significantly change developers’ choices. But if it can truly integrate tiered reasoning, long-context caching, multi-agent collaboration, and batch processing into its API, the focus of competition will shift from model response quality to complex-task completion rates.
This is also where Moonshot AI needs to prove itself:
- Can the 1M context maintain stable recall across real-world codebases and long documents?
- Can Max reasoning deliver a measurable improvement in task success rates?
- Can Swarm reduce the total time required for complex tasks instead of increasing call costs?
- Can developers reliably control Agent mode through the API?
- Is K3.1 sufficiently compatible with K3, and how much code will existing applications need to change during migration?
Before the official release, developers can treat kimi-k3-1 as a noteworthy pre-release signal rather than a production dependency. Only after the model documentation, official endpoints, rate limits, pricing, and evaluation data have all been made public will K3.1 be able to answer a fundamental question: Is it a minor iteration of K3, or a step by Moonshot AI toward becoming an agent platform?
Sources
- IT Home: Moonshot AI’s Most Powerful AI Model: Kimi K3.1 Frontend Identifier Leaked, Expected to Be Released Soon: Reported information about the
kimi-k3-1identifier, the K3.1 teaser, and the million-token context. - Kimi API Open Platform: The official entry point for Kimi models and API services. Pricing, model status, and API capabilities are subject to the final information published on the official page.



