DocsQuick StartAI News
AI News105 Trillion Tokens: Xiaomi MiMo Takes the Top Spot
Industry News

105 Trillion Tokens: Xiaomi MiMo Takes the Top Spot

2026-07-27T06:02:53.764Z
105 Trillion Tokens: Xiaomi MiMo Takes the Top Spot

Xiaomi MiMo-V2.5 processed 10.5 trillion tokens last week, rising to the top of the global model usage rankings. This figure indicates that its adoption among developers is growing rapidly, but it does not mean that it ranks first in model capabilities, user base, or commercial revenue.

10.5 Trillion Tokens: Xiaomi MiMo Takes the Top Spot

Xiaomi MiMo-V2.5 became the world’s most-used AI model by call volume last week.

On July 27, Xu Jieyun, Special Assistant to the Chairman of Xiaomi Group and Deputy General Manager of the Strategic Marketing Department, revealed that all five of the world’s most-called models last week came from China. Among them, Xiaomi MiMo-V2.5 rose to first place, reaching 10.5 trillion Tokens in weekly usage, up 12% week over week.

That is an extraordinary level of throughput. A week contains 604,800 seconds, so 10.5 trillion Tokens works out to an average of roughly 17.36 million Tokens processed per second. Of course, this is merely a mathematical conversion to help illustrate the scale. It does not represent the real-time throughput of a single Xiaomi cluster, nor can it be used to directly infer GPU count, revenue, or actual user numbers.

Even more notable is the pace of growth: MiMo-V2.5 only entered public beta on April 23, yet it reached the top of the usage rankings just over three months later. Its weekly usage was still around 3.94 trillion Tokens in mid-June, but had grown to more than 9 trillion Tokens by mid-July. In the latest week, it crossed the 10 trillion threshold.

Chart showing the growth in Xiaomi MiMo-V2.5’s weekly usage, from its April public beta to surpassing 10.5 trillion Tokens in July

This No. 1 Ranking Is, First and Foremost, a Victory in Developer Distribution

Usage rankings can easily be packaged as a new kind of “model capability leaderboard,” but the two are not the same.

What 10.5 trillion Tokens demonstrates is that MiMo-V2.5 has entered a large number of real or semi-real-world request pipelines, and that developers are willing to keep using it. It reflects a market choice shaped by a combination of price, availability, context length, generation speed, and task fit—not a test score.

In other words, it is more like cloud-computing instance usage than a processor benchmark. The chip with the highest benchmark score does not necessarily sell the most; similarly, the model with the strongest benchmark results does not necessarily handle the most API traffic.

There are at least three direct reasons why MiMo-V2.5 has climbed so quickly.

1. Pricing Adjustments Lowered the Barrier to Migration

Xiaomi previously optimized its Token Plan, and the MiMo-V2.5 family has also undergone substantial price adjustments. In the large-model API market, price is not a secondary metric. It is a core factor determining whether a model can enter the production traffic pool.

For a chatbot, each answer may consume only a few thousand Tokens. But in scenarios involving Agents, code analysis, search summarization, and batch data processing, a single task may involve dozens of tool calls. System prompts, conversation histories, retrieval results, and tool outputs may also be repeatedly included in the input. Even a small difference in per-inference pricing becomes an entirely different cost equation when multiplied across billions of weekly calls.

Lower prices therefore do more than attract users looking for free or discounted access. They allow applications previously limited to smaller models to experiment with longer contexts and more complex reasoning pipelines, while also encouraging developers to gradually promote backup models into primary routing roles.

2. MiMo-V2.5 Targets Precisely the Scenarios That Consume the Most Tokens

Xiaomi positions MiMo-V2.5 as offering stronger reasoning, more reliable Agent performance, longer context, better instruction following, and improved understanding of ambiguous instructions. These capabilities align closely with the application types currently driving the fastest Token consumption.

Traditional Q&A follows a “one question, one answer” pattern. An Agent is more like assigning the model a work order: it plans first, then searches, invokes tools, checks the results, retries after failures, and finally generates an answer. It is not unusual for a single user request to expand into a dozen—or even dozens of—model calls in the background.

Long context also substantially increases input volume. Asking a model to read a code repository, hundreds of pages of documents, or a complete customer-service history can push a single request into the tens or even hundreds of thousands of Tokens. Once MiMo-V2.5 enters batch workloads in these scenarios, usage can grow much faster than the number of users.

Accordingly, 10.5 trillion Tokens does not necessarily mean a comparable influx of new users. It may instead indicate that existing users are entrusting longer and more complex tasks to the model.

3. A Model Family Is More Easily Integrated into Complete Product Pipelines

The MiMo-V2.5 public beta did not consist of a single text model. It included MiMo-V2.5, V2.5-Pro, the V2.5-TTS Series, and V2.5-ASR, covering text, speech synthesis, and speech recognition while emphasizing omni-modal perception and understanding.

This is especially important for Xiaomi. Xiaomi is not merely building a web-based chat interface; it has hardware entry points spanning smartphones, vehicles, smart homes, and IoT devices. If speech recognition, intent understanding, tool invocation, and speech output can all be connected through the same model family, it becomes easier from an engineering perspective to manage latency, cost, and release cycles than when stitching together services from multiple vendors.

However, the figures disclosed so far concern model usage. They do not reveal how much came from Xiaomi’s internal products, third-party developers, or aggregation platforms. Nor can leaderboard traffic be treated as equivalent to model penetration across Xiaomi’s hardware ecosystem.

All of the Global Top Five Are Chinese Models—the Signal Matters More Than the Ranking

Another change in this week’s ranking is that all five of the most-used models came from China.

In the previous week, Tencent Hy3, MiMo-V2.5, DeepSeek-V4-Flash, MiniMax M3, and Zhipu GLM-5.2 already occupied the top five positions. In the latest reporting period, MiMo-V2.5 continued to grow and rose to first place. The precise order may frequently change with free promotions, pricing adjustments, and new releases, but the expanding overall share of Chinese models in high-frequency API use cases is no longer a one-week anomaly.

There is no mystery behind this. As the capability gap between models narrows, competition naturally shifts from “who can build it” to “who can make it affordable and reliable for developers.”

Leading proprietary US models remain highly competitive in complex programming, frontier reasoning, and high-value knowledge work. But for large-scale tasks such as content generation, customer service, translation, information extraction, search summarization, and lightweight Agents, developers do not look only at the upper limit of capability. They care more about which model can complete the same 10,000 tasks at a lower total cost, with more stable latency and fewer failed retries.

The most aggressive advantage of Chinese models at present is their combination of “strong enough” capabilities and aggressive pricing. Competition among MiMo, DeepSeek, MiniMax, GLM, and Tencent’s models is also continuously driving down per-unit inference costs in China’s API market.

Our assessment is that Chinese models taking all five top positions is more noteworthy than any single model briefly reaching No. 1. The former indicates that developer preferences are creating a clustering effect, while the latter may be influenced by promotions, free quotas, or routing strategies.

But This Is Not a “Global AI Market Share Table”

One crucial caveat must be applied to these figures: so-called global model usage rankings generally come from observable third-party model aggregation and routing platforms. They cannot cover all official APIs operated by model vendors, private deployments by cloud providers, on-premises enterprise inference, or internal calls made by consumer products.

For example, the substantial traffic generated by the official ChatGPT, Gemini, and Claude products may not be fully included under the same measurement methodology. Open-source models run by enterprises in dedicated clouds or self-hosted clusters may also be entirely invisible.

“Global No. 1” should therefore be understood more precisely as follows: MiMo-V2.5 ranked first among the model calls covered by this particular public dataset, not that it has surpassed every other model in total global usage.

Tokens are also not a sufficiently standardized unit of sales volume. Different models use different tokenizers, so the same Chinese passage may be divided into different numbers of Tokens. Whether a platform counts cached inputs, reasoning traces, and retry requests can also affect the final result. A model that uses more Tokens to complete the same task may appear to have higher usage, even though its actual efficiency is not necessarily better.

Free models can amplify this distortion in particular. When call costs approach zero, developers may route large volumes of batch jobs, experimental tasks, and even low-value requests to the model. Whether that traffic remains after the free period ends is a better indicator of whether the model has truly entered production environments.

The usage ranking therefore cannot, at a minimum, be used to directly draw the following conclusions:

  • MiMo-V2.5 has the world’s strongest model capabilities;
  • Xiaomi has the largest number of AI users;
  • All 10.5 trillion Tokens came from paid requests;
  • Usage can be converted directly into revenue using listed prices;
  • The ranking covers every vendor’s official channels and private deployments.

What Developers Should Examine: Look Beyond Per-Token Pricing

After MiMo-V2.5 took the top spot, developers may be tempted to immediately replace their production models with a cheaper and more popular new option. But a sound migration decision should not be based solely on comparing prices per million Tokens.

A more meaningful metric is the “cost of completing one valid task.” At a minimum, this includes:

  1. Task success rate: Whether the model can generate a compliant result on the first attempt;
  2. Average retry count: A cheaper model may ultimately cost more if it frequently fails;
  3. Input and output length: Whether it produces unnecessarily long reasoning processes or redundant answers;
  4. P50, P95, and P99 latency: Fast average performance does not guarantee stability during peak periods;
  5. Tool-call reliability: Whether function names, parameter structures, and invocation order remain consistent;
  6. Effective long-context performance: Supporting a long context window does not mean the model can accurately retrieve information buried deep within it;
  7. Structured-output compliance rate: How often it violates JSON, schema, or field constraints;
  8. Rate limits and service availability: Whether large-scale batch processing is prone to congestion.

For teams that already use multi-model architectures, MiMo-V2.5 is better introduced into the candidate routing pool first rather than immediately replacing every primary model. Low-risk tasks such as summarization, classification, extraction, and rewriting can be migrated initially, followed by shadow-traffic testing for coding, Agent, and long-context workloads.

A genuinely meaningful comparison does not involve asking several models the same brainteaser. Instead, enterprises should replay their own historical requests: keep prompts, temperature, maximum output length, and tool environments fixed, then measure success rates, latency, and total Token costs across 1,000 or 10,000 tasks.

If MiMo-V2.5 can maintain the cost-effectiveness implied by the rankings in real-world business environments, its usage will represent more than the outcome of promotions—it will translate into sustained developer retention.

After 10.5 Trillion, Xiaomi Must Prove Retention

From its public beta in April to reaching No. 1 in July, MiMo-V2.5’s growth curve has been remarkably steep. At the very least, it shows that Xiaomi is no longer relying solely on smartphones and vehicles to tell its AI story. It has gained measurable external adoption in the foundation-model API market.

The next set of challenges will be even harder.

First, can Xiaomi establish a sustainable inference-economics model after optimizing prices? Second, can stability under high concurrency keep pace with usage growth? Third, can MiMo-V2.5’s Agent, long-context, and omni-modal capabilities convert leaderboard traffic into long-term production workloads?

Xiaomi Pengcheng has officially announced that it will hold a technology launch event on July 30. Attention will now focus on whether Xiaomi discloses more information about its model technology, infrastructure, and deployment progress. However, before official details are released, this rise to the top of the usage ranking should not be prematurely interpreted as a signal of a new model launch.

This No. 1 ranking is meaningful, but its significance lies primarily in “distribution and adoption,” not in being an “intelligence champion.” As model capabilities gradually become commoditized, persuading developers to spend real Tokens on your model is itself a hard metric. Whether MiMo-V2.5 can remain in production pipelines after subsidies and hype fade will be the real test of its next stage.

References

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: