Huawei Qiyao Tops Both Time-Series Model Rankings

The Ascend-native time-series pre-trained model Qiyao, jointly developed by Huawei and East China Normal University, recently topped both the TIME-leaderboard and GIFT-Eval Pretrained rankings. It is designed not for chatting, but for continuous data forecasting and anomaly detection in scenarios such as power grids, industry, and finance.
Huawei’s Qiyao Tops Two Time-Series Model Leaderboards, Marking a Tough New Phase for Foundation Models
On October 11, Huawei Cloud announced that Qiyao, an Ascend-native time-series pretrained model jointly developed by Huawei and East China Normal University, had taken the top spot on the international TIME-leaderboard and in the Pretrained category of GIFT-Eval.
This is not another industry model release wrapped in the language of large models. Time-series models process continuous data such as electricity loads, equipment vibrations, traffic flows, weather indicators, sales curves, and financial prices. Their output is not a passage of text, but a forecast of what may happen over the next hour, day, or week. They rarely appear in the spotlight of consumer AI, yet they have a direct bearing on energy dispatch, industrial downtime, and inventory decisions.
Huawei says Qiyao ranked first in both international evaluations, indicating that China’s time-series foundation models have entered the global first tier. For now, however, that conclusion rests largely on leaderboard rankings and official disclosures. Public information has not yet provided the model’s full parameter count, training-data composition, individual scores on each dataset, or its error margins relative to other models on the leaderboards. In other words, reaching the top is a signal worth watching, but it cannot yet be taken to mean that Qiyao is comprehensively superior across all real-world applications.

Why Time-Series Models Need Their Own Foundation
Over the past few years, foundation models have focused mainly on text, images, and speech. Time-series data also has a sequential structure, but poses a more difficult challenge: it typically comes from different devices, sampling frequencies, and business systems, with pronounced differences in scale, noise, missing values, and periodicity.
Even when two datasets are both curves, power-grid load may be sampled every 15 minutes, vibration sensors on industrial equipment may generate thousands of data points per second, and financial data may be affected by breaking news and trading rules. A language model can map different sentences into a relatively consistent token space. Time-series models, by contrast, face inputs more like a collection of instrument panels with completely different scales.
Traditional approaches often build a separate model for each task:
- Power companies train one model for load forecasting;
- Factories train another for equipment fault detection;
- Financial institutions build specialized models for different assets and cycles;
- A change in region, equipment, or sampling period usually requires another round of tuning.
This approach can work in a single setting, but it is costly to develop, difficult to transfer, and heavily dependent on domain experts. The goal of a time-series foundation model is to pretrain on large-scale time-series data from multiple domains, allowing the model to learn trends, cycles, sudden changes, seasonality, and relationships between variables. It can then adapt to specific tasks through zero-shot inference or a small amount of fine-tuning.
This can be understood as a shift from “training a separate forecasting engineer for every factory” to “first training a time-series expert who understands general patterns, then quickly familiarizing it with a particular production line.” That is also why Qiyao is described as a time-series pretrained model, rather than simply a model for one specific forecasting task.
Qiyao’s Core Selling Points: Zero-Shot Use, Low-Cost Fine-Tuning, and Ascend-Native Design
According to Huawei Cloud, Qiyao can be used directly for various zero-shot time-series forecasting tasks and also supports low-cost fine-tuning. For developers and businesses, these two capabilities matter more than leaderboard rankings alone.
Zero-shot does not mean the model needs no business context at all. It means there is no need to retrain a complete model for every task. Developers provide historical observations, a time window, and a forecast horizon, and the model attempts to generate the future sequence. For businesses with limited data, frequent changes, or too few accumulated fault examples, this capability can significantly shorten the validation cycle.
For example, a regional power grid may want to forecast its load curve for the next 24 hours. A traditional project might require cleaning years of historical data, selecting a model architecture, engineering holiday and weather features, and running multiple rounds of training. With a foundation model, the process could start by providing standardized historical load data and related variables, then evaluating zero-shot performance before deciding whether to fine-tune the model with local features such as weather, temperature, and holidays.
Low-cost fine-tuning addresses a different need: businesses may want more than generic forecasts; they may want the model to understand their own equipment, production lines, and business rules. If every adaptation requires updating the full set of parameters, compute, storage, and maintenance costs can quickly rise. Parameter-efficient fine-tuning or lightweight adaptation lets businesses update only a smaller set of parameter modules while retaining the foundation model’s general capabilities.
But “low cost” in time-series applications cannot be measured by training memory alone. Real deployments also need to account for data integration, feature engineering, model monitoring, drift detection, rollback, and inference latency. Even if fine-tuning is inexpensive, rewriting the entire preprocessing pipeline every time a data source changes could still outweigh the algorithmic gains.
Huawei’s emphasis on Ascend-native design indicates that Qiyao’s competitive edge is not limited to its model architecture; it also lies in its training and inference pipeline. Time-series forecasting often involves processing large numbers of windows in batches, while production environments may also handle requests across multiple variables, time granularities, and tenants. Operator compatibility, memory utilization, data movement, batch-processing efficiency, and inference stability all affect the final cost.
For government and enterprise customers with Ascend infrastructure, this hardware-software integration may be more attractive than a standalone set of model weights. If the model can be trained and deployed reliably on domestic computing infrastructure, businesses can also gain greater control over their supply chains, data compliance, and long-term operations.
What Topping Both Leaderboards Does and Does Not Tell Us
TIME-leaderboard and GIFT-Eval are both widely followed evaluation frameworks in time-series model research. Qiyao’s leading position on both at least shows that it is competitive across several forecasting tasks covered by public benchmarks, rather than succeeding on just one test set.
The Pretrained category in GIFT-Eval is particularly relevant because it assesses how pretrained models generalize across multiple datasets. This differs from a specialized model repeatedly tuned for a specific dataset and more closely reflects the real test for a foundation model: can it transfer to an unseen data distribution without extensive task-specific training?
But first place on a leaderboard is not the finish line for real-world deployment, for at least three reasons.
First, public datasets differ from production data. Real industrial data often contains sensor drift, misaligned timestamps, equipment replacements, missing intervals, and outliers. Its quality is far more complicated than that of research benchmarks. An error-rate advantage on clean data may not fully translate into gains in a production system.
Second, forecast accuracy is not the only metric. Power-grid dispatch cares about accurately capturing peaks; industrial maintenance cares about how far in advance faults can be detected; financial applications must also account for risk and stability. A model with a low average error may still be unsuitable for critical applications if it fails during extreme events.
Third, deployment cost matters just as much for foundation models. Businesses ultimately need to know how much compute each forecast requires, whether the model supports on-premises deployment, whether latency meets real-time requirements, whether model updates are controllable, and whether errors can be explained and rolled back. Leaderboards generally cannot answer all of these questions.
The right way to interpret Qiyao’s results, then, is that Huawei and East China Normal University have achieved an important milestone in the general capabilities and engineering of time-series foundation models, not that the time-series forecasting problem has been completely solved.
Three Gaps to Close Between Research Leaderboards and Production Systems
1. Standardize Inputs and Outputs
The input to a time-series model is not simply a string of numbers. Production environments need clear definitions for timestamps, variable names, sampling frequencies, missing-value indicators, units, and historical window lengths. Without common standards for data interfaces across systems, it is difficult to reuse model capabilities at scale.
A genuinely usable product should package resampling, normalization, outlier handling, and window segmentation for developers, while allowing users to inspect those steps instead of hiding all data cleaning inside a black box.
2. Estimate Uncertainty
Businesses usually need more than a single forecast curve; they also need to know how confident the model is. Power-load forecasts could include upper and lower bounds, industrial equipment forecasts could indicate risk ranges, and inventory systems could estimate the probability of a stockout.
If a model can only produce a seemingly precise number without indicating the likely error range in extreme scenarios, managers will find it difficult to connect that model to automated decisions. The next stage of competition among time-series foundation models may shift from point-forecast error alone to prediction intervals, anomaly detection, and risk calibration.
3. Address Concept Drift
Time-series data is not static. Seasonal changes, equipment aging, policy adjustments, and shifts in consumer behavior can all gradually invalidate previously useful patterns. Once a model is deployed, data distributions and forecast errors need to be monitored continuously to determine when retraining or fine-tuning should be triggered.
This means a time-series model project is inherently not a one-time deployment, but an ongoing operating system. The teams that standardize model updates, evaluation, staged releases, and anomaly rollback are more likely to turn leaderboard performance into lasting business value.
China’s Time-Series Foundation Models Enter the Engineering Race
Qiyao’s arrival also reflects how competition among China’s large models is expanding beyond general-purpose language models into more specialized areas that are closer to the core of industrial systems.
Users can quickly assess the performance of a language model through conversation. The value of time-series models is usually buried in system metrics: one fewer equipment shutdown, a slight reduction in reserve power, an anomaly detected several hours earlier, or more stable inventory turnover. These gains are less visible to consumers, but are more likely to have a direct impact on enterprise procurement.
Huawei’s advantage is that it combines cloud services, computing platforms, and industry customers. If Qiyao can provide clearer information about its capabilities, evaluation details, deployment options, and industry use cases, it could follow a path distinct from that of research-only teams: not just releasing weights or papers, but embedding the model in production workflows across power, manufacturing, transportation, and finance.
Developers will still care most about usability: whether the model is open, which time granularities and variable types it supports, how long its context window is, whether it offers probabilistic forecasts, how it handles missing values, and whether it can be deployed on platforms other than Ascend. Public information currently provides few details on these points, and subsequent product announcements will determine Qiyao’s practical impact.
For businesses, a more realistic approach than replacing an existing forecasting system as soon as a model tops a leaderboard is to choose a clearly defined use case for a controlled comparison. For example, they could split historical data into a strict out-of-time test set, compare their existing model with Qiyao’s zero-shot and fine-tuned results, and evaluate accuracy, latency, compute cost, and performance in anomalous scenarios side by side.
If Qiyao can continue to prove itself against these real-world metrics, its significance will be more than a place on two leaderboards: it could help shift time-series forecasting from “reinventing the wheel in every industry” to a development paradigm based on foundation models and industry adaptation. For companies building AI infrastructure, that may be more worth watching than yet another chatbot.
Huawei has not yet publicly disclosed Qiyao’s full technical report, parameter count, open-source plans, or a unified API access method. If model weights, evaluation code, or a cloud service become available, developers should pay close attention to data licensing, commercial-use restrictions, cross-hardware compatibility, and the stability of long-horizon forecasts.
Conclusion
Qiyao, jointly released by Huawei and East China Normal University, arrives at a pivotal point in the transition of time-series foundation models from research to industry. Topping both the TIME-leaderboard and GIFT-Eval shows that Chinese teams can compete on equal footing with research groups around the world. Its Ascend-native optimization also ties model performance to the domestic computing ecosystem.
But the real test lies beyond the leaderboards: whether the model can handle messy data, provide trustworthy uncertainty estimates, adapt to industries at low cost, and remain stable over long periods of operation. Qiyao has earned a place in the industrial arena. The next challenge is to prove that it can become time-series infrastructure businesses are willing to rely on over the long term.
References
- IT Home: Huawei and East China Normal University Jointly Develop a Time-Series Forecasting Foundation Model, Topping Two Leading International Time-Series Benchmarks — Public information about Qiyao, TIME-leaderboard, GIFT-Eval, and its zero-shot forecasting and low-cost fine-tuning capabilities.



