DocsQuick StartAI News
AI NewsXiaomi’s Next-Generation Humanoid Robot Makes Its Debut
Industry News

Xiaomi’s Next-Generation Humanoid Robot Makes Its Debut

2026-08-19T23:07:23.491Z
Xiaomi’s Next-Generation Humanoid Robot Makes Its Debut

Xiaomi’s next-generation humanoid robot made its debut today at the 2026 World Robot Expo. According to a video released by the company, the robot can understand visitors’ intentions under the control of a large language model and independently perform real-time interactive actions such as handing over flowers, shaking hands, fist-bumping, and making a heart gesture. It will be open to the general public starting August 20.

Xiaomi’s New-Generation Humanoid Robot Makes Its Debut: Large Models Begin Taking Over “How to Move”

Xiaomi’s new-generation humanoid robot made its debut today at the 2026 World Robot Expo, demonstrating for the first time through a publicly released video how it interacts with visitors in real time.

From handing over flowers and shaking hands to fist bumps and finger-heart gestures, the movements shown in the video are not complex. Xiaomi emphasized, however, that these actions are not played back one by one according to a fixed script. Instead, they are driven by a large model that enables the robot to understand, make decisions, and complete them autonomously. Starting August 20, the expo will be open to the general public, and visitors will be able to interact directly with the robot.

The significance of this debut lies not in how many social gestures the robot can perform, but in Xiaomi’s attempt to move large models beyond “answering questions” toward “understanding the environment and taking action.” For humanoid robots, this also represents a step from demonstrating physical capabilities to demonstrating embodied intelligence.

Xiaomi’s new-generation humanoid robot interacts with visitors at the “Embodied Garden” booth at the 2026 World Robot Expo

Not an Action-Library Demonstration, but a Complete Interaction Pipeline

Over the past few years, humanoid robot demonstrations at product launches have focused largely on walking, running, jumping, doing somersaults, or performing standardized tasks such as carrying and grasping. These demonstrations primarily prove the robot’s mechanical structure, joint control, and motion-planning capabilities.

The interactive video released by Xiaomi showcases a different category of capability: when facing a real person, the robot first identifies what the person is doing, then determines how it should respond, and finally converts that decision into an appropriate physical movement.

Take “handing over a flower” as an example. The robot must complete several layers of processing: recognizing that a visitor is approaching or reaching out, understanding that the object in the person’s hand may have the characteristics of a gift or an invitation to interact, determining when to extend its arm and the bouquet, controlling its arm, wrist, and fingers to grasp and transfer the object, and maintaining its balance throughout the process. Handshakes, fist bumps, and finger-heart gestures are similar. They may appear to involve only changes in hand position, but in reality they require visual recognition, intent inference, action selection, whole-body coordination, and safety control.

This pipeline differs from the way traditional robots execute preset programs. A conventional solution often works like this: “When condition A is detected, execute action B.” An embodied system driven by a large model is closer to: “Based on the current environment and context, determine what should be done now.” The former resembles a vending machine, with strictly defined inputs and outputs; the latter is more like an operating system capable of temporarily selecting a response based on the situation at hand.

Of course, this does not mean that the robot already possesses general intelligence in open environments. Interactions at an expo typically take place in a clearly defined venue, within a limited range of actions and under controllable crowd conditions. The tasks the robot faces are far simpler than those found in homes, factories, or public spaces. Xiaomi’s publicly released information also does not disclose specific parameters such as the model name, on-device computing power, perception system, action-planning framework, or control cycle. Therefore, the more accurate description at this stage is that Xiaomi has demonstrated a productized prototype involving large models in the real-time interaction of a humanoid robot, rather than having already solved general-purpose embodied intelligence.

What Large Models Really Change Is the Robot’s “Action Entry Point”

The hardware of humanoid robots has been advancing for a long time. Motors, reducers, joint modules, controllers, and battery systems continue to improve, enabling robots to walk more steadily, run faster, and carry heavier loads. But hardware alone still leaves robots capable only of executing actions designed in advance by engineers.

The addition of large models changes the entry point through which actions are generated.

Without large models, engineers must break down a large number of scenarios into rules: shake hands when an outstretched palm is detected, perform a fist bump when a fist approaches, and make a finger-heart gesture when a specific hand sign is detected. The more rules there are, the more complex the system becomes. When a visitor’s movements vary slightly, or when the scene includes occlusion, noise, or simultaneous interaction with multiple people, a rule-based system can easily fail.

Large models can process vision, language, and context within a single decision-making framework. If a visitor says, “Shake hands with me,” the robot can select the action directly from the language input. If the visitor says nothing but extends a fist, the system can also combine visual information to infer that the person wants a fist bump. If the visitor first offers a flower, the robot must further infer that this is an invitation to interact rather than an ordinary movement of an object.

For developers, this change can be understood as a shift from “calling fixed action APIs” to “letting the model select action APIs.” The underlying robot still requires numerous reliable control modules, but the upper layer no longer needs a separate set of rules for every natural interaction scenario. Instead, the model maps environmental inputs to action goals, which are then handed over to the motion-control system for execution.

This architecture generally forms a layered system:

  • The perception layer processes inputs from cameras, microphones, force sensors, joint states, and other sources, identifying people, objects, poses, and spatial relationships.
  • The decision layer is handled by a vision-language model or embodied large model, which understands human intent, determines the current task, and generates action goals.
  • The planning layer breaks high-level intentions such as “shake hands” or “hand over a flower” into executable trajectories, poses, and timing sequences.
  • The control layer handles millisecond-level joint control, balance control, collision detection, and safety protection.

Large models do not directly replace motor controllers. Allowing a language model to control dozens of joints directly would be both unstable and unable to meet real-time and safety requirements. A more realistic approach is for the large model to handle relatively slower, high-level decisions, while the “cerebellum” or controller handles fast, continuous, and deterministic movement execution. Xiaomi has not disclosed its specific technical architecture, but its description of the robot “autonomously understanding, making decisions, and completing actions” at least indicates that the model has been placed within a key link of the perception-to-action pipeline.

Xiaomi’s Latest Demonstration Is Different from the 2022 CyberOne

Xiaomi has previously entered the humanoid-robot field. In August 2022, Xiaomi unveiled the full-size humanoid bionic robot CyberOne, demonstrating capabilities including bipedal walking, environmental perception, and emotion recognition. At the time, the industry’s attention was focused more on whether Xiaomi could make the transition from consumer electronics to robot hardware, as well as whether the robot itself could stand and walk stably.

Four years later, Xiaomi’s new-generation humanoid robot has made another appearance, with the focus shifting from “Can the robot move?” to “Can the robot understand before it moves?” This reflects a change in the competitive focus of the humanoid-robot industry.

Early competition centered on the robot body itself: whose joints were more flexible, whose walking was more stable, and whose robot could perform more difficult dynamic movements. As the hardware supply chain gradually matures, purely physical demonstrations are becoming increasingly difficult to turn into lasting differentiation. A robot that can walk, run, and do somersaults may have good motion control, but that alone does not prove that it can work continuously in real-world scenarios.

The key questions in the next stage are whether a robot can understand a task, adjust its strategy when the environment changes, transfer the same capability to different scenarios, recover when errors occur, and operate continuously at an acceptable cost.

By placing the new robot at the “Embodied Garden” booth and arranging for members of the public to participate in interactions, Xiaomi is effectively testing something more difficult than a product-launch demonstration: whether the robot can perform consistently when faced with nonstandardized human input. Visitors will not extend their hands or fists according to an engineer-defined rhythm, nor will they stand in the same position every time. Their reaction speeds, heights, range of motion, and willingness to interact will all differ, and these variables will directly expose the robot’s perception and decision-making capabilities.

For this reason, the robot’s performance after the booth opens on August 20 will be more worth watching than a carefully edited video. Whether it misinterprets gestures, has noticeable delays, can handle multiple people approaching at once, moves gently enough, and experiences latency or degradation after continuous interactions will all be more technically meaningful than whether it can “make a finger heart.”

Several Hard Problems Remain Before Genuine Deployment

The first is real-time performance.

Large models are good at understanding complex semantics, but model inference generally requires computing power and time. A human handshake is an immediate response. If the robot needs to wait several seconds to decide whether to extend its hand, the interaction will feel mechanical. Humanoid robots must balance the use of cloud-based large models, edge computing, and local control: complex tasks can be delegated to the cloud, while actions involving safety and low latency must be performed locally as much as possible.

The second is action reliability.

It is not difficult to demonstrate one successful action. The challenge is to keep succeeding under different angles, lighting conditions, crowds, and obstacles. For a robot, shaking hands is not simply a matter of extending its hand. It must control the force, determine the state of contact, and stop promptly if the other person suddenly withdraws their hand. If handing over a flower is involved, it must also handle dropped objects, occlusion, and failed grasps.

The third is the safety boundary.

Large models can generate flexible action strategies, but they may also produce plans that violate physical constraints or safety rules. Therefore, high-level models must be subject to strict action-space limitations. All actions involving human contact require collision detection, speed limits, force limits, and emergency-stop mechanisms. There is still an entire engineering system between the model “understanding the intent” and the robot “being able to execute it safely.”

The fourth is the data loop.

Embodied-intelligence models require large amounts of real-world interaction data. Language models can learn from internet text, but robots must learn in the physical world about object weight, friction, occlusion, changes in lighting, and differences in human behavior. In theory, every interaction at the booth can become training and evaluation data, provided that the company can effectively collect, clean, and label the data, then feed failure cases back into the model and control system.

The fifth is commercial value.

Expo interactions are well suited to helping the public quickly understand embodied intelligence, but they do not directly prove industrial value. For manufacturing customers, what really matters is whether the robot can reliably perform tasks such as loading and unloading, quality inspection, material handling, and assembly; whether it can connect to existing production lines; whether it can operate continuously for thousands of hours; and whether its maintenance costs are lower than those of human labor and traditional automation equipment.

From this perspective, the capabilities Xiaomi has publicly demonstrated so far are closer to a “human-robot interaction interface” than to a mature productivity tool. They can help Xiaomi validate multimodal understanding, action planning, and human-robot collaboration, and can also help it accumulate data for future home services, retail services, and smart manufacturing. However, more information and on-site validation are still needed before large-scale sales become possible.

Humanoid Robots Enter the “Brains Matter” Stage

For some time, the main strengths of China’s humanoid-robot industry have been its supply chain and motion control. Servo motors, reducers, controllers, and full-robot manufacturing capabilities have continued to mature, pushing the industry from laboratory prototypes toward mass-produced prototypes. Companies such as Unitree, UBTECH, Astribot, and LimX Dynamics are also advancing productization in areas including dynamic movement, industrial manipulation, and complex task execution.

But physical capability is gradually becoming a basic threshold. As hardware structures and control algorithms continue to improve, walking, climbing stairs, carrying objects, and even highly dynamic movements may all be quickly matched by competitors. The real differentiator will be whether a robot can unify vision, language, spatial relationships, and bodily movement, and execute reliably in real-world environments.

The signal from Xiaomi’s latest debut is clear: it is not content to treat a humanoid robot as a piece of hardware that can walk. It wants to extend the capabilities of large models into the physical world. For a company with businesses spanning smartphones, automobiles, home appliances, and operating systems, this direction also fits its long-term ecosystem strategy. If robots can understand human instructions and perform tasks across different devices and spaces, they could become a new interaction endpoint within Xiaomi’s smart ecosystem.

However, ecosystem collaboration remains only a potential benefit for now. Xiaomi has not yet announced the new-generation humanoid robot’s specific price, mass-production timeline, battery life, payload, degrees of freedom, model size, chip configuration, or actual application customers. Securities Times, citing information from WRC2026, reported that the robot is approximately 1.70 meters tall, is currently at the prototype stage, and has roughly half of its degrees of freedom concentrated in its hands. After its domestic debut, it is also reportedly scheduled to be exhibited in Europe. The increase in hand degrees of freedom indicates that Xiaomi places considerable importance on fine manipulation and human-robot interaction, but whether this can translate into stable grasping and operational capabilities will require more test data.

For developers, what is worth watching next is not whether Xiaomi will add several more actions, but whether it will open real development interfaces, simulation environments, or datasets; whether it will allow third parties to integrate task-planning and skill modules; and what protocols it will use between the large model and the robot-control system. If the robot can only perform closed, predefined actions inside an exhibition booth, its primary value will be brand promotion. Only if developers can build new skills around it can it potentially become a platform product.

Conclusion

The debut of Xiaomi’s new-generation humanoid robot today represents a shift in the way humanoid robots are demonstrated: from showy physical movements toward natural interaction with ordinary people, and from preset action scripts toward large models participating in understanding and decision-making.

This is a meaningful step, but it should not be overinterpreted. Handing over flowers, shaking hands, fist bumps, and finger-heart gestures show that the robot can integrate large-model capabilities into a visible physical interaction. They do not, however, answer the questions of generality, reliability, safety, and cost that will determine commercialization.

The real dividing line may not be whether the robot can complete one successful interaction at an exhibition booth, but whether it can continue completing tasks after leaving the booth, in factories, stores, homes, or public spaces. Its performance after the public opening on August 20 will give the outside world a better sense of the prototype’s actual capabilities. Whether Xiaomi goes on to disclose further technical details and a mass-production roadmap will determine whether this debut was merely a product preview or a signal that its embodied-intelligence business is entering a substantive phase.

Sources

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: