DocsQuick StartAI News
AI NewsXianglu enables AI to understand what’s cooking in the pot in real time.
New Model

Xianglu enables AI to understand what’s cooking in the pot in real time.

2026-08-23T15:04:04.696Z
Xianglu enables AI to understand what’s cooking in the pot in real time.

On August 23, Xianglu Robotics unveiled its multimodal cooking AI foundation, CookingMuse 厨启, at the 2026 World Robot Conference, along with the 3K Vision AI cooking robot and the embodied cooking robot 现炒方舟. The core breakthrough is that the vision model now goes beyond identifying ingredients to assessing dynamic conditions inside the wok and adjusting cooking parameters in real time.

Xianglu Lets AI See and Understand What’s Cooking in the Wok in Real Time

On August 23, Xianglu Robotics made the global debut of three cooking robot products at the 2026 World Robot Conference: the multimodal cooking AI foundation CookingMuse, the 3K Vision AI Cooking Robot, and the embodied cooking robot “Fresh-Cook Ark.”

The focus of this launch was not simply that “robots can stir-fry,” but that the control paradigm for cooking AI has changed. The system now attempts to continuously observe changes inside the wok through cameras, then dynamically adjust heat, cooking time, and movements based on the state of the ingredients. Many previous cooking robots were more like automated devices that executed fixed scripts. Xianglu, by contrast, aims to build a cooking system capable of adapting to conditions on the fly.

This may look like nothing more than adding cameras to a wok, but it addresses one of the most difficult problems in food-service automation: the ingredients and real-world conditions for the same dish are rarely truly identical.

The Xianglu 3K Vision AI Cooking Robot uses industrial cameras to capture real-time images during cooking

From “Following a Recipe” to “Making Decisions Based on Conditions”

According to Xianglu, the 3K Vision AI Cooking Robot is equipped with three 40-megapixel global-shutter industrial cameras that capture images inside the wok at 30 FPS. The images are processed by an AI model trained on millions of dish-related data samples, allowing it to assess changes in the ingredients in real time and make cooking decisions accordingly.

The key is not the cameras’ resolution, but whether visual information genuinely enters the closed-loop control system.

Traditional automated cooking equipment typically relies on preset programs: heat for a specified number of seconds in Stage 1, stir a specified number of times in Stage 2, add seasonings in Stage 3, and then proceed to plating. Such systems are highly efficient under standardized conditions, but real kitchens are not standardized. Slightly thicker slices of meat, vegetables with higher water content, incompletely thawed frozen ingredients, or deviations in the order in which ingredients are added can all cause the same program to produce different results.

Fixed scripts solve the problem of “repeating actions,” but they do not necessarily ensure “consistent results.”

CookingMuse takes an approach closer to perception and control in autonomous driving. Instead of merely receiving an instruction indicating which step to execute, the model continuously reads images from inside the wok, determines the current state of the ingredients, and decides what to do next. For example, it may assess whether the ingredients have been heated sufficiently, whether their surface color has changed, whether there is too much moisture in the wok, whether stir-frying should continue longer, or whether the heat should be reduced to prevent overcooking.

In terms of product positioning, CookingMuse is more like an AI foundation for cooking scenarios than a standalone machine. It must connect visual perception, culinary knowledge, cooking reasoning, and equipment control, ultimately enabling cooking devices in different form factors to share a common ability to “understand the kitchen.”

Three Industrial Cameras Do More Than Simply “See Clearly”

The 3K robot uses three 40-megapixel global-shutter industrial cameras to capture images at 30 FPS. Two technical aspects of this configuration are particularly noteworthy.

The first is the global shutter. Conventional rolling-shutter cameras expose an image line by line, which can cause localized distortion during rapid stirring, wok tossing, or ingredient movement. For a system that needs to identify the position, shape, and motion of ingredients in the wok, this increases uncertainty in visual assessments. A global shutter captures a more consistent image of a single moment, making it better suited to high-speed motion.

The second is continuous observation at 30 FPS. This does not mean the model must perform a complete culinary inference 30 times per second. Rather, it gives the system a sufficiently dense stream of signals showing how conditions change over time. Many changes during cooking are continuous: moisture gradually evaporates, ingredient colors shift slowly, and the temperature in the wok fluctuates over short periods. Compared with taking photos at only a few fixed points in time, a continuous video stream is better suited to capturing these processes.

However, a vision system still faces many challenges in genuinely understanding conditions inside a wok. These include reflections from oil, steam occlusion, overlapping ingredients, and changing lighting. The wok itself may also remain in motion. The model must determine not simply “what is in the wok,” but “what stage of cooking these ingredients are currently in.”

This means CookingMuse must handle more than object detection. It also requires temporal understanding and state estimation: what changes in the color, position, and shape of the same piece of meat mean at different points in time, and whether similar-looking frames should prompt different actions when their preceding and subsequent context differs.

For developers, this is more like a real-time multimodal control problem than a conventional image-recognition task. The model’s output cannot stop at “broccoli detected” or “this dish is Kung Pao chicken.” It must be translated into executable equipment parameters, such as heating power, stirring speed, movement duration, and the appropriate time to plate the dish.

In-Process Control Is More Valuable Than Post-Process Inspection

Xianglu describes this capability as moving quality control forward from “after the process” to “during the process.” This is a crucial distinction.

Restaurant chains have traditionally relied on weighing, photography, spot checks, and customer feedback to assess results after the fact. Such methods can identify problems, but they cannot salvage a dish that has already been plated. A truly valuable automation system should intervene as soon as a deviation appears: if it detects that ingredients contain more moisture, it should extend a heating stage; if the temperature in the wok rises too quickly, it should reduce the heat in time; if an ingredient container is placed incorrectly, it should pause or adjust the process.

According to available information, the 3K can dynamically adjust parameters such as heating power and cooking duration with millisecond-level responsiveness. Even when ingredient specifications are inconsistent, thawing conditions differ, or ingredient containers are misplaced, the system attempts to maintain relatively consistent results.

Of course, “millisecond-level adjustment” more accurately refers to the responsiveness of the control system; it does not mean the model rethinks the entire dish every millisecond. A cooking model generally needs to be divided into multiple layers: the lower layer handles sensor acquisition and actuator response, the middle layer manages temperature, movement, and timing, while the upper-layer model understands the state of the dish and determines the strategy. Only in this way can the system potentially balance real-time performance with the complexity of inference.

This is also a fundamental difference between cooking AI and chat models. If a chat model takes a few extra seconds to respond, it is usually only a user-experience issue. If a cooking robot is delayed at a critical point, however, the result may be burnt food, undercooked ingredients, or a safety hazard. Whether CookingMuse can ultimately be deployed in practice therefore depends not only on how “smart” the model is, but also on whether edge inference, equipment control, exception handling, and safety redundancy are sufficiently robust.

Fresh-Cook Ark: Compressing the Back of House into a Mobile Unit

If the 3K addresses “how to cook a dish more consistently,” then Fresh-Cook Ark addresses “whether an entire back-of-house kitchen can be turned into an unmanned module.”

Xianglu defines Fresh-Cook Ark as a mobile, miniature, unmanned kitchen built around embodied-intelligence robots. It covers fresh-cut ingredient storage and preservation, on-demand cooking, bowl dispensing and meal serving, cooking-fume treatment, and equipment self-cleaning. Its goal is to create a complete closed loop from ingredient intake to finished-meal delivery.

Its potential lies not in replacing a single action performed by a chef, but in redefining the boundaries of food-service spaces. Traditional restaurants require fixed kitchens, exhaust-system construction, staffing, and relatively complex operational management. If all these capabilities can be packaged into a mobile unit, such systems could theoretically be deployed in shopping malls, business parks, exhibitions, airports, and even temporary event venues.

What also distinguishes this type of product from an ordinary vending machine is its emphasis on “cooking to order,” rather than simply reheating pre-prepared food. For consumers, freshly stir-fried food offers greater immediacy. For operators, the real challenge is maintaining consistency in ingredient storage, cooking cadence, hygiene management, and equipment maintenance.

Whether Fresh-Cook Ark can become a replicable commercial unit will ultimately depend on several metrics: daily meal output, per-serving cost, equipment failure rate, cleaning and maintenance time, and actual consistency across different combinations of dishes. Getting the robot to complete the workflow is one thing; enabling it to operate continuously during peak hours and at low cost is another.

Xianglu Is Really Betting on a “Closed Loop of Cooking Data”

From CookingMuse and the 3K to Fresh-Cook Ark, Xianglu has not launched three isolated pieces of hardware, but a vertical system that runs from models to equipment and then feeds data from the equipment back into the models.

Models need large volumes of real-world cooking-process data to learn how to distinguish between conditions involving different ingredients, cookware, heat levels, and environments. The more devices deployed, the richer the collected data becomes—including videos from inside the wok, temperature curves, motion trajectories, and final cooking results. Only with richer data does the model have an opportunity to progress from “following recipes” to “understanding how to adjust under different conditions.”

This path resembles the development logic of general-purpose foundation models, but it is also more closed and resource-intensive. General-purpose models can obtain vast amounts of training material from text and images on the internet, whereas cooking models require real-world process data labeled with time, temperature, actions, and results. A single image can tell the model, “This is a plate of stir-fried beef,” but it cannot directly tell the model when the beef should be turned, when excess moisture in the wok should be reduced, or how the heat should change to achieve the desired texture.

The competitive moat in cooking AI may therefore lie not in the specifications announced at a product launch, but in three areas:

  1. Whether a company has enough real-world cooking-process data, rather than merely recipes and images of finished dishes;
  2. Whether it can convert model assessments into reliable equipment-control instructions;
  3. Whether it has established a closed loop in which cooking results feed back into the model, enabling the system to continuously correct itself.

Xianglu previously partnered with Haier Robotics and proposed jointly developing a fully automated AI cooking robot for the home. If this direction continues, the data accumulated in commercial kitchens and the personalized needs of home kitchens could complement each other: commercial scenarios can provide stable, large-scale cooking data, while household scenarios introduce greater variation in ingredients, tastes, and equipment conditions.

However, home kitchens are far more complex than commercial kitchens. Users do not have standardized cookware, ingredient quantities vary, kitchen space is limited, tolerance for cooking fumes and noise is lower, and safety responsibilities extend directly into the home environment. The fact that commercial robots can operate with relatively standardized woks, lighting, and workstations does not mean they can be introduced directly into ordinary households.

This Is Not a “Chatbot for the Kitchen”

It is important to make clear that the value of CookingMuse does not lie in allowing a user to say, “Make a home-style dish,” and having the model generate a recipe. That is merely the easiest layer of cooking AI to demonstrate.

The harder part is enabling the system to understand real-world uncertainty and act within it. The model must simultaneously process vision, equipment status, culinary knowledge, and safety constraints. It must also handle extreme situations, such as a sudden fire in the wok, dropped ingredients, a camera obscured by steam, missing critical ingredients, or actuator failures.

In other words, CookingMuse is closer to a domain-specific perception–decision–execution system than a question-answering model wrapped in a kitchen-themed shell. Its upper limit depends on the model’s capabilities, but its baseline performance depends on the reliability of the engineering system.

From an industry perspective, the significance of this launch is that vision models are beginning to move from “identifying objects in the kitchen” to “understanding states during the cooking process.” The former can already be used for dish recognition, inventory management, and operational guidance. The latter is what truly involves the core control challenges of automated cooking.

If this step can operate reliably in real-world commercial settings, its impact will extend beyond a single cooking robot. It may push food-service equipment from fixed-program automation toward flexible automation based on real-time perception. A single machine would no longer need to serve only one standardized dish; instead, it could dynamically adapt to different ingredients and on-site conditions.

However, Xianglu’s current disclosures still focus primarily on the product launch and technical roadmap. Little public information is available about CookingMuse’s specific model architecture, training-data scale, inference deployment method, supported range of dishes, or the mass-production timeline and commercial pricing of Fresh-Cook Ark. These metrics will determine whether this is merely a robot launch with strong demonstration value or an infrastructure product that can genuinely enter the food-service supply chain.

At least in terms of direction, Xianglu has chosen the right problem. What food-service automation needs most is not another set of fixed actions, but machines that can see real-world conditions, understand changes, and make adjustments before errors occur. AI in the kitchen is finally moving from “reciting recipes” to “watching the wok.”

What This Means for Developers

Products such as CookingMuse also give AI developers a clear window into the industry: competition among vertical multimodal models is shifting away from model leaderboards and toward closed-loop performance metrics in the real world.

In the future, a cooking model may no longer be evaluated by whether it can generate a plausible-looking recipe, but by whether it can:

  • Reliably identify ingredient states in a video stream;
  • Distinguish between visually similar cooking stages that lead to different outcomes;
  • Continue operating safely when sensor data is incomplete or the camera view is obstructed;
  • Translate natural-language objectives into executable equipment strategies;
  • Continuously improve through every real-world dish it produces, rather than merely scoring highly on offline datasets.

This is the reality AI must face as it moves from software into the physical world: model outputs are no longer pieces of text, but heat, speed, temperature, and movement. An error is no longer merely an irrelevant answer; it may ruin an entire wok of food or even create risks to equipment and personal safety.

Xianglu’s launch has not answered every question, but it has brought the central technical challenge of cooking robots to the forefront: Can a machine truly understand what is happening inside the wok and translate that understanding into timely, controllable, and verifiable actions? If the answer gradually approaches “yes,” cooking robots will no longer be merely automated robotic arms in the back of house. They will become vertical AI systems with genuine perception and decision-making capabilities.

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: