DocsQuick StartAI News
AI NewsAiqiu Goes on Sale, ELA Lets Robots “Read Facial Expressions”
New Model

Aiqiu Goes on Sale, ELA Lets Robots “Read Facial Expressions”

2026-08-19T09:07:10.853Z
Aiqiu Goes on Sale, ELA Lets Robots “Read Facial Expressions”

The AIqiu emotional-interaction humanoid robot goes on sale today starting at RMB 98,000. Its key selling point is the proprietary ELA model, which integrates emotion perception, language generation, and action decision-making into a single interaction pipeline.

AIQ Goes on Sale, ELA Lets Robots “Read the Room”

On August 19, Sichuan Embodied Humanoid Robot Technology Co., Ltd. officially launched its emotionally interactive humanoid robot, AIQ, with prices starting at RMB 98,000. More noteworthy than the price is the company’s attempt to add a long-overlooked piece to the embodied intelligence puzzle: robots must not only understand what people say, but also determine the emotion behind it and decide what language, facial expressions, and movements to use in response.

To achieve this, AIQ unveiled an “Emotion-Language-Action” foundation model, abbreviated as ELA. The company claims it is the world’s first model of its kind and has integrated it with a proprietary multimodal affective computing engine to connect emotion perception, dialogue generation, and robotic action decision-making.

This narrative may sound like adding an “emotional intelligence module” to a large language model. But if ELA truly works as described, the problem it must solve is far more complex than conversation. A person saying “I’m fine” may genuinely mean they are fine—or that they do not want to continue talking. The robot must look beyond the words and simultaneously interpret speaking rate, pitch, pauses, facial expressions, and body posture, then decide whether to move closer, step back, offer comfort, or simply remain silent.

Front view of the AIQ emotionally interactive humanoid robot, highlighting its dragon-lizard-inspired design and 3D-projected face

ELA Is Not Just a New Prompt for a Chat Model

Most current companion robots are still essentially “voice assistants in movable shells”: speech is converted to text and sent to a language model, the model generates a reply, and a movement is then selected from a preset action library either randomly or through keyword matching. A robot may wave while saying, “It’s great to see you,” without necessarily knowing whether the other person is actually happy.

ELA aims to tighten up this loosely connected pipeline. Based on currently available public information, its interaction process can be roughly divided into four layers:

  1. Multimodal signal acquisition: Collecting information such as speech, facial expressions, voiceprints, acoustic-field data, and potentially body posture;
  2. Emotional-state modeling: Determining the other person’s emotion type, intensity, and direction of change from multidimensional signals;
  3. Language and behavior decision-making: Using the conversational context to determine the reply, tone of voice, and actions the robot should perform;
  4. Embodied expression: Presenting those decisions through voice, the projected face, head orientation, arm movements, and body language.

The third step is crucial. Emotion-recognition models are nothing new, and sending an emotion label to a language model is not difficult. The real challenge is ensuring that verbal responses and physical movements do not contradict one another. If a robot says, “I understand,” while suddenly moving rapidly toward the user—or keeps asking questions when the user is clearly irritated—the interaction will quickly descend into the uncanny valley.

This is also where ELA differs in purpose from common VLA models—Vision-Language-Action models. VLA generally deals with instructions such as, “There is a cup on the table; bring it to me,” and is evaluated mainly on recognition accuracy, trajectory quality, and task completion. ELA faces questions more like, “This person has fallen silent—should I say something now?” The former pursues success in physical tasks, while the latter must also account for social boundaries, emotional continuity, and individual differences.

In other words, VLA answers “What should I do?” ELA must also answer “Should I do it, when should I do it, and how conspicuously should I do it?”

Emotion Is Not a Label That Can Be Reliably Read

The direction AIQ has chosen is valuable, but it is also much more difficult than the marketing language suggests.

Human emotions do not form a set of clear, stable categories. The same facial expression can mean entirely different things across cultures, relationships, and situations. Faster speech may indicate excitement or anxiety. Avoiding eye contact may signal impatience—or simply that the user is looking at their phone. Relying on a single facial frame or one sentence makes it easy for “affective computing” to become little more than mechanical labeling.

A genuinely practical ELA system therefore needs at least three capabilities:

  • Temporal continuity: It cannot judge only the current second; it must observe changes in speech, facial expressions, and behavior over time;
  • Expression of uncertainty: The model should not confidently reach a conclusion every time. It must be able to output states such as “possibly feeling down, but with insufficient evidence”;
  • Personalized calibration: Some people naturally speak softly, while others habitually show little facial expression. The system must establish long-term individual baselines rather than applying the same thresholds to every user.

Long-term memory is equally important. If a companion robot treats the user as a stranger every time it starts up, its supposed “emotion” amounts to nothing more than an immediate performance. But once it begins continuously storing voiceprints, facial data, family relationships, and emotional records, it immediately encounters issues involving privacy, consent, and data security.

For developers, the hardest part of such a system is not connecting more sensors, but establishing an explainable and reversible state-management mechanism. For example, can users view the emotional inferences the robot has stored? Can incorrect records be deleted? How is data isolated among family members? These questions may attract less attention than degrees of freedom or model parameter counts, but they directly determine whether the product can enter homes, hospitals, and educational settings.

A Projected Face Is a Practical Solution—and a Trade-Off

AIQ uses neither simulated silicone skin nor a simple flat display installed on its face. Instead, its face is created using 3D ultra-short-throw projection. The immediate benefit is that its expressions can be defined in software, allowing developers to swap character appearances, adjust gaze and emotional style, or even switch the robot’s entire visual persona for different scenarios.

This approach is easier to engineer than a complex mechanical face. Mechanical eyebrows, eyelids, and mouths require numerous miniature actuators, increasing cost, noise, weight, and failure rates. Projection transfers some of that mechanical complexity to graphics rendering and optical systems. For a product starting at RMB 98,000, it is a relatively realistic method of controlling costs.

Projection, however, comes with its own limitations. Ambient light affects display quality, while viewing angles, facial-surface calibration, and projection focus all require ongoing adjustment. If lip movements, audio, and physical actions are misaligned by tens or hundreds of milliseconds, people will notice the anomaly more readily than they would on a conventional display.

Emotionally interactive robots are particularly sensitive to latency. If a utility robot pauses for two seconds, users may assume it is planning a route. If a companion robot pauses for two seconds and then smiles, users will simply feel that its “reaction is wrong.” ELA’s practicality therefore cannot be judged solely by offline recognition accuracy. It must also be evaluated on end-to-end latency from perception to action execution, as well as how it arbitrates conflicts among different modules.

Illustration of facial-expression changes on AIQ’s 3D ultra-short-throw projected face, showing calm, happy, and concerned states

At RMB 98,000, This Is Not a Household Appliance

In terms of hardware, public reports indicate that the highest-end AIQ configuration has 32 degrees of freedom, offers a choice of three-fingered or five-fingered dexterous hands, uses a side-removable plug-in battery, and has a rated battery life of approximately three hours. Reports also describe it as “135 mm tall,” which is clearly inconsistent with the humanoid form factor and previously disclosed information, making it more likely to be a unit or data-entry error. Earlier materials list specifications of approximately 1.43 meters and 38 kilograms. Before purchasing, buyers should still rely on the manufacturer’s final specification sheet and contract.

RMB 98,000 is already less expensive than many research-grade humanoid robots, but this is still not a consumer-electronics product for the mass household market. Once dexterous hands, batteries, maintenance, model services, and scenario customization are included, actual deployment costs will most likely exceed the starting price.

In the short term, AIQ is better suited to the following markets:

  • Reception and guided-tour services in science museums, exhibition halls, and branded retail stores;
  • Companionship, reminders, and simple interactions in eldercare and wellness facilities;
  • Affective computing and embodied intelligence research in schools and laboratories;
  • Cultural tourism projects, IP operations, and live performances;
  • Service scenarios that require highly humanlike expression but not heavy physical labor.

These scenarios share one characteristic: users are willing to pay for “expressiveness” and a “sense of character,” rather than calculating return on investment solely in terms of handling efficiency. AIQ’s dragon-lizard-inspired design and simultaneous development of a character persona and content IP also indicate that the team does not intend to compete head-on with industrial humanoid robots. It is more like a programmable embodied character than a general-purpose laborer sent into a factory to tighten screws.

This positioning is smart. China’s humanoid robot companies are collectively competing in walking, grasping, box handling, and factory training, and the market has already become highly homogeneous. By avoiding direct competition in payload and locomotion performance and focusing its resources on emotional expression and culturally distinctive characters, AIQ has at least established a clear identity.

The Biggest Gap Right Now Is Verifiable Information

As of August 19, the manufacturer had not provided sufficient detail about ELA’s model architecture, training data, parameter scale, or evaluation methods. The claim that it is the “world’s first” currently comes primarily from the company itself, while reproducible benchmarks and third-party testing remain unavailable.

What developers really need to see is not a robot completing a scripted conversation at a launch event, but answers to the following questions:

  1. Is ELA a unified end-to-end model, or a combination of multiple models and rule-based systems?
  2. Which modalities does its emotion recognition cover, and can it operate reliably in noisy, backlit, and multi-person environments?
  3. Are language and actions generated jointly, or is text generated first and then matched to an action library?
  4. Does inference run locally, on an edge server, or in the cloud? Which capabilities remain available without an internet connection?
  5. What is the end-to-end response latency, and does action planning have an independent safety layer?
  6. Does the model provide an SDK, APIs, or fine-tuning capabilities? Can third parties integrate their own characters and actions?
  7. What metrics does the manufacturer use to measure “more accurate emotional responses”: classification accuracy, user satisfaction, or long-term interaction retention?

The answers will determine whether ELA represents a genuinely new model paradigm or is merely a new system name for a repackaged emotion classifier, language model, and action library. At this stage, the more prudent assessment is that AIQ has placed affective computing at the core of the robot’s decision-making pipeline, but it has not yet disclosed enough evidence to demonstrate that ELA achieves an end-to-end breakthrough at the model level.

Robots Need Emotional Intelligence, but They Must Not Pretend to Understand People

AIQ’s significance does not lie in enabling robots to perform a few more smiling motions. Rather, it expands the evaluation criteria for embodied intelligence from “Was the task completed?” to “Was the interaction appropriate?” As robots enter homes, eldercare, education, and public services, correct actions will be only the baseline. Appropriate social behavior will become a new threshold for viable products.

However, emotionally interactive robots are also especially prone to creating illusions about their capabilities. A system may recognize sadness in someone’s voice and generate comforting words, but that does not mean it truly understands sadness. It may remember a user’s habits, but that does not mean it has formed an equal relationship with the user. If product design deliberately blurs this boundary, technological novelty will quickly turn into ethical controversy.

ELA is therefore worth watching—but even more important is how it acknowledges its own uncertainty. A mature emotionally interactive robot should not rush to demonstrate “I understand you” in every interaction. It should know when to ask, when to step back, and when to hand the issue back to a human.

AIQ has already connected “perceiving emotion—understanding context—generating language—executing actions” into a complete product pipeline, and it has begun testing the market with a starting price of RMB 98,000. The real test is not how many expressions it can display at a launch event, but whether it can consistently deliver responses that are inoffensive, natural, and sufficiently stable over weeks and months of real-world interaction.

If it succeeds, ELA could become an important branch of embodied intelligence. If it fails, it will be little more than a robot with a face that is better at changing expressions.

References

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: