Embodied data infrastructure is now being sold by the unit.

The embodied intelligence data foundation, launched at RMB 5,100, recently went on sale, signaling a shift in robot training data from custom collection to standardized supply. The price is not the point; the real change is that data is beginning to take on a reusable, verifiable, and tradable product form.
Embodied AI Data Infrastructure Is Now Being Sold as an Off-the-Shelf Product
Robot training data is finally being sold not only through custom projects, but also as a standardized product.
Recently, a product positioned as an “embodied AI data infrastructure platform” entered the market at an introductory price of RMB 5,100. It attempts to package the data collection, governance, management, and training support capabilities required for embodied intelligence into an infrastructure product that can be purchased directly. Its goal is not merely to sell a number of videos showing robotic arms performing tasks, but to build full-stack physical AI infrastructure spanning everything from data production to model deployment.
The significance of this development does not lie in whether RMB 5,100 is cheap. For teams that routinely invest millions of yuan in building data collection facilities and purchasing robots, this amount barely qualifies as the cost of a trial. What deserves more attention is that robot training data is taking on a clearly defined commercial form: it now comes with specifications, deliverables, and quality requirements—as well as a public price.
Over the past year, discussions in the embodied intelligence industry have focused primarily on robot hardware, dexterous hands, and end-to-end models. By 2026, however, more and more teams have discovered that what truly slows down R&D is often not the model architecture, but the fact that the data is simply not ready.
Why Embodied AI Data Has Been So Difficult to Sell as a Standardized Product
The text, code, and image data used by large language models can at least be copied, segmented, and cleaned in batches. Embodied AI data, by contrast, comes directly from the physical world. It records what a robot saw in a specific environment, what instructions it received, how its joints moved, and what happened after it executed an action.
A seemingly simple data sample involving “placing a cup into a tray” may include all of the following:
- Video feeds from multiple RGB or depth cameras;
- Robot joint positions, velocities, and torques;
- End-effector poses and gripper states;
- Natural-language task instructions;
- Teleoperation inputs from the operator;
- Timestamps, device parameters, and scene metadata;
- Outcome labels indicating whether the task was completed and at which step it failed.
The problem is that different robots represent this information in different ways. Even among robotic arms, some record six-axis joint angles, while others have seven degrees of freedom; some control the end-effector pose, while others directly output joint commands. Camera mounting positions, sampling frequencies, coordinate systems, and gripper definitions also vary.
It is as if every cloud provider used a different network protocol while keeping its log format undisclosed. Even after the data has been collected, it is difficult to transfer it directly to another robot or training framework.
As a result, embodied AI data has historically been delivered primarily as a project-based service: a customer first specifies the scenarios and actions it needs, after which the service provider builds the site, deploys the equipment, recruits data collection operators, and finally performs annotation and acceptance testing. This model can deliver data, but it is difficult to scale. Every time the robot platform, sensor, or task changes, the entire collection process may need to be rebuilt.
A standardized data infrastructure platform is intended to eliminate precisely this kind of duplicated effort.

RMB 5,100 Should Buy More Than Just a Hard Drive
Judging by its product positioning, a so-called “data infrastructure platform” should not be understood simply as a storage device preloaded with data. Its real value lies in standardizing the interfaces between collection, governance, training, and validation.
A data infrastructure platform that can be integrated into an R&D workflow must address at least four categories of issues.
First, Standardizing Data Structures
Robot training does not end with dragging videos into a model. Visual frames, action sequences, language instructions, and robot states must be aligned on the same timeline, while coordinate systems and control variables must also be clearly defined.
If an infrastructure platform can only read data generated by the vendor’s own devices, it is closer to a proprietary collection tool than genuine infrastructure. A standardized product should allow developers to connect robots of different brands and configurations, and convert raw records into a unified or mappable data structure.
Second, Managing Data Quality
The “volume” of embodied AI data is easily overstated. A robot operating continuously for ten hours does not necessarily produce ten hours of valid training data. Equipment pauses, occlusions, sensor drift, timestamp misalignment, and failed actions can all produce segments that are unusable for training.
In March this year, the Beijing Innovation Center of Humanoid Robotics disclosed that Phase I of its data facility covered more than 30 typical scenarios and over 120 mainstream robotic devices. It had also established standards for collection, annotation, and quality inspection, achieving a data acceptance rate of more than 95%. Conversely, this figure also shows that even at a professional facility, quality control remains an independent engineering discipline rather than a quick cleanup step after collection.
If a standardized platform merely counts files but cannot expose metrics such as success rate, synchronization error, dropped-frame rate, trajectory completeness, and task distribution, developers are still buying a black box.
Third, Integrating the Training Toolchain
The fact that data can be read does not mean it can be used directly for training. Development teams must still perform data segmentation, sampling, normalization, augmentation, version management, and separation of training and validation sets.
A more practical issue is that the data structure must be compatible with existing imitation learning, vision-language-action model, and policy learning frameworks. Otherwise, teams that purchase “standardized data” will still have to spend weeks writing conversion scripts, defeating the purpose of an out-of-the-box product.
Fourth, Preserving Traceability
When a robot model behaves abnormally, engineers need to trace policy outputs back to the training samples: which device collected the data, what sensor configuration was used at the time, whether the data was manually corrected, and which data version it belonged to.
This is similar to version control in software development. Without data lineage and change records, teams will struggle to reproduce experiments or determine whether model degradation was caused by code, hyperparameters, or a data update.
RMB 5,100 is therefore merely an eye-catching market label. Whether the product is worth buying depends on how much it actually delivers across these four areas, as well as how subsequent scaling, adaptation, and data services are priced.
The Industrial Foundation for Standardized Supply Is Already in Place
This commercialization attempt did not appear out of nowhere.
In September 2025, Shanghai launched China’s first standardized embodied intelligence dataset platform, “Pujiang X,” while also publishing standards for humanoid robot datasets covering dataset classification and coding, annotation specifications, quality evaluation, and formatting requirements. The platform aims to connect data collection, governance, training, and validation, addressing the lack of unified production and distribution rules for multimodal data.
A humanoid robot embodied manipulation dataset case study disclosed in the same year revealed another side of large-scale production: a training facility spanning more than 5,000 square meters, over 100 robots with different configurations, and dozens of real-world application scenarios, collectively generating more than one million data records and approximately 2.5 PB of real-robot data. The data includes robot joint poses, images, and text, and is governed through a combination of simulation data, manual review, and model validation.
By 2026, the production of embodied AI data had clearly developed into a specialized industrial chain. Upstream companies provide robots, sensors, computing power, and storage; midstream companies build collection facilities, teleoperation systems, and annotation and quality inspection platforms; downstream companies train foundation models and develop robots for applications in industry, the home, healthcare, and eldercare.
Since the beginning of this year, a large number of data collection jobs have also appeared on the open recruitment market. Using first-person-view devices, motion-capture equipment, or teleoperation systems, data collectors repeatedly perform tasks such as folding clothes, opening cabinet doors, and sorting parts. Although this may look like manual labor, every action is in fact constrained by camera positions, execution sequences, and quality rules. If a sensor malfunctions, the view is obstructed, or time synchronization fails, the entire recording may become unusable.
This explains why embodied AI data cannot simply adopt the production logic used for internet data: webpages can be crawled repeatedly, but physical actions must be performed again; if text cleaning fails, a script can be rerun, but if real-robot data collection fails, the site, robot, and operator must all be allocated again.
The Real Competition Is Not About “Who Has More Data”
Embodied intelligence companies like to emphasize the number of hours of data they possess, but hours alone are not a sufficiently reliable metric.
Twenty thousand hours of static, repetitive, narrowly distributed data may not be as valuable as two thousand hours covering different objects, lighting conditions, robot configurations, and failure cases. For policy models, edge cases are often more valuable than repetitive successful trajectories: What should happen if the cup is occluded? What if the drawer gets stuck? How should the robot recover if an object slips after being grasped?
Going forward, data products will need to be compared across four dimensions:
- Coverage: How many scenarios, tasks, robot configurations, and sensor combinations are included.
- Effectiveness: Whether the data can reliably improve the target model, rather than merely passing format checks.
- Transferability: Whether one dataset can be used with different robot platforms, and how much additional adaptation is required.
- Reproducibility: Whether users can track data versions and reproduce quality evaluations and training results.
There is also an easily overlooked issue: standardization does not mean forcibly reducing every robot to the same representation.
The dynamics of bipedal humanoid robots, wheeled robots, and fixed robotic arms differ significantly. A sensible standard should unify metadata, time synchronization, coordinate descriptions, and quality metrics while preserving platform-specific information. If contact forces, foot states, or dexterous-hand joint data are discarded for the sake of a unified format, the result is merely a lowest common denominator that is not conducive to training high-performance models.
The core capability of a data infrastructure platform is therefore not the invention of a universal format, but the establishment of a clearly defined common layer and device-specific extension layer so that different data can be discovered, read, compared, and converted.
What Practical Value Does It Offer Development Teams?
For leading companies that already have complete collection and training platforms, a standardized product priced at around RMB 5,100 is unlikely to replace their internal systems. It is more likely to serve roles such as data exchange, rapid validation, or edge-node deployment.
The organizations most likely to benefit are small and medium-sized robotics teams, university laboratories, and application integrators. These teams often have specific scenarios and robot platforms but lack the budget to build a data platform from scratch. In the past, they had to spend considerable time synchronizing camera feeds with joint logs, managing collection tasks, and filtering failed trajectories. They may now be able to obtain these capabilities through a standardized product.
However, several questions still need to be clarified before procurement:
- Which robots, sensors, and training frameworks are supported?
- Does the RMB 5,100 price cover hardware, a software license, or data services?
- Are data format documentation, export interfaces, and secondary development capabilities provided?
- Can data be stored offline, and how are intellectual property rights and privacy responsibilities allocated?
- Are device adaptation, storage expansion, and ongoing services charged separately?
- Is data quality accepted based only on file integrity, or is it validated through model performance?
If these questions do not have clear answers, the low price may simply be a customer-acquisition tactic, while subsequent adaptation and services account for the bulk of the cost.
Data Is Becoming a Commodity, but “Buy and Train” Is Still a Long Way Off
Our assessment is that the launch of embodied AI data infrastructure as a commercial product is a positive signal, but at present it is better viewed as the starting point for standardized supply rather than evidence that the industry’s problems have already been solved.
First, it demonstrates that market demand is changing: robotics companies are no longer satisfied with outsourcing batches of data collection tasks. They want systems that can continuously produce, govern, and reuse data. What suppliers sell is no longer limited to labor and facilities, but now includes toolchains, standards, and infrastructure.
This approach has the potential to reduce duplicated investment across the industry. For common tasks such as grasping, placing, opening and closing doors, and sorting, unified data formats, quality metrics, and evaluation methods could allow the same batch of data to be reused by more robot platforms and models, spreading out the cost of collection.
However, it will not immediately make robot data as standardized as cloud computing resources. Embodied models must ultimately contend with specific hardware, specific scenarios, and long-tail problems in the real world. Standardized data can shorten the cold-start phase, but it cannot replace hardware adaptation, on-site collection, and closed-loop validation.
More precisely, the industry is shifting from “every company building its own roads” to “laying public roads first, then completing the last mile.” RMB 5,100 is not the true total cost of embodied intelligence data, but it may represent the first time robotics teams can purchase part of their data infrastructure in the same way they would buy standardized software or a development kit.
Only when data can be clearly priced, accepted against defined standards, and circulated across devices can embodied intelligence truly possess an industrial foundation comparable to the data pipelines of the large-model era. That is what makes this product launch most worthy of attention.
References
Due to the domain allowlist restrictions for links at the end of this article, original URLs are not provided for the following Chinese sources:
- QbitAI: “Embodied AI Data Infrastructure Goes on Sale at an Introductory Price of RMB 5,100: A New Approach to Robot Training Data,” covering the product launch, introductory price, and positioning as full-stack physical AI infrastructure.
- Science and Technology Daily: “China’s First Standardized Embodied Intelligence Dataset Platform Unveiled at the Pujiang Innovation Forum,” introducing the “Pujiang X” platform and humanoid robot dataset standards.
- National Data Standardization Technical Committee: “Typical High-Quality Dataset Case Study | Humanoid Robot Embodied Manipulation Dataset,” disclosing the scale of the training facility, number of robots, data volume, and governance processes.
- Beijing Economic-Technological Development Area: “Robot Data Facility Builds an Embodied Intelligence Data Factory,” introducing the scenario coverage, equipment configurations, and quality management system of Beijing’s humanoid robot data facility.
- CBNData: “I Became ‘Fuel’ for Robots: RMB 30 an Hour to Fold Clothes, While Selling Data Commands a Valuation of RMB 15 Billion,” presenting the current state of embodied AI data collection jobs and the commercial value chain.
- EqualOcean: An interpretation of “2026 Outlook for China’s Embodied Intelligence Data Collection and Data Industry,” outlining the embodied AI data industry chain and its underlying infrastructure.



