Simple AI releases 2,000 hours of robot training data collected without robot arms
Founder Xiaofei Li is applying an autonomous-driving data playbook to manipulation, with limited parity claims built on more than 10 times as many human demonstrations.
By RuntimeWire Staff · Published
Primary source: PR Newswire
Why it matters
Robot teleoperation makes manipulation data accurate and expensive. HiFi-UMI suggests portable human capture can substitute for it in some post-training pipelines, though Simple AI used more than 10 times as many demonstrations.

Xiaofei Li, founder and CEO of Simple AI, released a 2,000-hour robotics dataset on August 20th that eliminates the need for a target robot or robot teleoperation during data collection. The HiFi-UMI system records people performing two-handed manipulation tasks with portable sensor-equipped grippers, then converts those demonstrations into trajectories for training real robot arms.
The release turns a lesson from Li's autonomous-driving career into a robotics product. Public investor material says Li earned bachelor's and doctoral degrees from Tsinghua University and helped develop an L4 autonomous-driving architecture before founding Simple AI. His current bet follows the data-loop approach that shaped autonomous vehicles: collect behavior in real environments, automate quality control, feed the resulting data into models and test the models back in the physical world.
In an August 20th announcement, Li said Simple AI wanted to determine whether data fidelity could make robot-free demonstrations suitable for deployment-oriented training. The accompanying technical report, first posted on July 28th, reports approximate performance parity with conventional robot teleoperation across three policy architectures and four tabletop task groups.
A data factory without the robot arm
Robot teleoperation produces useful training data because every recorded action already reflects a specific machine's geometry, sensors and physical limits. It also ties each hour of collection to a robot arm, a control rig and an operator. That makes expansion across homes, hotels, warehouses and other varied environments slow and capital-intensive.
The Universal Manipulation Interface, a robot-free human-demonstration system introduced in a 2024 paper by researchers affiliated with Stanford University, Columbia University and Toyota Research Institute, offered another route. Its handheld grippers let people demonstrate tasks in natural settings without transporting a robot to every location. Simple AI's research focuses on the remaining fidelity problems, including trajectory accuracy, synchronization, camera coverage and the relative position of two grippers.
HiFi-UMI places a stereo camera and inertial sensors on the operator's head and two non-parallel fisheye cameras on each hand module. Simple AI reports approximately 3 millimeters of local end-effector accuracy, sensor synchronization below 40 microseconds and about 200 degrees of visual coverage around each hand. A shared hardware trigger aligns the cameras and inertial sensors.
After capture, Simple AI reconstructs each trajectory and attempts to replay it in simulation. Simple AI reports approximately 98% pass rates for both trajectory reconstruction and simulation replay validation in its release materials. This automated rejection step matters commercially because cheaper collection loses much of its advantage if engineers must manually repair or inspect every episode.
The HiFi-UMI-2K dataset contains 2,000 hours and more than 482,000 episodes collected across over 110 scenes, according to its dataset card. It includes synchronized multi-view video, bimanual trajectories, gripper states, language annotations, subtask boundaries and quality-control metadata. Simple AI released it under a CC BY 4.0 license, which permits commercial use with attribution.
HiFi-UMI-2K represents less than one-tenth of the larger corpus Simple AI says it has processed. The project's materials claim more than 20,000 hours and 4.32 million episodes collected across over 480 scenes. Opening 2,000 hours gives outside researchers enough material to train and test models while leaving most of the reported corpus unreleased.
The parity claim has a denominator
The report evaluates three policy backbones spanning vision-language-action and world-action-model approaches, across four bimanual tabletop tasks. Compared with policies post-trained on robot teleoperation data, HiFi-UMI-only policies produced aggregate success-rate differences of -2.5, +3.1 and -0.6 percentage points, respectively.
On a precision-insertion task, the strongest HiFi-UMI-only policy reached 85% success. The teleoperation baseline had an apparent environmental advantage because its demonstrations were collected in the evaluation scene, while the robot-free examples came from other locations.
Those results require a careful denominator. The vision-language-action comparisons used about 3,200 HiFi-UMI trajectories per task and roughly 300 teleoperation trajectories. Simple AI therefore used more than 10 times as many robot-free demonstrations. The relevant commercial test is whether those additional human demonstrations are sufficiently cheaper and faster to produce. Per-trajectory data efficiency remains unresolved.
The robot used for evaluation also shared the capture system's gripper and wrist-camera configuration. Its arm kinematics differed, but the test did not span radically different hands, sensor layouts or robot bodies. The authors limit their conclusion to the tested models, tasks and conditions. HiFi-UMI provides evidence that the real-robot post-training anchor can be removed in this setup, rather than a general finding that manipulation models no longer need robot-generated data.
The pre-training results point toward a second use for Li's data operation. Simple AI says pre-training one policy on 4,000 hours from the broader corpus reduced average offline action-prediction error by 41% across 10 unseen tasks. On the evaluated StarVLA-QwenPI setup, it increased real-robot success by 18.1 percentage points when the amount of task-specific post-training data was held constant.
Li is building the data loop before the household robot
Simple AI is based in Beijing. Li has framed embodied AI as an industry moving from technical demonstrations toward repeatable products and deployment, a transition he previously watched unfold in autonomous driving.
Capital has followed that thesis, although the public financing record is difficult to reconcile. An April 2026 report described a separate several-hundred-million-yuan round led by Linear Capital and Puhua Capital, with Junshan Capital, Shunwei Capital and BV Baidu Ventures participating. Separate June financing reports described a Pre-A round of several hundred million yuan backed by Didi, Plum Ventures, Col Capital, Linear Capital, CCV and Puhua Capital, but the available material does not establish a lead investor or how the two rounds relate.
The HiFi-UMI release shows where part of that capital is going. Simple AI is assembling capture hardware, reconstruction software, annotation, simulation validation, model training and robot deployment into a single production chain.
The competitive field is already splitting between robot-heavy and robot-free collection. The DROID dataset uses robot-based collection and teleoperation equipment, while commercial suppliers such as Trillion Robotics are selling portable handheld UMI hardware. Simple AI is betting that tighter sensing and automated replay can make portable demonstrations precise enough for the final, task-specific stage of training.
Releasing 2,000 hours under a permissive license is also a distribution move. If researchers adopt the format, reproduce the results or build training tools around HiFi-UMI, Simple AI's collection method gains credibility beyond its own robots. The larger corpus Simple AI says it has processed could then become more valuable as an extension of an emerging public standard.
Li's immediate achievement is narrower and more useful than the broad promise of a household robot. Simple AI has built a sizable data-production system and published a benchmark showing that, under controlled conditions, people carrying grippers can generate enough high-fidelity supervision to train a robot without operating one during collection. The economics of that substitution, and its ability to survive different hardware and contact-heavy tasks, will determine whether HiFi-UMI becomes infrastructure or another strong laboratory result.