A single task burns a lot of tokens and a lot of compute, and capability on site has not grown in proportion.
The Capability–Deployment Paradox
Products ship faster every year. Getting them to work on a real site has not got any easier. Will that gap close on its own once the models get better?
What follows works through three angles, compute, data, and how that data gets collected, and along the way explains why we built microNature the way we did.
The Demo Is Improving. The Deployment Is Not
Demos out of the frontier labs look better every year. Success rates on real sites have not followed. Over the past few years the two curves have not converged at all. They are still pulling apart.
Split the same span of time into three layers (FIG.A2) and it gets fairly clear which layer the problem is stuck in.
Hardware is doing fine. Robot density in Chinese manufacturing stands at 470 units per 10,000 employees, third in the world, and new installations in 2023 came to 276,000 units, about 51 per cent of the global total (IFR). Capacity, cost and reliability all move forward quarter by quarter. The bottleneck has shifted to the brain, and not for want of parameters. What is missing is scene data and physics models built for a specific domain. Nobody has fully digitised the patch of real world a robot will actually work in, and nobody has handed it to the model in a physically accurate form. The next three sections take those in turn.
Cause One: Tokens and Compute Go Up, Capability Does Not Follow
The mainstream route for a general-purpose robot runs pixels → semantics → inference → action. You compress a continuous physical world into discrete tokens, then ask a large model to work out what to do. A single task burns a lot of tokens and a lot of compute, and what you get back on site does not rise anywhere near in proportion.
Cause Two: Plenty of Data, Almost None of It Useful
People point at the size of video datasets as evidence that physical intelligence is not short of data. That is the wrong measure. An hour of video may hold less usable physical knowledge than a thousand words of text. Video records correlation between pixels, not a physical quantity anything can act on. No force, no contact constraint, no link between an action and its outcome. The total really is enormous. The usable fraction is tiny.
Measured instead as high-quality embodied data fit for training, the shortfall comes out roughly like this.
Cause Three: Four Ways to Get Data, None of Them Faithful to the Physics
Four paths supply embodied data today. Each one falls short on physical fidelity, and each falls short differently.
| Path | How it works | Where the physics breaks down |
|---|---|---|
| Teleoperation | A person drives a real robot remotely while the motion is recorded | Physically true, but slow and expensive, and force, heat and contact are hard to capture in sync |
| Ego-centric video | Captured by a head-mounted or on-board camera | No global geometry, no environment parameters, heavy occlusion, and most of the physical state never observed at all |
| Synthetic data | Generated in bulk by a general-purpose simulation engine | Scenes are mostly templates and the physics does not match the real site, which is where most of the sim-to-real gap comes from |
| Video learning | Behaviour priors learned from human or web video | No action labels and no physical quantities, so what gets learned is what looks right, not what is physically achievable |
The Teleoperation Maths, and Why Throwing People at It Does Not Work
Of the four paths (A4), the one with the highest physical fidelity is teleoperation, and its cost will not move. Roughly CNY 2,000 an hour (Gartner). Against a requirement of 10 million hours, collection alone runs to about CNY 20 billion, before you have hired the operators, and that team is very hard to scale. As a business the numbers do not add up.
Put the Three Together: Robots Still Have No Scaling Law
The three sections above (A2 / A3 / A4) are each true on their own. The trouble is they are true at the same time, and they compound.
The total is enormous, the usable physical knowledge inside it is tiny, and an hour of video may not beat a thousand words of text.
Teleoperation, ego-centric video, synthetic data, video learning. All four sit below the fidelity threshold.
Build the missing layer of real-world physical infrastructure
The gap is a supply problem. Scene data and domain-specific physics models are both still too thin on the ground. Hardware is already out in front, waiting on a layer of infrastructure built from the real physical world, one that turns the deployment site into a digital base a robot can be trained on, validated against and scored in, and brings the cost of getting data down to between 1/20 and 1/200 of what it is now. That layer is what the microNature architecture is built for.
See the microNature architecture →Sources
- IFR, World Robotics. Robot density in Chinese manufacturing reached 470 units per 10,000 employees in 2023, third worldwide, against a global average of 162. New installations in China in 2023 came to 276,000 units, roughly 51 per cent of the world total.
- Gartner and industry estimates. Teleoperation data costs roughly CNY 2,000 per hour. The shortfall in high-quality data for humanoid robots runs to four orders of magnitude.
- Gasgoo Auto Research Institute. High-quality embodied data in existence stands at about 500,000 hours against roughly 10 million hours needed for general capability, a shortfall above 99 per cent. The volume of physical AI data is about 1/20,000 of the corpus available to language models. Simulation can bring the cost of a single data item down to between 1/20 and 1/200.
- Frost & Sullivan. The market for embodied intelligence solutions in China is projected at roughly CNY 142.6 billion by 2030.
Note: FIG.A3 and FIG.A5 illustrate trend and relative position. They are not measurements, and specific figures follow interviews and first-hand data.