SheetMN-A1
TitleThe capability–deployment paradox
Issue2026.08
Scopecompanion sheet

The Capability–Deployment Paradox

Products ship faster every year. Getting them to work on a real site has not got any easier. Will that gap close on its own once the models get better?

Insights Three dimensions · compute / data / paths Reads with MN-00
In brief Two things are in short supply: data from real scenes, and physics models built for a specific domain. That shortage is what opens the gap between what a system can do in a demo and what it can do on site. It is the problem microNature works on.

What follows works through three angles, compute, data, and how that data gets collected, and along the way explains why we built microNature the way we did.

← Back to the main file · microNature — Physical Intelligence Infrastructure
00Contents
A1Phenomenon

The Demo Is Improving. The Deployment Is Not

Demos out of the frontier labs look better every year. Success rates on real sites have not followed. Over the past few years the two curves have not converged at all. They are still pulling apart.

TIME → CAPABILITY LAB DEMO capability FIELD DEPLOY capability THE GAP data ↑ · tokens ↑ · deployment ≈ TODAY
FIG.A1lab demonstration against field deployment, and a gap that keeps widening

Split the same span of time into three layers (FIG.A2) and it gets fairly clear which layer the problem is stuck in.

TIME → READINESS hardware capacity and performance iterating by the quarter · density 470 per 10,000 workers the brain · foundation models improving · not keeping up with hardware scene data · domain-native physical modelling close to zero · barely supplied at all the real bottleneck sits here
FIG.A2three layers of supply moving at different rates, hardware fast, models in the middle, scene infrastructure all but absent

Hardware is doing fine. Robot density in Chinese manufacturing stands at 470 units per 10,000 employees, third in the world, and new installations in 2023 came to 276,000 units, about 51 per cent of the global total (IFR). Capacity, cost and reliability all move forward quarter by quarter. The bottleneck has shifted to the brain, and not for want of parameters. What is missing is scene data and physics models built for a specific domain. Nobody has fully digitised the patch of real world a robot will actually work in, and nobody has handed it to the model in a physically accurate form. The next three sections take those in turn.

A2Compute

Cause One: Tokens and Compute Go Up, Capability Does Not Follow

The mainstream route for a general-purpose robot runs pixels → semantics → inference → action. You compress a continuous physical world into discrete tokens, then ask a large model to work out what to do. A single task burns a lot of tokens and a lot of compute, and what you get back on site does not rise anywhere near in proportion.

TOKEN / COMPUTE INPUT → deployment capability 10× 100× 1000× grey bars = scale of input (indicative) capability growth ≈ flattening input ↑1000× capability ↑ limited
FIG.A3input grows by orders of magnitude, deployment capability flattens, and the two do not move together
What this means A bigger model and more tokens do not turn into deployment capability by themselves. Compression mostly throws away contact, force, friction, deformation and material, and those are precisely what decide whether a motion can be completed. The compute goes into recovering meaning. The physics never gets recovered at all.
A3Data

Cause Two: Plenty of Data, Almost None of It Useful

People point at the size of video datasets as evidence that physical intelligence is not short of data. That is the wrong measure. An hour of video may hold less usable physical knowledge than a thousand words of text. Video records correlation between pixels, not a physical quantity anything can act on. No force, no contact constraint, no link between an action and its outcome. The total really is enormous. The usable fraction is tiny.

Measured instead as high-quality embodied data fit for training, the shortfall comes out roughly like this.

≈ 500,000 hours high-quality embodied data in existence industry estimate ≈ 10 million hours data needed for general capability industry estimate gap > 99% about 20× apart for comparison: the physical corpus is around 1/20,000 the size of a language model corpus, and physical intelligence has no ready-made foundation of the kind language models started from
FIG.A4high-quality embodied data, what exists against what is required, a gap above 99%
Where language models started The internet spent thirty years accumulating high-quality text that any team could pick up and start training on. Physical intelligence has no such stockpile. No internet, no Transformer. Without a foundation of real physical data, scaling physical intelligence is just as far off.
A4Paths

Cause Three: Four Ways to Get Data, None of Them Faithful to the Physics

Four paths supply embodied data today. Each one falls short on physical fidelity, and each falls short differently.

PathHow it worksWhere the physics breaks down
TeleoperationA person drives a real robot remotely while the motion is recordedPhysically true, but slow and expensive, and force, heat and contact are hard to capture in sync
Ego-centric videoCaptured by a head-mounted or on-board cameraNo global geometry, no environment parameters, heavy occlusion, and most of the physical state never observed at all
Synthetic dataGenerated in bulk by a general-purpose simulation engineScenes are mostly templates and the physics does not match the real site, which is where most of the sim-to-real gap comes from
Video learningBehaviour priors learned from human or web videoNo action labels and no physical quantities, so what gets learned is what looks right, not what is physically achievable
PHYSICAL FIDELITY fidelity threshold that deployment requires Teleoperation true but costly · will not scale Ego video physical state unobservable Synthetic data physics not aligned to the site Video learning pixel correlation only
FIG.A5all four sit below the threshold (relative and indicative, not measured)
What they share All four are missing the same thing: Physical quantities aligned to the real deployment site. That is a gap in supply, not an engineering detail inside any one of them. Nobody has modelled the physics of a real scene, so there is nothing there to supply.
A5Cost

The Teleoperation Maths, and Why Throwing People at It Does Not Work

Of the four paths (A4), the one with the highest physical fidelity is teleoperation, and its cost will not move. Roughly CNY 2,000 an hour (Gartner). Against a requirement of 10 million hours, collection alone runs to about CNY 20 billion, before you have hired the operators, and that team is very hard to scale. As a business the numbers do not add up.

COST PER HOUR OF DATA (CNY) ≈ CNY 2,000 teleoperation capture Gartner · labour-intensive ≈ CNY 10 – 100 domain-native simulation cost per item falls to 1/20 – 1/200 a 20 to 200 fold difference one condition: the simulation has to be built on the physics of the actual domain. A generic template saves you nothing
FIG.A6the cost of acquiring data, teleoperation set against domain-native simulation
The condition Only simulation brings that cost down, and only when its physics matches the real site. When it does not, the data is unusable and the saving never arrives. That is what domain-native means.
A6Conclusion

Put the Three Together: Robots Still Have No Scaling Law

The three sections above (A2 / A3 / A4) are each true on their own. The trouble is they are true at the same time, and they compound.

① Tokens and compute

A single task burns a lot of tokens and a lot of compute, and capability on site has not grown in proportion.

② Data efficiency

The total is enormous, the usable physical knowledge inside it is tiny, and an hour of video may not beat a thousand words of text.

③ Data accuracy

Teleoperation, ego-centric video, synthetic data, video learning. All four sit below the fidelity threshold.

④ Result Stack all three and the scaling law that arrived for language still has not arrived for robots. The gap keeps showing up on site, in one form after another.
Where this leads

Build the missing layer of real-world physical infrastructure

The gap is a supply problem. Scene data and domain-specific physics models are both still too thin on the ground. Hardware is already out in front, waiting on a layer of infrastructure built from the real physical world, one that turns the deployment site into a digital base a robot can be trained on, validated against and scored in, and brings the cost of getting data down to between 1/20 and 1/200 of what it is now. That layer is what the microNature architecture is built for.

See the microNature architecture →
App.Sources

Sources

  1. IFR, World Robotics. Robot density in Chinese manufacturing reached 470 units per 10,000 employees in 2023, third worldwide, against a global average of 162. New installations in China in 2023 came to 276,000 units, roughly 51 per cent of the world total.
  2. Gartner and industry estimates. Teleoperation data costs roughly CNY 2,000 per hour. The shortfall in high-quality data for humanoid robots runs to four orders of magnitude.
  3. Gasgoo Auto Research Institute. High-quality embodied data in existence stands at about 500,000 hours against roughly 10 million hours needed for general capability, a shortfall above 99 per cent. The volume of physical AI data is about 1/20,000 of the corpus available to language models. Simulation can bring the cost of a single data item down to between 1/20 and 1/200.
  4. Frost & Sullivan. The market for embodied intelligence solutions in China is projected at roughly CNY 142.6 billion by 2030.

Note: FIG.A3 and FIG.A5 illustrate trend and relative position. They are not measurements, and specific figures follow interviews and first-hand data.