Robonaissance

Robonaissance

The Robot Kept Its Job

Three 2026 papers train dexterous robot hands from human video and keep the robot’s own demonstrations in. The two that tested cutting them, or their stage, did worse.

Hugo's avatar
Hugo
Oct 10, 2026
∙ Paid

This is Episode 10 of Reading the Frontier, a series of close readings of what the frontier labs publish.


On 20 April 2026, at Sequoia’s AI Ascent in San Francisco, Jim Fan of NVIDIA described how robots get their training data, and he was not kind about it. Three years of teleoperation, he said, had been the golden era: VR headsets tuned for streaming latency, rigs he likened to medieval torture devices. The whole apparatus was capped at 24 hours per robot per day, and in practice, he said, closer to three, because robots throw tantrums.

Then he gave his team’s count. Their model, EgoScale, was pretrained on about 21,000 hours of egocentric human video, recorded by cameras on people’s heads, with no robot data at all. The step that taught it to act took 50 hours of motion capture from people wearing gloves and 4 hours of teleoperation, under a tenth of one percent of the mix. Within a year or two, he predicted, teleoperation would shrink to almost nothing and egocentric video would become the main diet of robot learning. Then he asked the room for a moment of silence for it.

The four hours stayed in the recipe, by his own count and in his team’s paper. A recipe, here, is the data and the training stages a paper uses to build its model. EgoScale’s authors, Fan the last of fifteen, call the stage that holds them critical: it anchors what the model learned from video to this robot’s sensing and actions. The hand that recipe drives has twenty-two degrees of freedom, twenty-two independent ways to move.

EgoScale and two other 2026 papers, HuRo and EgoWild2Dex, train dexterous hands from human video, dexterous being the papers’ own word, and all three keep the robot’s own demonstrations in the training data.

Human video took the hours. On the dexterous hands, the robot kept its job.

That contradicts a belief, that video of people can now do robot data’s job and the robot hours left in the recipes are a vestige on the way out. Phantom, Figure and Sunday each say a version of it for their own tasks, and the reader supplies the rest by dropping the hand each result was tested on. A small share of the hours is a fact. The step from a small share to no job is the one that fails.

The claim comes in two strengths, and no written form names the hand

Teleoperation does not scale, because a person drives one robot for a few hours a day, and the claim that video can replace it is older than the papers that test it. Between March 2025 and April 2026, six parties made some version of the claim public. Phantom, Figure and Sunday said no robot data at all, for their own tasks. A second paper on humanoid policies, Skild and Fan claimed scalability with a small robot set, or a shrinking share of teleoperation.

The word substitute appears too. HumanNet, a one-million-hour video corpus, and HumanScale, a pretraining study built on it, use it for egocentric video as a source for pretraining, and HumanNet’s own limitations bound it to priors, not one-to-one replacement. HumanEgo, a policy trained on video from smart glasses, carries it over to policy learning. EgoScale’s authors say something different: “Rather than replacing robot data, large-scale human demonstrations dramatically amplify its effectiveness.”

None of these written forms says which hand it means. A dexterous hand, the papers’ word, has articulated fingers counted in degrees of freedom: EgoScale’s Sharpa hands have 22 (the paper adds that its models also transfer to robots with fewer), HuRo’s robot has two hands of 15, and EgoWild2Dex’s has 6 per hand with the finger count unstated. A parallel-jaw gripper, the phrase Phantom and HumanEgo use, offers a pose and one opening width.

Four kinds of data keep coming up, and they are different things. Video only means recordings of people at work: some from a camera on the head, Phantom’s from a camera fixed in the scene, HumanNet’s from internet video. Sensorized human data adds trackers and gloves, or a glove or exoskeleton built to match a robot hand. Robot data is what a robot recorded, and teleoperation, the prediction’s word, is a person driving the robot to record it. Rodney Brooks argued in September 2025 that visual data alone cannot teach dexterity, on grounds of touch and force.

Each paper bounds its own claim by its hand. The headlines drop the hand.

EgoScale’s one demonstration is not one

Start with the shirt. EgoScale folds a shirt with a success rate of 0.88 after seeing the robot do it once. That sounds like learning from video. What else is in the recipe?

The simplest case is the one-shot setting. For each task the robot is shown one demonstration of its own, and 100 human demonstrations of that task come with it, for each bottle shape separately in the bottle task; the paper’s account of this experiment does not say how those were captured. The same recipe unscrews bottle caps at a success rate of 0.55 across three bottle shapes. Its ordinary tasks, for contrast, get 100 teleoperated robot demonstrations each, 20 for shirt rolling.

Now add the layers one at a time. First comes pretraining: 20,854 hours of head-camera video and no robot. Then the aligned stage. As the paper describes its setup, a person in Manus gloves, with Vive trackers recording wrist pose, is filmed by the same camera configuration as the robot’s, at matched viewpoints, and performs each of 344 tabletop tasks about 30 times, while the robot is driven through the same task about 5 times. The total is about 50 hours of human data and 4 of robot data. What the stage supplies is a translation into this robot’s terms: what its cameras see and what its hands do, on tasks a human also did. The 21 tracked keypoints of the human hand are retargeted into the robot hand’s joint space, its 22 degrees of freedom, within joint limits.

Now take pieces away. The paper reports that models omitting either the large-scale human pretraining or the aligned stage fail in the one-shot setting. And the stage goes whole: version 1 of the paper, from February 2026, reports no comparison that takes out only the robot trajectories. Whether the four robot hours or the fifty human hours do the anchoring is not in its experiments.

By the hours, the robot is a rounding error: 4 against 20,854 is 0.019%, and the 54 hours of the aligned stage are 0.26% of the 20,908 hours of both. Figure 1 draws these hours. The robot’s post-training demonstrations come on top, in counts the paper does not convert to hours. By the paper’s own account the stage is the anchor.

Two panels of horizontal bars. The first, to one scale, shows 20,854 hours of human pretraining video as one long bar with the 54 hours of the aligned stage as a sliver at its end. The second enlarges the aligned stage: about 50 hours of human data and 4 hours of robot data.

Caption. Hours are the paper’s own. EgoScale’s post-training robot demonstrations (100 per ordinary task, 20 for shirt rolling, 25 per bottle, 1 in the one-shot setting) lie outside the 4 hours and have no stated hours.

User's avatar

Continue reading this post for free, courtesy of Hugo.

Or purchase a paid subscription.
© 2026 Robonaissance · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture