Robonaissance

Robonaissance

The Mirror Test, Part 7: Living in the Past

A robot trained in a simulator without delay fell over at once. Brains work on news that is about a tenth of a second old. Acting turns small errors into a path.

Hugo's avatar
Hugo
Oct 08, 2026
∙ Paid

The Mirror Test, Part 7: Living in the Past

A robot trained in a simulator without delay fell over at once. Brains work on news that is about a tenth of a second old. Acting turns small errors into a path.

This is Part 7 of The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.


In 2018 Jie Tan and seven colleagues at Google Brain, X and Google DeepMind set out to teach a four-legged robot called Minitaur to walk by training it in a physics simulator. The simulator, Bullet, had a convenient property. “In Bullet, the motor commands take effect immediately and the sensors report back the state instantaneously.”

The first attempt was trained in that world, with no model of the robot’s actuators and no delay. In simulation, the learning system “did not come up with an agile gait. Instead, a slow walking gait was learned in simulation.” Then they put it on the real robot. “Moreover, the real Minitaur fell to the ground immediately due to the reality gap.”

The team then measured the robot they had. “The PD servo running on the microcontroller has a lower latency (3ms) while the locomotion controller executed on TX2 has a higher latency (typically 15-19ms).” They named “inaccurate actuator models and lack of latency modeling” as “two major causes of the reality gap,” and rebuilt the simulator with both. This time a gallop emerged. It reached roughly 1.34 meters per second in simulation and 1.18 on the real robot.

Fifteen to nineteen thousandths of a second is not much. The authors explain why it mattered: “This instantaneous feedback makes the stability region of a feedback controller in simulation much larger than its implementation on hardware. For this reason, we often see a feedback policy learned in simulation starts to oscillate, diverge and ultimately fail in the real world.” The paper puts the fall down to the reality gap as a whole, the missing actuator model and the missing delay together.

The same year, Xue Bin Peng and colleagues trained a Fetch robot arm to push a puck. A policy trained without varying the simulated physics succeeded in 51 percent of simulated episodes and in none of 10 trials on the real arm. A policy trained with memory and with 95 randomised physical parameters succeeded 89 percent of the time over 28 real trials. One of those parameters was the time between actions, “a simple model of the latency” of the physical system. Holding just that one fixed in training cut real success to 29 percent.

Neither robot was asked to do anything clever. They were asked to walk and to push. What made it hard is that their outputs did not end the computation. Each action moved a body, and the body’s new state was the next input.

What changes when intelligence has to act through a physical system, with delays and errors that cannot be left out?

An error becomes a path

Start with what goes wrong when a learned policy drives its own inputs.

The obvious way to teach a robot is to show it what an expert does and train it to copy. Stéphane Ross, Geoffrey Gordon and J. Andrew Bagnell explained in 2011 why that is not enough. Problems like imitation learning, “where future observations depend on previous predictions (actions),” break the usual assumptions of statistical learning. A copier that is wrong with some small probability on the expert’s situations can be wrong far more often on its own, “because as soon as the learner makes a mistake, it may encounter completely different observations than those under expert demonstration, leading to a compounding of errors.” Over a task of T steps the mistakes can grow with T squared, so a task twice as long can go wrong four times as badly.

They showed it in a video game, a kart race on a track floating in space, where a kart can fall off at any point. A learner trained only on the expert’s laps did not improve as more laps were collected, because the laps “do not help the learner to learn how to recover from mistakes it makes.” Their remedy, called DAgger, keeps “collecting a dataset at each iteration under the current policy,” so that the learner is trained on the situations its own driving creates. After 15 iterations, it produced “a policy that never falls off the track.”

Delay is the second way a loop bites. Kevin Black, Manuel Galliker and Sergey Levine put it plainly in 2025: “While a robot is ‘thinking’, the world around it evolves according to physical laws.” The large models now being used to control robots think slowly. A 3-billion-parameter model called π0 aims at a control step of 20 milliseconds. Another group’s version of the 7-billion-parameter OpenVLA, optimised for speed, achieved, as the paper reports it, “no better than 321ms of latency on a server-grade A100 GPU.” If the robot simply waits for the next batch of actions, the pauses “change the dynamics of the robot,” and training no longer matches what the robot meets. Delay turns into the first problem again: the robot ends up in states its training never showed it.

A table of delays. Minitaur quadruped, 2018: locomotion controller 15 to 19 milliseconds, motor servo 3 milliseconds. Pi zero on an RTX 4090: prefill alone 46 milliseconds against a 20 millisecond control step. Speed-optimised OpenVLA on an A100: no better than 321 milliseconds. Gemini Robotics: about 250 milliseconds end to end. Human motor system: sensorimotor feedback delays on the order of 100 milliseconds, fastest reflex loop 10 to 40 milliseconds, motor cortex command to muscle force about 40 milliseconds.
Delays reported for robot systems and for the human motor system. Each row times a different interval on different hardware, so the numbers are not a ranking. Sources: Tan et al. 2018; Black, Galliker and Levine 2025 (including an OpenVLA figure they report from another group); Gemini Robotics Team 2025; Franklin and Wolpert 2011.

Three repairs

The field’s answers fall into three families, and all three are in use now.

The first is to train on your own states. In a November 2025 report, Physical Intelligence wrote that “Policies trained with imitation learning are known to suffer from compounding errors,” citing Ross’s 2011 paper, and adopted “a form of such interventions, called human-gated DAgger.”

The second is to put the physics, including the delay, into training. That is what the Minitaur team did when it rebuilt its simulator with an actuator model and the measured latency, and the gallop emerged. Peng’s group varied 95 parameters. DeepMind’s soccer-playing humanoid, published in 2024, “added random time delays (10 to 50ms) to the observations to emulate latency in the control loop.”

The third is to stack loops. Minitaur’s own control has two layers with different delays: a motor servo at 3 milliseconds and a locomotion controller at 15 to 19. The soccer robot’s learned policy acts at 40 times a second, on top of a position controller that, the authors believe, “hides model mismatch from the agent by applying fast stabilizing feedback at a high frequency.” With direct motor-current control, they found, “the sim-to-real gap was too large.”

The problem has not gone away. A June 2026 report from Luqia Technologies tested vision-language-action models on a real UR5 arm. In one of its tests, fed the true state at every step, “the predicted trajectories closely match the ground-truth trajectories.” Run in a loop on its own outputs, “progressive deviation from the ground-truth trajectory appears over time.” Their summary: “offline stability does not guarantee closed-loop stability.” It is one preprint. It also puts the whole problem in one line.

User's avatar

Continue reading this post for free, courtesy of Hugo.

Or purchase a paid subscription.
© 2026 Robonaissance · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture