Moving a Task From One Robot Body to Another | Aurora Feng, Neural Motion

Moving a Task From One Robot Body to Another | Aurora Feng, Neural Motion

San Francisco, September 22, 2026

Aurora Feng asked the foundation model labs a question I had not thought to ask. What do you do with the data you collected on your last generation of hardware?

The answer, more often than not, is nothing. They dump it, or they leave it somewhere and stop looking at it. The gap between one hardware revision and the next is wide enough that demonstrations recorded on the old arm will not train a policy on the new one, so teams that spent months and money on high quality teleoperation data write it off and collect again.

The usual framing treats cross-embodiment as a coordination failure between companies, the kind of thing a standards body eventually fixes. Aurora's conversations put it inside a single company, between two versions of the same robot, on a roadmap written by one team. That widens who needs this.

Aurora is founder and CEO of Neural Motion, and she was my guest on the latest episode. The full conversation is on YouTube.

The ceiling is set at collection time

Start with the arithmetic she walks through on the show. A lab collects millions of hours across five arms. Each arm has 200 tasks. The model learns those 200 tasks, because that is what it was fed. Ask for the 201st and the policy has nothing to reach for.

Two hundred tasks covers a lot of useful work, and she says so. Her point is that somebody fixed the number before training started, on the day they decided what to record. No amount of compute or architecture work moves it, because the ceiling sits in the corpus, and the corpus closed when collection stopped.

Realistically, you can never collect the entirety of tasks in the world. There are always going to be out-of-distribution tasks.

Every lab building a robot foundation policy stands under the same ceiling at a different height. The standard way out is to collect more, either through teleoperation, which scales slowly and costs a lot, or through egocentric human video, which scales fast and brings a problem I will come back to.

Co-training helped, and it did not remove the ceiling

The field's answer so far has been to pool everything. Per the Open X-Embodiment paper at ICRA 2024, the collaboration assembled data from 22 robots across 21 institutions, covering 527 skills and roughly 160,000 tasks, and the RT-X models trained on it showed positive transfer across platforms. Octo and OpenVLA followed the same logic. Throw every robot dataset you can find into one pile, train one model, and let it work out what generalizes.

Aurora's objection is that nobody knows why it works when it works. She frames it as two distinct ways at the problem, one from the data side and one from the model side, and the field has spent nearly all of its attention on the model side. Co-training hands the model an unexamined mixture and hopes. Nobody knows what happens if you fix the data recipe first, before anything enters the pre-training pile, because the foundation model teams have not had the time to run that experiment.

Neural Motion is built to run it. The model takes a task recorded on one robot and produces that same task on a robot it has never seen, video and action together, with no retraining and nothing collected on the new body. If that works, the mixture stops being whatever your vendors happened to collect and becomes something you choose.

Retargeting solves a different problem

Smart people have told me cross-embodiment is a solved kinematics problem. You have the poses and the joint positions, you retarget through the URDF, you are done. I put that to her directly.

Kinematics solves correspondence in configuration space. What you need is correspondence in observation space and action space, and those are different problems. Two robots performing the same task look different to a camera and move through different action spaces, and neither gap closes because you mapped one joint configuration onto another.

If all I cared about was moving an end effector from embodiment A to embodiment B, then sure, classical retargeting could be useful. But that's not the goal.

The goal is data a policy can learn from. Moving an end effector correctly and producing an episode worth training on are two different bars, and only the second one has a customer.

She gives real-to-sim-to-real the same treatment, and the argument runs on losses. Going into simulation, you simplify the real environment into rules, and the simplification costs you something. Coming back, you have to re-complicate it to fit reality, and that costs you again. Close the loop and the data that comes out the far end is often too far from what went in to be usable. Her approach keeps teleoperation data as teleoperation data throughout. Only the body changes.

Human video without a robot reference

The cheap answer to the data bottleneck right now is human video. Hands do everything, task coverage is enormous, and nobody is paying an operator to drive an arm at human speed.

Aurora's objection is that human data leaves the robot embodiment undetermined. You get wide coverage of what tasks exist and no grounding in what a robot doing them should look like, which means no reference to measure the human data against. Scale it to millions of hours and the marginal value of each additional hour keeps falling, because the missing piece was never volume.

Her read is that the ordering matters. Teach a model the mapping between task and embodiment first, and human video becomes usable as one more source body. Skip that step and you are scaling something the model cannot anchor.

The community came before the company

Aurora is not a researcher. She started Saturday Robotics because the events available in San Francisco were networking events where everybody pitched, and she wanted a room where people debated research instead, with no VCs raising and no founders selling. Three months in, by her count, the group had 2,600 researchers in it.

Then the room started returning things to her. The frustration she kept hearing in different forms became the thesis. One of the reading club's first guest lecturers turned out to have written the paper she had been carrying around, having gone through thousands on her own, and the rest of the team formed around him.

She built the community before she built the team, and the team came out of the community. For a first-time founder recruiting PhDs out of top programs, that is a distribution advantage built in three months with no capital, and I think founders underrate it.

What I will be watching

The argument is good. The evidence is early, and Aurora says so on the show. Her framing is a hypothesis, and she uses the word. Fix the data recipe before pre-training and the result should improve. Nobody has measured it.

The test that matters is whether a policy trained on converted data performs like a policy trained on data physically collected on that body. Rendering an episode that looks right and producing an episode that trains well are separate bars, and only the second one is worth anything commercially. That measurement is what the technical report needs to carry, and it is what I will read it for.

The second thing I will watch is whether the advantage survives contact with labs holding far more compute and capital. The structural claim is that converting data is cheaper per unit of task diversity than collecting it. If that holds, the constraint moves from collection to compute, and Aurora expects compute to be the bottleneck left standing in five years. Being right about that and being the one who benefits from it are different outcomes.

A few things to remember

The ceiling is set at collection time. Whatever your robot recorded is the neighborhood your policy works in, and that was decided before training started.

The problem lives inside single companies. Labs cannot reuse demonstration data across their own hardware generations, so the buyer and the problem already share a roof.

Retargeting solves configuration space. The gaps that matter are in observation space and action space, which it does not touch.

Human video needs a robot reference before it is worth scaling. Coverage without grounding has falling marginal value

Aurora is on LinkedIn, and Neural Motion is at neural-motion.org. They are hiring around in-context learning, cross-embodiment and pre-training for robot foundation models.

---
Also see the episode on Apple Podcasts and Spotify