Variation Kills Automation: Teaching Robots Stiffness | Sze Yuan Cheong & Elijah Ong, Devol Robots

Variation Kills Automation: Teaching Robots Stiffness | Sze Yuan Cheong & Elijah Ong, Devol Robots

Sze Yuan Cheong and Elijah Ong are the co-founders of Devol Robots, which builds the AI that lets industrial robots handle contact: insertion, assembly, and the other jobs factories still do by hand. Sze, the CEO, started and ran factories for a decade before he touched a robot. Elijah, the CTO, turned down a Stanford PhD to help a three-person startup in Austin build a force-controlled arm from scratch. Disclosure: Optim VC is an investor in Devol and I advise the company.

Summary

Devol Robots is named after George Devol, inventor of the Unimate, the first industrial robot arm. Sze's view is that the field hasn't changed in the 60 years since: robots still repeat the motion they were programmed to make. "We wanted to be the other Devol," he told me, "the one that changes the status quo."

On a factory floor, that status quo means the end of the line is automated and people do almost everything between the machines. Sze explains why in three words: variation kills automation. Devol's bet is that robots can handle variation once they learn how stiff or soft to be at each moment of contact, and that robot AI trained on pixels is blind to that information.

What factories still do by hand

In a large plant, the end of the line is automated: packaging, palletizing. The work between the machines is still manual, whether that means plugging parts together, assembling a phone, or moving metal parts from station to station. "All the things that have a ton of variation, humans are doing all of that," Sze said.

Factories have a standard workaround. If a fixture puts a part in the same absolute position every time, a machine can take the task, though someone still retunes it daily to correct drift. The approach breaks when SKUs change often, because every variant needs jigs and fixtures designed, integrated, programmed, and fine-tuned. Sze called that work "a huge industry by itself," and one with a shrinking workforce, since engineers his age go into computer science and AI. "No one really wants to do that work anymore."

Twenty jigs and ten microns

Devol's live deployment is with an optics manufacturer. The big processes on that line are already automated: washing, ultrasonic cleaning, coating. "What is automatable is, for now, already automated," Sze said. The manual work sits in between. A lens passes through about 20 steps, each with its own jig for every lens variant, and people move the lenses from jig to jig.

The jigs don't match. A washing jig is a loose, hand-tuned fit meant to take everything, and the next jig may have 10 to 20 microns of tolerance. A small glass lens going from one to the other tends to stick, and pushing a stuck lens scratches it, which means a reject. Sze said an operator needs six to nine months of training to do the transfer. "I couldn't do it."

What a camera can't see

I put a scenario to Elijah: a robot arm resting on a table, then the same arm pressing into it with 50 newtons. A camera records the same picture twice. Pixels, he said, give a robot spatial understanding and tell it nothing about the interaction. "The robot is sort of touching the table, but the image doesn't tell you what it is actually doing."

The missing quantities are stiffness and damping. In classical robot control, stiffness lets a controller regulate force and position together, and damping does the same for velocity. Sze made it concrete. When you fish your keys out of your pocket you do it without looking, fast, and correctly every time, yet you could never describe the path your fingers took. That feel is what stiffness and damping capture.

A 3 a.m. phone call

Elijah spent his graduate years on robotic grasping. Stanford offered him a PhD in the same direction and he turned it down. "I felt that direction wasn't leading anywhere at the time, and that I had to get out there and try to find the answer myself."

He found it at a startup in Texas working on impedance control, where he helped build an arm from scratch around series elastic actuators. The work convinced him that impedance (forces, stiffness, and damping together) is the physical representation a robot needs to understand contact. He called Sze at 3 a.m. to say they could finally solve robotic manipulation.

Their first attempt, a model trained directly on impedance data, didn't learn well. Elijah traced the problem to how robot data reaches a neural network, flattened into plain vectors. In Euler angles, a pose that rotates to 180 degrees can snap back to zero, so the numbers jump while the physical motion stays smooth. "Robot data actually lives in curved space," he said, and each signal should be projected onto its own manifold before a model learns from it. He sees the flattening as one reason vision-language-action models need so much data to learn a single pick-and-place.

An Ethernet plug, phase by phase

I asked Elijah to walk through plugging in an Ethernet cable. Carrying the cable toward the port is free-space motion, so the arm stays soft. Near the port it stiffens so it can be guided in, with enough damping to slow down and react when the plug touches the port's edge. During the slide it holds position tightly along the insertion axis. Devol's model learns that sequence as a compliance schedule, which the team calls a visuo-impedance map: for each moment and direction, how much stiffness and damping to command.

Elijah said about 1,000 hours of pretraining is the minimum they have found for a model to pick up this kind of physical understanding. In the benchmark from their first paper, IWM 1.0, Sze said Devol's model got 20 demonstrations per task and two leading comparison models got 100. By his account Devol's model succeeded more than 90% of the time and the comparators landed near 30% and 40%.

From one demonstration to production

When a manufacturer calls with a new insertion task and already owns a supported robot, an operator demonstrates the task once and the model identifies the object and the objective. The robot then spends about an hour trying the task on its own. "It fails a bunch of times, and it succeeds," Sze said, and that data is used to post-train a policy grounded in the physics of the task. A customer without a robot can collect 20 to 30 minutes of data with UMI grippers and see a working policy before they deploy.

Most installed arms are position-controlled, and Elijah says the model runs on them too. On a force-controlled arm, or an industrial arm with a force-torque sensor at the end effector, the model reads impedance directly. Everywhere else it uses what he calls derived impedance: it extracts from the trajectories where the motion was soft and where it was stiff, and uses that signal to set the robot's acceleration and deceleration.

The model breaks, Elijah said, on something like threading a needle. The thread is too soft to return a usable force signal, so the task falls back on position control and vision.

The bitter lesson, and the lesson AI forgot

I put the bitter lesson to them: scale beats hand-built structure, and the VLA labs have billions of dollars. Elijah doesn't argue against scale. "I'm not saying scaling is bad. It's that we need to scale in the right paradigm." He wants the field to treat manipulation as a robotics control problem and start from what a robot needs to know about the physical world. With a physical prior and the right geometry, he argues, scaling laws hold and cost far less.

Sze's version comes from classical robotics, which always worked by finding the correct abstraction and building on it. "That lesson has been forgotten." Stiffness and impedance are Devol's abstraction, and the architecture exists to carry it. He previewed a forthcoming paper on a new architecture, Devol One, that feeds the output of a latent world model into the action decoder instead of running the two as separate routes.

Their closing takes pointed the same way. Elijah: "We really shouldn't scale from pixels anymore." Sze thinks AI researchers are slowly rediscovering what classical robotics already knew, and that the field's effort goes into planning while execution on the robot goes unsolved: "dare I say, no one is actually trying to solve it." Elijah's example is the action head. Diffusion, flow matching, and autoregressive heads all run open loop. "I ask the robot to go there, and it just goes."

Don't build everything yourself

Sze's advice to robotics founders comes from Devol's early years. They started out convinced that the best model required owning the whole stack, so they built their own actuators, drivers, and custom firmware. "As a startup, you can't really do that. We learned that the hard way." His advice now is to work with other companies and pick the one problem you are good at.

Devol wants to hear from researchers, roboticists who care about control, and manufacturers with a contact problem to solve. Sze extended the invitation to anyone interested in robotics, experience or not. The team, he said, is built by looking for outliers.

Full episode available on:
YouTube
Spotify
Apple Podcasts

Learn more about Devol Robots: www.devolrobots.ai