A Point of View on Physical AI

A Point of View on Physical AI

What I believe about robotics and physical AI right now.

September 29, 2026. Image credit: Chestnut Robotics.

Every week someone asks me what I think about a data company, a humanoid round, or whether the market has run ahead of the machines. I answer the same questions in DMs, at dinners and on calls, and the answers drift a little each time. This page fixes them in one place. Most of the evidence comes from the founders and researchers covered on this newsletter.

The market

1. We are late in a hot cycle, and the cycle will turn

Physical AI raised about $36 billion in 2025 and passed $55 billion seven months into 2026, per Adrian Macneil's count at Actuate, with the average round up from $31 million to $48 million even after setting aside Waymo's February raise. Vijay Chattha of VSC reminded a room in September that two or three years ago robotics was a dirty word and founders flew around the planet to close a seed round. Now the category has a name and a queue.

Most of this money is buying something real. The people who built the last wave are seeding the next one. The Diffusion Policy and ALOHA authors running Sunday, Sanja Fidler launching Veeda, Sebastian Thrun teasing a stealth robotics company two decades after Stanley. Three-dimensional sensors, actuators and compute have crossed the price line where a small team can put a full autonomy stack on a $4,000 arm. Open hardware and tooling let that team go from unboxing to model training in days. Chao Cao of Sancho calls this the eighties-and-nineties-of-computing moment.

Valuations at the top of the stack have also priced in a future the deployments do not yet support, and I expect this to resolve the way it resolved in self-driving. Gary Chen of Raise Robotics drew the comparison: 2015 brought enormous capital and unrealistic timelines, the reckoning came around 2020, and Waymo works today all the same. Bilal Zuberi described the capital side at AUTONOMOUS 2026: venture has split into small funds writing lottery tickets and multi-billion-dollar platforms for whom a large check is an optionality bet. When timelines slip, optionality capital moves on. James Hardiman gave the mechanics for the companies caught in between: raise twenty million when the business needed five, scale burn to match, miss the growth the next round required, and unwind.

So my base case is a correction in the next two to three years, concentrated in the layers priced for physical AGI, followed by consolidation: acquisitions of good teams by strategics and by better-capitalized peers, down rounds, a few shutdowns, and then the survivors work. The companies that raised what they needed and kept burn flat will buy the market when it arrives. Capital that came with the narrative will leave with it. The founders and engineers will stay, because the problems are real and the tools keep getting cheaper.

2. Valuations are a layer question

Rounds at the model and humanoid layer are priced for general-purpose autonomy. Rounds in tooling, data infrastructure and applied deployment are priced for revenue, often revenue that already exists. The correction lands on the first group. Peter Ludwig of Applied Intuition said at Dirty Jobs that models are less of a moat every quarter, and Lukas Pankau of Industrial Next puts models at ten to fifteen percent of an AI-powered production line. If the model is fifteen percent of the value and gets easier to build every quarter, the premium for owning one should fall over time, and the premium for the other eighty-five percent should hold.

Humanoids and form factors

3. The humanoid timeline gap

The teams building humanoids are among the best in the field, and the hardware has improved faster than I expected. My disagreement is with the timeline the rounds assume. The money is priced for general-purpose autonomy in unstructured environments within a few years. The benchmarks say something else. LIBERO, the standard manipulation benchmark, is considered solved at 95 to 99 percent. LIBERO-Plus perturbs seven properties of the same scene, a camera position or a lighting change, and performance on the same models falls to as low as 30 percent. That is a nudge inside a simulator, before sim-to-real. RoboDojo's real-world set scored the best of ten leading policies at 12.8 percent, and on one hardware platform nine of ten models scored zero.

Deployment tells the same story from the other side. A humanoid can log a terabyte of data a day and still not run a shift. Fan-Yun Sun of Moonlake collected a number from vendors at AUTONOMOUS: deploying one robot into one environment takes a median of about a year in the United States.

My prediction, dated: by 2031, the robots doing paid work will be dominated by wheeled, fixed and purpose-built machines, and general-purpose humanoids will be a small share of units and a smaller share of hours. Some of the companies priced on the opposite assumption will re-price. Some will become excellent hardware and component businesses.

4. Form factor follows the task

Generalize the software, specialize the hardware where the task allows. James Kuffner of Symbotic, whose 20,000 robots move eight million boxes a day, made the point gently: a wheeled robot moving a sixty-pound box on a flat floor beats the anthropomorphic form on speed, cost and efficiency. Adarsh Kulkarni of Foundry made it as the joke that got the biggest laugh at AUTONOMOUS, that getting a humanoid to do your dishes means dishwashers suck. Gary Chen and Conley Oster of Raise Robotics made it from a shipyard: their applications are ones where human physiology is the limiting factor, so a machine shaped like a human inherits the same reach, payload and fatigue constraints that created the job. Cranes exist because people cannot lift steel beams.

The form factors I find myself backing are designed around a workflow: mobile manipulators for overhead work, in-van robots for parcels, retrofit autonomy on excavators, machine-tending robots in cleanrooms, dexterous hands where the task pays for dexterity. If the intelligence transfers across bodies, the body becomes a procurement decision.

5. Where humanoids make sense, and the best argument against my own position

Hazardous environments now: nuclear cleanup, confined spaces, anywhere a person should not be. Human-designed spaces later, once reliability catches up: hospitals, hotels, homes built around stairs and door handles. The home is a hard market. Sunday's 99.1 percent laundry folding across 31 unseen homes is a real result, and it took the authors of Diffusion Policy and ALOHA, a paid army of glove wearers and a narrow task to get there.

The strongest argument against me comes from Dyna-2. Dyna pre-trained on a million hours of first-person human video and found that dexterous hands, the harder control problem on paper, were the cheaper embodiment to post-train: ten minutes of robot data for a bottle cap, against hours for some gripper tasks. If pre-training corpora are human hands, then human-shaped hardware is the closest thing to what the model already knows. Deepak Pathak made the opposite point at AUTONOMOUS with a $4,000 arm, no hand, no force sensor, cooking omelettes from a single wrist camera, and called the wait for better hands and touch an excuse. I hold both results at once. The data economics may end up favoring humanoid hands. That still leaves the reliability gap, the deployment year and the price of the machine.

Data

6. The data wars ended in portfolio management

For two years the argument was which source wins: teleoperation, egocentric human video, simulation or deployment. At Actuate the practitioners refused the premise. Naveen Kuppuswamy of Flexion compressed it to an adage, all data is good data and some data is better data, and reframed the field as a buffet: ego video is the appetizer table, cheap and plentiful, and the mistake is filling your plate with only that. Deepak Pathak's version is that every source fails on at least one of scalability, diversity and closeness to the robot, so Skild pre-trains on video and simulation, post-trains on teleoperation, and lets deployment data compound. Nobody serious defends a single-source strategy anymore. The argument is about sequencing sources against a budget.

7. Collection is commoditizing, and the value is moving up the stack

The hours number is falling from both ends. Rhoda post-trains a new factory task in ten to twenty hours of robot data on top of video pre-training, in four or five days. Applied Intuition trains driving policies from zero human demonstrations in simulation. Ludwig's summary was careful: you still need real data to ship, and the amount keeps shrinking. On the supply side, egocentric video was quoted at three to forty dollars an hour at Actuate, with vendors offering tens of thousands of hours in a single day, and Jason Ma of Dyna called diverse human video a commodity he sees no reason to produce in house.

That is the environment a data company sells into. Raw hours are a race to the bottom. What holds value is everything that makes an hour usable: ingestion pipelines that keep training-launch time flat as data grows tenfold, episode-level quality control, label repair, evaluation that proves a dataset improves a policy, rights clearance, hand-pose extraction, and conversion of data across robot bodies. Foxglove estimates a third to half of engineering time in robotics goes to undifferentiated infrastructure. In the data huddle I ran at Dirty Jobs, a founder with robots in about 150 factories said 95 percent of what they collect is useless, and the pipe is the constraint. Someone is going to manage robot data the way asset managers manage portfolios: what to keep, how to weight it, when to retire it, and when to convert it. That job outlasts the current scramble to collect.

This is where my view on data companies lands. They have a place. The technical moat in collection is thin. The operating moat is thick: access to factories nobody else can enter, rights to footage nobody else can license, a cost structure a competitor cannot match, a fleet that generates data as a byproduct of paid work. Execution beats the algorithm.

8. How I evaluate a data company

  1. Who pays today, and for what? Haomiao Huang of Matter opened a panel with the South Park Underpants Gnomes: collect data, question mark, profit. The companies that fill in the question mark have a customer with a problem other than the data. Lumafield sells inspection and ends up with the best dataset on manufactured objects anywhere. A glove gets worn because it solves something on the floor. A lab buying hours is a customer too, and a concentrated one.
  2. What is the access or cost moat? Thousands of factories across dozens of industries, a facility in a low-cost region where arms run around the clock, cameras on private property with commercial rights attached. If the answer is a data-collection app, there is no moat.
  3. What happens when the biggest customer builds it in house? The buyer universe is a handful of well-funded labs. The answer I accept is operational: running collection across thirty countries or a hundred operators is a business a lab would rather not run, and a quality or evaluation layer is something a lab would rather buy than staff.
  4. Does cost per validated hour fall with volume? If each capture stays bespoke, this is a services business with services margins. If the pipeline automates, it is a platform. Ask what share of the pipeline a model handles today versus at launch.
  5. Who owns what? Factory footage holds trade secrets and worker images. The pipeline is a security surface. Customers who hand over data are starting to ask whether that means they own the model, and nobody has a shared definition yet. A company with a clear answer here will win contracts a better-funded rival cannot.

The upgrade path I look for runs from collection to curation, evaluation and conversion. A company that starts by selling hours and ends by selling the certainty that those hours improve a policy has moved from commodity to infrastructure.

9. Deployment data is the one source competitors cannot buy

Every panel this year converged on the same loop: sales, deployment, data, a better model, more sales. Kulkarni's framing at AUTONOMOUS is the one I like. Before physical AGI, the most valuable asset is a diverse, high-quality dataset from real production, and after the models get good, the weights are worthless without a facility built to run them. The companies I like best get paid to collect. Their robots are on a line doing work a customer needed done, and the data is a byproduct. Collect-then-sell has to find the buyer twice.

10. Contact is the hardest data, and force is where reliability lives

Cameras see nothing when a person seats a connector, runs a hand along a cable, or finds a part in a bin. When I asked a room of fifty physical AI founders at Dirty Jobs which data is hardest to collect, the answer was contact: forces and pressures that ricochet off a part while the video shows nothing happening, aligned to the robot frame. Pulkit Agrawal of Eka made the case at Actuate with a raspberry. Vision proposes the plan, forces do the control. You can watch tennis forever and never develop a serve, because the data never contains what the ball feels like on the racket. Ashwin Balakrishna of Physical Intelligence conceded from the other side that closing the last mile of dexterity still needs well-designed robot or robot-like data, and Kuppuswamy's Sufi parable named the risk: ego video is the streetlamp we search under because the light is there.

So I back the capture layer for touch: sensorized gloves that workers wear while doing their normal job, instrumented hands that share their kinematics with the glove, and force-aware control that treats stiffness and damping as commands. I spent a year looking for the right haptics team before I found one. The evidence that touch closes the reliability gap is early and comes mostly from single companies, and I hold this position with that caveat attached.

11. Cross-embodiment matters, and it is not yet proven

Aurora Feng of Neural Motion asked the foundation labs a question I had not thought to ask: what do you do with the data collected on your last generation of hardware? Mostly nothing. The gap between hardware revisions is wide enough that demonstrations recorded on the old arm will not train a policy on the new one, so teams write off months of teleoperation and collect again. That puts the problem inside single companies, which widens who needs it. Her framing of the ceiling is great: whatever tasks a robot recorded is the neighborhood its policy works in, and that was decided at collection time, before training started.

Dyna-2 supplies the mechanism that makes transfer plausible. The knowledge that crossed the embodiment gap in their experiments was how objects and scenes behave under contact, learned from video prediction, while action labels described one body. Skild's omni-bodied brain and Open X-Embodiment's positive transfer across 22 robots point the same way. What has not been measured is the bar that matters commercially: whether a policy trained on converted data performs like one trained on data collected on that body. Rendering an episode that looks right and producing one that trains well are different bars. Until that number exists, cross-embodiment is a conviction with evidence pending.

Models

12. Model capabilities converge; the value lands in the product

The chief scientist of a model company said it about his own company. Jason Ma expects models to get larger, data to get more plentiful, and capabilities across labs to end up more or less similar, with value captured in the product or service delivered to the end customer. Ludwig says the average model is rising fast and getting cheaper, and reaching the top requires your own data generation and a research team. Pankau puts models at ten to fifteen percent of a production line. Chao Cao's evidence is the robotaxi race: Tesla gathered orders of magnitude more driving data, and Waymo runs the service.

I take this to mean the model layer is where I buy tools, and the moats I pay for sit around it: the hardware platform, the data and evaluation layer, the inference layer that serves whichever policy wins, the observability layer that says what the fleet did today, and the deployed application that owns the customer.

13. The architecture argument, and why I do not join a camp

VLAs uptrain a vision-language model onto actions. World-action models start from video generation and predict the scene and the actions together. Sancho goes straight to 3D geometry and reasons at test time, on the argument that pixels are an expensive representation for contact and that a robot cannot collect its way to every situation it will meet. Eka adds force as a first-class input. At Actuate the panel on this converged: the divide is mostly about which base model you start from and how you spend a fixed data and compute budget, and the winners will craft the mixture for their workload. I invest around the fight. Whichever recipe wins, someone has to serve it, adapt it to the arm, prove the data improved it, and watch it in production.

14. Simulation and reinforcement learning

Simulation has become cheap enough that generating data is often easier than collecting it, and reinforcement learning is the technique with a track record of producing superhuman systems, which needs failure at scale and gets it only in sim. Ludwig expects RL to do much of the work of solving construction sites over the next few years. Eka's raspberry pick came entirely from sim. The counterweight is Kulkarni's wet ping pong paddle: simulating wet rubber or oil on a ball costs absurd compute, and hiring people to collect that data is cheaper. Nobody on the manufacturing panel picked simulation as the thing physical AI scales on. My position is that simulation wins the free-space and driving problems first and the contact problems last, and that the right fidelity is whatever answers the question you are asking.

Deployment

15. Performance before generality

Below a threshold of reliability and speed, nobody pays, however general the robot. Industrial robots sit at 99.99 percent with zero generality; learned robots are general and slow. Eka's framing at Actuate was that scaling video and teleoperation moves the generality axis without touching performance. A RaaS founder at Dirty Jobs said teams celebrate 70, 80 or 90 percent, and if people are going to depend on the system it needs five nines. Going from 95 to 99 percent costs something like ten times more, and at 99.9 someone still certifies the system under a different ISO standard, so the certification bill arrives after the training bill.

The live investment question for me is superhuman-narrow versus general-modest. I lean toward the companies that pick a threshold a customer will pay for, hit it, then widen.

16. Deployment is the eval

Dyna's two models both pass at close to 100 percent in the office. At customer sites neither had seen, graded by the customer's own operators, one passed 46 percent of the time and the other 87. The only way to see that gap is to measure where the work happens. The frontier of difficulty is an M4 that becomes an M3, a torque rating that shifts ten percent, a bird landing on an RTK antenna, a motor coupler that only fails in October. Jagdeep Singh of Rhoda drew the line between a clip that proves a minute and a floor that demands months through variety and mess.

Scale is operational before it is technical. Symbotic's million autonomous miles a day rest on teleoperation fallbacks and service logistics as much as on models. Dyna's deployment layer answers one question around the clock, what did the robot do today, and that layer is what took them from a pilot to a customer ordering more robots at more sites.

17. The pilot is dying as a unit of commerce

An LOI is not revenue. You cannot factor it, raise on it, or buy hardware against it. Nate Williams of UNION ran a room on this at Dirty Jobs, and the buyers agreed with the sellers: a corporate venture arm said proofs of concept convert five to ten percent of the time and prefers a full commercial agreement with a conditional test phase, a RaaS CEO refuses to run pilots at all, and a manipulation founder wants ten million dollars on the table before anyone calls it a pilot. Whatever you agree to, sourcing has a 400-question form waiting, and the Series A bar moved from one million in ARR three years ago to ten million that sometimes fails to clear it now. The founders getting through hire industry people early, agree on widgets per day instead of picks per minute, and treat sales as a team sport.

18. Infrastructure and inference

Every company runs the same loop of collect, triage, mine, train, evaluate and deploy, and most take weeks per iteration. A terabyte per robot per day, bad networking at every site, and value concentrated in the interventions and errors. Factory Wi-Fi was never built to be a utility and now has to be one. On inference, my view is that once a fleet is large enough, its robots should share a nearby GPU cluster over Wi-Fi or 5G instead of each carrying an onboard computer, and that hosting customers' fine-tuned private policies is where the margin will live, the way it does in language models. Safety and uptime will keep a lot of control on the robot, so the answer is a split, and the split is a product.

Demand and geography

19. Labor is the demand signal, unprompted, in every vertical

Nobody on any stage this year had to be asked about labor. Average tenure for a person lifting fifty-pound boxes in frozen storage is eighteen days. Half of US construction workers retire within a decade. Field agriculture runs on guest workers, and fields go unharvested. The skilled-trades version is sharper: every customer Raise Robotics visits has a best person who is sixty, and the knowledge of how the work gets done walks out with them. A shipbuilder posted for a welder and got two applicants in two weeks, reposted the job as robot welding technician, and got thousands. The market keeps answering the replacement question before anyone asks it.

20. Reshoring is the missing middle, and China sets the pace

The US is good at the high end, jet engines and chips, and makes food and packaged goods at scale because shipping them across an ocean makes no sense. Eduardo Torrealba of Lumafield calls the gap between them the missing middle: products that cost five to five hundred dollars to make, which come home with automation or stay where they are. The know-how for mass-producing complex electromechanical systems has largely left the country, which is why, in Kulkarni's example, a hundred billion dollars cannot buy ten million drones on demand. Meanwhile one of China's top robot actuator makers builds actuators as a side business to motors for automatic mahjong tables, and volume in the middle is what gets you practice with process. Pari Singh of Flow put 2025 and 2026 as the tipping point where Chinese robotics, drones and EV autonomy lead. I read that as pressure that forces the West to do what it does well, and as a reason to expect the best robot hardware to keep arriving from China for a while.

Where OPTIM plays

21. Three layers and one flywheel

I invest in three layers: hardware platforms, data and training infrastructure, and deployed autonomy. Underneath all three is one flywheel: sales, then deployment, then data, then simulation, then scale. I pass on humanoid hardware at late-stage marks, on data collection without a customer, and on science projects without commercial intent.

22. The studio

I build robots in my own studio: arms assembled from parts, policies trained on a single GPU, rollouts that fail. I wrote about it here. It changed how I read a pitch. I now ask what the office success rate is and what the customer-site number is, how many hours of data the last task took, and what broke last week.

What would change my mind

Positions are only useful if they can be wrong. These are the markers I am watching.

  • On the cycle: if 2027 funding stays above the 2026 pace and deployments grow in proportion, revenue receipts rather than demos, then the heat was warranted and this page is too cautious.
  • On humanoids: a general-purpose humanoid fleet running paid shifts at a customer for six months, at task success above 99 percent without teleoperation fallback, at a cost per hour below the wheeled or fixed alternative.
  • On benchmarks: a general policy that holds above 90 percent under LIBERO-Plus-style perturbations and above 50 percent on RoboDojo's real-world set.
  • On data companies: a frontier lab bringing collection fully in house at scale and a leading vendor's revenue falling as a result, or per-hour pricing stabilizing instead of falling.
  • On touch: vision-only policies reaching production reliability on contact-rich tasks, insertion, cable routing, assembly, at industrial cycle times.
  • On cross-embodiment: a public result where a policy trained on converted data matches one trained on natively collected data within a few points, or the opposite, a foundation policy absorbing the embodiment gap through scale alone.
  • On models: a lab holding a durable capability lead that shows up as pricing power for two years. That would mean the model layer is a moat after all.

Not zero-sum

Peter Ludwig refused the zero-sum framing on stage: the markets are large enough that every company in the room can succeed together. Navin Chaddha's arithmetic says the physical world is a thirty trillion dollar economy, and ten percent of it moving to AI is five times the revenue of the entire software industry. A room of fifty founders handing each other a microphone ran on the same assumption. So does this newsletter, and so does this page.

~Bogdan

References