Out of the Demo, Into the Field: Notes from Dirty Jobs 2026

Out of the Demo, Into the Field: Notes from Dirty Jobs 2026

September 23, 2026 - San Francisco

Jay Kapoor opened the third Dirty Jobs Summit with a list of jobs: the 3 a.m. shift in the factory, the road crew pouring an underpass in the heat, the people working thousands of feet down in a mine, the harvest that has to come in the week it ripens. His point was that the people doing those jobs have been showing up in smaller numbers every year, and that the last software wave mostly forgot them. VSC Ventures has been investing in the category for five years, since before it had a name. Jensen Huang eventually supplied one.

The format is the thing to know about this event. Three hundred seats, one attendee per company and one per fund, four keynotes, and then "huddles": three rooms running at once, fifty people to a room, one moderator, a microphone that travels, and anyone at the table can talk. The agenda page says it plainly. They do not do panels. I moderated the data supply chain huddle with the Bright Data team at two o'clock, then sat in Nate Williams' room on getting from pilot to purchase order at three. The theme Jay set for the day, out of the lab, out of the demo, into deployment, held in every room I was in.

For friends who missed it, here is what happened. The keynotes were public and on the record, so those speakers get their names. The huddles were peer conversations with the microphone passing around the room, so I have kept those contributions unattributed and described people by what they build.

Peter Ludwig: the road stays the road

Peter Ludwig, co-founder and CTO of Applied Intuition, told the room he had taken a Tesla Cybercab from Sunnyvale to the venue that morning. It is a camera-only system, and it works on roads. Nobody gets away with cameras alone on a dusty construction site. That contrast framed his whole talk. The road stays the road, so a map survives. On a job site every bucket of earth changes the terrain, which changes how you map, how often you update your priors, and what the system has to do on its own to stay safe.

Applied has been at this since 2017, and Ludwig was candid that simulation, where the company started, is now a small fraction of what it does. Ten years of investment there, and he still said there is no substitute for real data. The exception he is excited about is reinforcement learning. The research is decades old and the systems were never fast enough to run the epochs RL needs. In the last twelve to eighteen months his team built simulators fast enough to change that, and he expects RL to do much of the work, over the next few years, of getting these environments solved, by which he means a machine can do anything a person can do in them. Applied's TerraZero work, published in July, trains driving policies from zero human demonstrations and topped a long-tail planning benchmark doing it. His summary of where that leaves data was careful. You still need real-world data to ship a production system. The amount you need is shrinking.

Models, in his account, are getting easier to build and are less of a moat every quarter. The average is rising fast and getting cheaper, and he pointed to the recent robot games in China, where fifty or a hundred companies showed capabilities that were frontier work not long ago. Getting from average to the top takes your own ability to generate data, because pure licensing hits a ceiling, and a research team, because anyone doing group projects on public models will struggle to keep up.

On the gap between a demo and a system running close to 24/7, he offered the boring examples that matter. A component you assumed was reliable fails every thirty hours, and now that is your mean time to failure. A vendor subsystem changes how the site's operators do their jobs, and now those people need retraining. He has teams in Australia running and supervising deployments, and said there is no way to learn those things except by being there.

On trust, he said operational safety is a core value of the company, that trust with a customer is earned through consistency and lost with a single mistake, and that the hard judgment is when to take the human out of the loop, neither too early nor too late. Applied has a clean safety record, and he noted, without naming anyone, that competitors' mistakes have opened doors. He also refused the zero-sum framing outright: the markets are large enough that every company in the room can succeed together. Applied works with companies that build machines and companies that operate them, and sells what he called the undifferentiated heavy lifting, software updates, health checks, diagnostics and data uplink.

Jagdeep Singh: what the floor proves that the conference room cannot

Jay introduced Jagdeep Singh, co-founder and CEO of Rhoda, as the person in Silicon Valley who has done the most deployments of frontier technology into skeptical buyers, and steered the conversation toward selling. Singh's opening distinction was between making something work on video and making it work on a production floor. A model that performs on the dataset sitting on your laptop meets variety and mess on the floor, and it has to keep performing through that variety for months. A clip proves a minute.

His example was box decanting on an automotive line. Boxes arrive with a strap. The robot lifts by the strap, tips the contents into a container, and sorts the plastic bags, paper and cardboard that come out. Every box arrives with the strap in a different place at a different angle. You cannot motion-plan it, and you cannot teleoperate your way to coverage either, because without the full diversity of strap configurations the policy guesses. The long tail lives inside single tasks, on a day-to-day basis.

The selling advice was practical. Being deployable means having answers before anyone discusses autonomy: what is your service and support strategy, your integration strategy, your supplier allocation, and can you pass the review a European carmaker's security department will run. There is a real trade-off between running as many proofs of concept as you can and running only the ones set up to convert, so Rhoda formalizes the business conversation before the POC starts. Where would this go, at what scale, through what process. It filters out some opportunities, which is the point.

Every POC is metrics-based. First a lab demonstration of the new task, then an agreed KPI: hit this cycle time at this level of autonomy and the POC passes. The executives who decide only read the summary report, so you never want one that says "did not pass" about something that was working. Jay added the buyer's-side complication: the customer may benchmark you against a human rate for a job humans no longer show up to do. Singh's answer was to stop arguing about picks per minute and agree on widgets per day.

The surprise on site was connectivity. Rhoda plans in the cloud and executes on the robot, so the shop floor needs a link out. Factory Wi-Fi is spotty and has never needed to be anything else, and plants are only now realizing that this generation of AI needs good connectivity as a basic utility. He expects on-device models to improve through quantization and distillation while the ceiling on what more compute can do keeps rising, so the harder the task, the more the cloud stays in the picture. Deployment also needs a sales team with relationships, application engineers who can define and stage tasks, and field support sized for robots in service most of the day, because a hundred people in a customer can say no while only one can say yes, and the people who say no include a procurement group he described as negotiating machines.

The story he told about emergent behavior explains his optimism anyway. A robot was moving ball bearings into containers on a customer floor when a worker tossed a stray object into the tray. The model reached out, picked it up, paused as if unsure what it was, and dropped it in the trash. Nothing in the training data covered it. Because Rhoda's models are pretrained on general video and carry a strong prior on how the world works, post-training for a new task now takes roughly ten to twenty hours of robot data and four or five days, and deployments happen within days rather than six months. Jay's closing line was that while the industry talks about hundreds of thousands of hours, Rhoda keeps saying under twenty, maybe under ten. Singh's calendar matched: 2025 was the year of building the model, 2026 is the year of POCs and deployments, and the research frontier is physical reasoning and in-context learning.

The data supply chain huddle

I ran this one with Shweta, who runs global solutions at Bright Data, and Bill, who has spent four years there talking to what he estimated as a thousand AI companies about their data pipelines. Bright Data has sold web data for twelve years, and their pitch from the main stage earlier came down to three things: the teams solving real physical problems need petabytes and narrowly specific data at the same time, collecting from the internet is getting harder as blocking spreads, and your expensive model people should spend their time on models rather than on plumbing.

I will be honest about the room. It took a while to warm up. Fifty people, a traveling microphone, and a first question that landed with silence. I had eight questions ready and expected to get through four, and that is about how it went.

The first question was a poll. If you could keep one kind of data, which would it be: teleoperation on your own robot, wearables and UMI-style devices, human video, simulation, or what your deployed fleet sends back. Almost no hands went up. The one enthusiastic hand came from a dexterity researcher who wanted tactile data, and when I asked what kind, gloves or instrumented hands, his answer was that everything we have today is a subset of what fingertips sense, so do not define it yet, ramp it up, and aim for human-level. A power infrastructure founder selling into data centers nominated a type nobody had listed: power flow sampled at higher frequency than any utility measures.

On whether the pipeline matters more than the data type, the hands went up. One founder converting data collected on a generic Chinese robot to his own machines said the trick is mapping wrist poses rather than joint angles, and that simulation has become cheap enough that generating data is often easier than collecting it. A founder who builds a curated library of skills for industrial robots, and who considers the model a byproduct of shipping useful skills, made the longest argument of the session. A pipeline that works for a small manufacturer should also work inside a European carmaker with more regulations than he can remember. Once the volume arrives, most of it adds nothing or confuses the model, so quality control is the work, and on a 21 billion parameter model every wasted episode is GPU time. A failure captured at the right moment and connected to the right context is, in his words, data gold. He added two things I have not heard said out loud at a conference: a data pipeline is also a security surface, and customers who hand over data are already asking whether that means they own the model, without any shared definition of what owning a model means.

Then I asked who in the room had ever bought data. One hand went up. A team building a sensor foundation model buys small specialized sets for use cases like oil rigs, wind turbines and HVAC, and calls it negligible next to the open-source corpus they pretrain on. Bill was visibly surprised: a company that sells hundreds of millions of dollars of web data a year, and one buyer in a room of physical AI founders. His own observation was that requests have shifted from "give me all of YouTube" to "give me egocentric video of one person sewing under specific lighting," which is exactly the data that is hard to find.

Two people pushed back from opposite directions. A founder in critical power does not want general data at all. He is hunting for black swan events on a customer's utility feed, which happen rarely, so usefulness beats volume and he sometimes simulates events to show a customer where the risk sits. It reminded him of the support-vector-machine era, when the work was picking the input that gave the most incremental learning. An investor whose fund backs deployed-data companies and avoids humanoids said his companies' flywheel is one deployment after another, and that whether that pace is good for venture returns is a question he is choosing to be patient about.

The loudest correction came from a supply chain and mobility investor based in St. Louis, who said everybody in the room is buying data whether they call it that or not. Road data, intersection data, air traffic control, weather, regulatory, address data, everything coming off a manufacturing machine, layered on geospatial. His rule was to buy what you cannot create. Nobody in the room owns a satellite or a fleet of agricultural drones. On addresses, he noted that every major shipper photographs every delivery and that 98 percent accuracy is considered unacceptable. Bill's example fit the same shape: a logistics customer predicting container arrival by stitching together ship position, weather and customs broker status from public sources. He was equally clear that power grid telemetry is not on the internet.

A founder of a spatial world model company asked the room what data is hardest to collect. The answer was contact. Forces and pressures that ricochet off a part while the camera shows nothing happening, aligned properly to the robot frame. The concrete example was a CNC cell where, every fifteen or so cycles, a person walks over and wipes shrapnel off the tool. Automate that one motion and you have lights-out production, and nobody has data on it. The same goes for what a machinist carries in his head: a CAD file gives you dimensions and thread and nothing about how fast the tool should move or when it needs maintenance. Going from 95 to 99 percent accuracy costs something like ten times more, and at 99.9 percent someone still has to certify the system under a different ISO standard than the one that covers motion, so the certification bill arrives after the training bill.

An investor who brokers startups into Fortune 500 innovation programs added the buyer's paradox. Corporates want the POC and refuse to share the data, because operational data is the last proprietary thing they hold.

The last exchange was about plumbing. A founder whose robots generate about a terabyte per robot per day said the main struggle is getting it off the machine fast enough to act on it, because a thermal problem in a compressor matters now, and every deployment's networking is bad regardless of 5G or anything else. A founder with robots in about 150 US factories answered: 95 percent of what they collect is useless, so they store everything and spend their compute and human annotation on the interventions and errors that carry outsized value, and redundant sensors that will not ship in the end product, a safety scanner, an air pressure sensor, tell them which moments deserve attention. The first founder agreed, said data compression now eats more of his team's time than he ever expected, and that Starlink makes more sense to him every month. Stale data loses most of its value, so the pipe is the product.

Nate Williams' room: from pilot to purchase order

Nate Williams of UNION, a three-time CRO before he was a fund manager, ran the three o'clock huddle on go-to-market. Two of his own rules set the tone. An LOI, MOU or IOI is not revenue, because you cannot factor it, raise on it or buy hardware against it, so skip from purchase order straight to master services agreement. His second rule is that sales is a team sport now. The old archetype where engineers mattered and sales was the bottom rung does not survive a market where Series A and B investors want revenue receipts.

The CEO of a robots-as-a-service company, a former investor himself, was blunt about corporate innovation teams. Many are effectively paid to waste time, and their KPI is having something cool to show when the CEO visits. His second point was uptime. Teams celebrate 70, 80 or 90 percent, and if people are going to depend on the system it needs 99.999. His company does not run pilots at all. They do not have the time, and they would rather take on performance risk in a real deployment than spend it proving things.

The corporate side answered. A venture and partnerships lead at a German automaker's North American arm said POCs and conditional purchase orders shift the cost of proving onto the vendor, which is why corporates like them, and that the questions have changed in two years, from admiring robots to "what does this do to my bottom line." A twenty-year veteran of a large electrical equipment maker's venture arm said 70 to 80 percent of their portfolio companies sign a partnership with a major international company, and credited a specific function for it: a guide, usually a former operator paid by the fund, who knows the timing, budget and decision process on the inside. His policy is zero POCs. They consume salary and time and convert five to ten percent of the time at best. He prefers a full commercial agreement up front with a light test phase subject to agreed conditions.

A founder eight years into a warehouse manipulation company, three of them in production, said the secret has been co-development with customers willing to put real money down before anything is deployed. He wants at least three customers with ten million dollars or more on the table to get off the ground, and said the company will do about 75 million in revenue this year while deploying ten sites and piloting five more. "Call it a pilot, call it whatever. I call it ten million bucks in the bank."

The unexpected star of the session was sourcing. Nate told the story of a portfolio CEO with a contract at a large telecom who expected to close in a month and had never heard of the sourcing department. Nate had negotiated two MSAs through sourcing as an operator, and each took six months. A former procurement executive who spent a third of his career at an aerospace company, then a utility, then an insurer, walked the room through what happens next. The salesperson believes the deal is locked because the engineers and executives love it. Then it reaches sourcing, where nobody has heard of it. You become a vendor of record in SAP or Oracle. You receive a 400-question form on data security, physical security and whatever your category demands, plus reps, certs and insurance, and if you cannot produce them you cannot win. A month goes by and everyone yells at procurement. His prescription was to hire industry people early, ideally a former procurement person who knows where the bodies are buried. For anyone selling to government, if you have not hired a lobbyist you have already lost, because you do not win RFPs, you crash them.

Nate's hiring advice was specific. The first sales hire is a full-stack seller who gives the founder scale: books the meeting, signs the NDA, follows up, keeps the engineering founder in engineering. A CRO is a different job, someone who demands product, feeds customer requirements back and keeps credibility with engineering. On the funding bar, one million in ARR was the Series A hurdle three years ago and ten million sometimes fails to clear it now, so seed to A has become two things in parallel: retiring technical risk and showing receipts on go-to-market experiments across channel, pricing, packaging and support. Founder-led sales, he said, is not founder-only sales.

The last stretch was about inertia. A founder asked how to sell into an incumbent sector when your value proposition is changing their workflow. Nate's answer came from selling software to utilities in 2010, when most decision makers wanted steak and golf: create a pincer by going to regulators like FERC and NIST and making the change important to the customer's next year, so they choose it without being asked. The procurement veteran added that utility math runs backwards, since utilities earn on capital assets rather than on electrons, so the winning path is often the tier two and three suppliers who can carry you into an offering. A construction robotics founder said his industry learns from fathers and grandfathers and does not change because mistakes kill people or businesses, so his pitch of business transformation got laughs. What worked was starting with a robot that is faster, cheaper and more accurate than a person, and letting the customer notice later that their business had changed. The session closed on incentives: know the approval levels, structure the contract so it does not need escalation, and lead with outcomes, because a CFO does not care about conversion rates.

Vijay Chattha, VSC's founder and Jay's partner, closed the day in conversation with Navin Chaddha, Managing Partner at Mayfield, seventeen times on the Midas List, at a firm that has backed more than 500 companies since 1969. Vijay's first question was the right one. Two or three years ago hardware and robotics were dirty words and founders in this room were flying around the planet to close rounds. Now physical AI is one of the hottest categories. What changed, and will it last.

Chaddha's answer was historical. Silicon Valley was the hardware valley, then spent twenty years as the software valley while software ate the world. Today hardware has eaten software: by his count eighteen or nineteen of the twenty largest companies by market cap are hardware and systems businesses, and he called this the golden era of systems. His market math was that the physical world is a thirty trillion dollar economy, and if ten percent of it moves to AI over the next three to five years that is a three trillion dollar opportunity, five times the roughly 600 billion in revenue the entire software industry makes today. Data is the constraint, because today's transformers were trained on an open internet and the physical world has no equivalent, so he sees room for data companies, model companies, middleware and applications, with reinforcement learning growing in importance as fleets run.

On founders, his list is the one he has used for years, and he said it matters more now. Mission and values come first, then team players whose language uses "we" more than "I," then a learning mindset, which he called ten to twenty times more important in the AI era because nobody, himself included, knows what will be announced in the next ninety days. On process, he described a blink moment in the first minutes of a meeting and warned his own team against the "cake soup problem," where the entrepreneur brings raw ingredients and the investor starts cooking a thesis of his own. Listen, evaluate whether they found the market, then analyze. He called venture a humbling business where you are wrong more often than right, and told founders to evaluate their investors the same way, on whether they invest in people or in vanity markets.

What I missed, and what was on the lawn

Lindon Gao's keynote on model company versus deployment company happened while I was in the hallway doing the thing this summit is designed for, so I have nothing to add beyond the pointer to our recent piece on Dyna-2 and its human-video scaling laws. The startup showcase ran in the courtyard all afternoon and, as Jay noted from stage, it had more robots than data vendors, which is not the norm this year. Formic's systems were out front, Autopilot had its autonomous machines running, and Presso, a VSC portfolio company, brought its dry-cleaning robot back for anyone who spilled wine on a suit during happy hour.

A few takeaways:

  • Ninety-five percent of fleet data is useless and the pipe is the constraint. A terabyte per robot per day, bad networking at every site, and the value concentrated in interventions and errors. The teams furthest along instrument the interesting moments with redundant sensors and spend their compute there.
  • The hours number keeps falling, from both ends of the stack. Rhoda post-trains a task in ten to twenty hours of robot data on top of video pretraining. Applied trains driving policies from zero demonstrations in simulation. Two serious operators reached the same conclusion from opposite directions.
  • The pilot is dying as a unit of commerce. A RaaS company refuses to run them. A corporate venture arm says they convert five to ten percent and prefers a full agreement with a conditional test phase. A manipulation founder wants ten million on the table before anyone calls it a pilot. Whatever you agree to, sourcing has a 400-question form waiting.
  • Trust math is asymmetric. Earned through consistency, lost in one mistake, per Ludwig. Certified after you hit 99.9 percent, at a cost nobody budgeted. The customer's stated human throughput, checked against who shows up.
  • The room did not treat this as zero-sum. Ludwig said every company present can succeed together, Chaddha's arithmetic says the market can hold them, and the huddle format itself, fifty people handing each other a microphone, ran on that assumption. It is the assumption this newsletter runs on too.

Thank you to Jay Kapoor, Vijay Chattha, Maggie Philbin and the VSC Ventures team for a third edition that kept the room senior and the conversations real; to Shweta Suri and Bill at Bright Data for co-hosting the data huddle and for the look on their faces when one hand went up; to Nate Williams for a session that felt like a sales offsite; and to the fifty people in my room who eventually started talking. If you were there and we missed each other, or you are building along any of these lines, reply to this email. That is how the best conversations from the day started.

References

Company figures, market statistics and customer anecdotes cited above were shared by speakers on stage and in the huddle rooms at Dirty Jobs 2026 and reflect their own reporting. Huddle contributions are paraphrased and are unattributed by design.