Physical AI Leaves the Lab
Foundation models have given robots general perception. The next test is not a demo video but a cost per task.
The robot in the video walks, picks up a tote and places it on a shelf. The question that matters is not whether it can do that once. It is whether it can do it ten thousand times without intervention, how much each pick costs including maintenance and supervision, and whether that number is lower than the alternative. Physical AI is entering the phase where those questions get answered.
What changed is the software. For decades industrial robots were programmed task by task, which made them precise but brittle. Foundation models trained on vision and language, and increasingly on robot demonstrations, give machines a general understanding of objects, instructions and scenes. A vision-language-action model maps what the robot sees and what it is told directly to motor commands. Change the task and you change the instruction, not the code.
Why now
- Robot foundation models from labs and chipmakers give general perception and language grounding.
- Humanoid and mobile manipulation platforms have entered paid pilots in logistics and automotive facilities.
- Simulation platforms generate training experience at a scale real-world trials cannot.
- Component costs are falling as actuators, sensors and compute reach higher volumes.
- Labour shortages in physical work are structural in several major economies.
What the pilots are showing
The honest answer is: early data, unevenly disclosed. Some operators report task completion rates and hours of autonomous operation; few publish cost per task. The pattern in what has been shared is that robots are useful in structured, repetitive settings first, and that the supervision ratio, how many robots one person can oversee, is the number that decides economics.
Physical AI readiness by environment. Readiness: Warehouse picking 78, Machine tending 72, Assembly 55, Inspection 70, Hospital logistics 48, Home tasks 20.
Humanoids versus everything else
Humanoid robots capture the imagination because they fit environments built for people. That is a real advantage: no need to redesign the factory. But it comes at a cost in complexity, balance and battery life. The practical view is that humanoids win where flexibility matters more than throughput, while wheeled and fixed platforms win where the task is known and volume is high. Both are accelerating, and the software stack underneath them is converging.
The supervision ratio, how many robots one person can oversee, is the number that decides the economics.
Who benefits, who is at risk
Beneficiaries: robot platform makers, sensor and actuator suppliers, simulation providers, and operators with repetitive, high-volume physical work. At risk: fixed-automation integrators that cannot move up the software stack, and workers in tasks that are repetitive, structured and measurable. Regulatory and labour pushback is a real constraint in shared human environments.
What happens next?
- The first multi-site commercial deployments publish uptime and cost-per-task data.
- Supervision ratios improve as fleet learning matures.
- Safety certification frameworks for robots in shared spaces take shape.
- Consolidation among humanoid startups as capital concentrates on those with deployment data.
Related topics
Sources & references
- 01Robot learning and vision-language-action literature — Academic and industry labs; see Sources pageresearch
- 02Company pilot announcements — Robotics vendors and operatorscompany
More from Artificial Intelligence
The Agentic Enterprise: From Pilots to Production
A Deep Dive on enterprise agent deployment: the adoption curve, the market map, the technology, the risks and the opportunities.
Open-Weight Models Are Becoming Strategic Infrastructure
Open-weight releases shape who can build, where models run and which countries control their AI stack. Governments and enterprises are treating them like infrastructure.
Small Reasoning Models Are the Quiet Revolution
An explainer on test-time compute, distillation and why small reasoning models make private, on-device and regulated AI deployments viable.