Physical AI: When intelligence meets the real world

Published on September 21 2026
In the first article, ‘Why world models are the future of autonomous AI‘, of this mini series on AI beyond Large Language Models, we closed with a question: ‘Given how things stand now, what happens next if this action is taken?‘ World models try to work out the answer. They give a system (human or machine) an internal way to simulate how an environment might evolve under different choices. Physical AI is what happens when that answer is handed not only to a person, but to a machine that can act on it. Where world models introduced AI as an internal […]

In the first article, ‘Why world models are the future of autonomous AI‘, of this mini series on AI beyond Large Language Models, we closed with a question: ‘Given how things stand now, what happens next if this action is taken?‘ World models try to work out the answer. They give a system (human or machine) an internal way to simulate how an environment might evolve under different choices. Physical AI is what happens when that answer is handed not only to a person, but to a machine that can act on it.

Where world models introduced AI as an internal simulator, physical AI brings that simulator out into the open: into warehouses, onto factory floors, onto roads, and into hospitals. It is the moment when artificial intelligence stops being only about text, images, and recommendations, and starts moving atoms as well as tokens.

The image shows an AI generated image depicting AI moving from the digital world into the physical world.
A conceptual illustration of AI moving from the digital world into the physical world. Created in Microsoft Designer.

What is physical AI?

Physical AI refers to artificial intelligence systems that operate in and interact with the physical world, rather than existing only in software or digital environments. More precisely, these are embodied or physically situated systems whose decisions must account for real-world dynamics, sensing uncertainty, and the consequences of action in a continuous loop. 

This distinguishes physical AI from what we might call ‘digital AI’: chatbots that generate text, image models that create pictures, recommendation engines that rank content. Those systems live on screens and servers. Their outputs are not trivial, but they do not directly steer vehicles, lift objects, or adjust valves. Physical AI closes the loop: it senses the world, interprets what it perceives, decides on an action, executes that action through motors or actuators, and then senses again to see what changed.

DIGITAL AI PHYSICAL AI
Where it lives Screens, servers, apps Robots, vehicles, instruments, smart spaces
Inputs Text, images, clicks Sensors: cameras, lidar, tactile, inertial units
Outputs Text, images, rankings, recommendations Motor commands, actuation, movement in the real world
Direct physical effect No Yes – moves objects, steers vehicles, adjusts equipment
Main constraint Accuracy, relevance, bias Safety, timing, uncertainty, physical consequences of errors

Comparison between Digital AI & Physical AI

Several elements are essential:

It also helps to set physical AI apart from older forms of automation. Traditional industrial robots follow fixed trajectories and rules in highly structured settings. They repeat the same motion thousands of times, in the same place, with the same parts. Physical AI systems, by contrast, are built to handle variation: new object positions, changing light, unexpected obstacles, and mixed human–machine workflows. The difference is between a machine that executes a script and one that adapts its behaviour as conditions change. 

Where physical AI already lives

Physical AI is not a single product or a distant promise. It is a pattern appearing across sectors wherever machines must act safely and flexibly in messy, real conditions. It is already deployed, often unobtrusively, in places many of us pass through every day.

➣ Warehouses and logistics

In modern fulfilment centres, AI coordinates fleets of autonomous mobile robots that navigate aisles, avoid workers, and coordinate with conveyor systems using on-board perception and planning. Amazon’s Sequoia warehouse system and its DeepFleet model have improved inventory identification, storage speeds, and robot travel efficiency across fulfilment centres. Drones manage warehouse inventory by flying between shelves and scanning barcodes or QR codes autonomously, turning a task that once required ladders and handheld scanners into a scheduled, software-driven operation.

➣ Factories and industrial systems

Humanoid and collaborative robots are entering real factory floors, in pilots and early deployments, at companies like BMW, Toyota, Schaeffler, and Tesla, performing tasks beyond fixed scripts. Industrial robots equipped with vision–language–action models handle bin picking, quality inspection, and assembly where conditions vary from one shift to the next. “Smart spaces” use fixed cameras and computer vision to optimise flows in factories and warehouses without requiring humanoid robots at all. Physical AI here operates at the level of the environment, not merely the machine.

➣ Transport and infrastructure

Autonomous vehicles run the core physical AI loop: they perceive the road, predict other users’ behaviour, plan a safe path, and execute it via steering, braking, and throttle. Drones inspect infrastructure such as power lines, pipelines, and bridges, combining perception with flight control to adapt to wind, obstacles, and changing viewpoints. In each case, the system must reason about real-world dynamics and act safely under partial observability and uncertainty.

➣ Healthcare and other domains

Surgical robots are beginning to add perception and assistance to what are still largely surgeon-controlled instruments. Cleaning, delivery, and inspection robots in hospitals, offices, and public spaces operate alongside people in dynamic environments, navigating corridors and doors while avoiding collisions. Again, the workflow is the same: perception, interpretation, decision, action, and renewed perception.

Across all these domains, physical AI is not a single gadget but a recurring design. In each case, the machines must understand and act in the world as it is, not as it was specified on a drawing.

Inside the machine: world models as the planning layer

If world models are the internal simulators that let a system ask ‘what happens next if…‘, physical AI is what we get when those simulators are embedded in machines that can act on their own answers. In the stack of a physical AI system, world models occupy a crucial middle layer between perception and action. 

At the lowest level, sensors and perception modules turn raw data into structured representations: objects, depths, semantics, and relationships. Above that sits the world model as an internal representation that predicts how the environment will change under candidate actions. A policy or planner then chooses actions based on those predictions. For example, ‘if I grasp here, the object will slide; if I approach from this angle, collision risk drops‘. Finally, low-level controllers translate chosen actions into motor commands. 

This gives physical AI systems a kind of internal rehearsal. They can mentally simulate action sequences before committing, testing alternatives in an internal model rather than only through trial and error in the real world. For an autonomous vehicle, world-model-like components predict how other road users and the vehicle itself (ego vehicle) will evolve over the next seconds, supporting safer planning. For a robot arm, they anticipate how candidate actions may alter object configurations, contact states, and scene geometry before physical execution.

Some systems use explicit world models that roll out latent states over time; others encode predictive structure more implicitly. The functional role remains the same: anticipating consequences of actions so the system can move from ‘react to what I see now‘ to ‘plan a sequence that leads to the desired outcome‘. 

Not every physical AI system uses an explicit world model, but where they are present, world models form the internal planning layer that predicts ‘what happens next if…’ and guides action.

The image shows an illustration of a physical AI stack
Physical AI stack: from perception to action, with world models as the planning layer

Human work and human judgement

Physical AI does not remove humans from the loop; it changes the shape of their work. In warehouses and factories, workers increasingly supervise fleets of robots, handle exceptions, and redesign workflows rather than performing every repetitive motion themselves. Engineers and technicians spend more time configuring behaviours, setting safety constraints, and interpreting system logs than manually tuning each trajectory. 

New roles emerge: robot coordinators, fleet managers, physical AI trainers, people who blend domain knowledge in logistics, manufacturing, or transport with an understanding of how these systems learn and fail. They decide where autonomy is appropriate, how tightly to constrain it, and how to intervene when conditions fall outside the system’s design envelope.

Physical AI systems propose or execute actions based on their internal models; humans still set goals and acceptable risk levels. When something goes wrong, the inquiry does not stop at ‘was the model right?‘ It must also ask whether the goals, constraints, and oversight were appropriate for this context. That’s where human and organisational judgement has to come in.

Models estimate outcomes; they don’t decide what is fair or worth sacrificing. Physical AI puts that distinction into sharp relief, because the consequences are physical and immediate. It pushes us to think carefully about where we delegate action, how we keep meaningful human control, and what kinds of mistakes we are prepared to tolerate in exchange for gains in safety, efficiency, or access.

Why governance and procurement must pay attention

The technology is moving faster than the rulebook. Physical AI introduces failure modes that matter for safety and accountability, such as compounding prediction errors, gaps between simulation and reality, and inconsistencies over long action horizons. Procurement and governance will need to ask not only ‘what does the model say?‘ but also ‘how was it tested in real conditions?‘ ‘what data and environments was it trained on?‘, and ‘who can override or shut it down, and how?‘

Standards, certification, and liability frameworks will have to catch up with systems that learn and adapt after deployment, not just follow static specifications. For public institutions and international organisations, this means developing ways to evaluate physical AI systems that go beyond feature lists and marketing claims, focusing instead on tested performance in relevant environments and clear lines of human oversight. 

As physical AI spreads, the stakes of getting governance and procurement right rise with it. Human-centred design and clear responsibility are not optional add-ons; they are part of what makes these systems trustworthy in the first place.

Intelligence that acts

Physical AI is best understood not as a replacement for human judgement, but as an extension of our capacity to act safely and effectively in complex environments. It takes the basic human move of imagining what might happen next if we do this and builds it into machines that can carry out those actions. 

Used thoughtfully, physical AI can help us move goods more efficiently, inspect infrastructure more safely, drive with fewer accidents, and perform delicate procedures with greater precision. It can free people from the most dangerous and repetitive tasks while creating new kinds of work that rely on human insight, ethics, and oversight.

The challenge ahead is to design the conditions under which the actions of these machines are acceptable. That is a task for engineers and companies, yes, but also for policymakers, procurers, and the public. Physical AI brings intelligence into the real world. How we shape it will determine what kind of world it helps to create.

Author: Slobodan Kovrlija


cross-circle