● Behind the technical choices regarding training methods lies a question of sovereignty: whoever controls physical data and simulation infrastructure will control the robots of tomorrow.
● Nvidia’s pro-simulation strategy has yet to prove itself, and players like Physical Intelligence are instead banking on the massive collection of real-world data.
Generative AI has been trained to write text and produce images, sounds, and videos. The next wave requires figuring out how to teach machines to act in the real world. Behind the term “Physical AI”—a term coined by Nvidia, which also sells the GPUs used to train it—various approaches are vying for prominence. Industrial robots, autonomous vehicles, smart workspaces: the possibilities are as numerous as the technical challenges. At the heart of the debate, one question divides researchers and engineers: should these systems be trained primarily using real-world action data, or should an internal representation of the world (a “world model”) be built first so that the machine can understand before acting?
Not only is there a lack of physical interaction data, but the data that does exist is difficult to acquire efficiently.
What is physical AI, and how does it differ from automation?
An automated factory arm that repeats the same movements down to the millisecond is not physical AI but programmed automation. A physical AI system, on the other hand, does not follow a script: it perceives, reasons, and adapts to what it encounters. Place an unexpected obstacle in front of a conventional robot: it will either stop or collide with it. Place it in front of a physical AI system: it will go around it, or ask (or even wonder) what it should do. A seminal 2025 article summarizes this logic in six building blocks: a body, sensors, the ability to act, to learn, to decide independently, and to interpret context. But behind this list lies the idea that intelligence isn’t in the code. It emerges from the interaction between the machine’s body, the environment, and the history of its past experiences. Physical AI therefore lies somewhere halfway between neuroscience and robotics.
Physical AI: The Problem of Training Data
In China, physical AI is no longer a futuristic concept: AGIBOT has deployed over 10,000 robots and released an open-source dataset collected under real-world industrial conditions. That’s the crux of the matter: while training a large language model requires a virtually unlimited supply of data, training a robot to grasp objects, detect obstacles, or navigate a hallway also requires data. Yet the article “Generative Artificial Intelligence in Robotic Manipulation” is clear: not only is there a lack of physical interaction data, but the data that does exist is difficult to acquire efficiently. Filming a robot fail at the same action a thousand times is time-consuming, costly, and potentially damaging to the equipment or the environment. Finally, there is the issue of planning long-term tasks—that is, how to sequence dozens of successive actions without losing sight of the objective. A robot that can only see remains one-eyed: it must also hear, touch, and understand what is being said to it in order to make a coherent decision in a changing environment. This is why the multimodal approach must be redefined.
World models vs. action models: two strategies for training robots
Researchers are considering two approaches to training physical AIs: the action model approach, which involves learning by doing, and the world model approach, which assumes that the AI has understood its environment before acting.
| Criterion | Action models | World models |
| Logic | Learning by doing, using real or simulated trajectories.
The robot repeats the action of cracking an egg thousands of times until it masters the exact amount of force required. |
Simulate the consequences before acting.
Before grasping the egg, the robot simulates several grips and predicts which one will prevent it from breaking. |
| Tools | Diffusion, GANs, reinforcement learning, imitation learning.
A chef cooks with sensors: their movements become the robot’s training data. |
High-fidelity simulation, digital twins, synthetic data.
A virtual kitchen generates millions of scenarios: slippery pots, flames that are too high, ingredients moved out of place. |
| Strength | Reliable for known tasks. Quick to deploy.
The robot can make a perfect omelet as long as the pan is always in the same place. |
Generalizes novel scenarios. Reasoning outside the distribution.
Did the recipe change at the last minute? It adapts the steps by reasoning based on what’s already cooking. |
| Limit | Fails outside its training distribution.
If the salt is moved 10 cm, the robot searches, hesitates, and gets the seasoning wrong. |
The gap between simulation and reality is significant, and the computational cost is high.
Steam and the actual texture of butter: the simulation never reproduces them perfectly. |
| Typical Use | Object manipulation, navigation in a known environment.
Fast food: same recipe, same station, a thousand times a day. |
Autonomous vehicles, general-purpose robotics…
Gourmet cuisine: improvising based on daily deliveries, managing multiple dishes simultaneously. |
In other words, world models provide the framework for reasoning, and action models are the driving force behind it. However, integrating the two remains an open problem, and the field has become industrialized before this problem has been solved. The dominant “ ” strategy—which involves replacing real-world data with computation and simulation—makes sense for companies that sell computing power, such as Nvidia. However, it is challenged by players such as Physical Intelligence, the startup founded by Chelsea Finn (Stanford) and Sergey Levine (Berkeley), which instead focuses on the massive collection of real-world data. Their results ultimately suggest that simulation does not yet resolve the issues of tactile data and deployment timelines.
Physical data: AI’s new black gold
This dilemma echoes the one that shaped the rise of large language models (LLMs): is more data better, or more computing power? Both. However, for physical AI, the debate is starting from scratch—but this time, the data isn’t on the internet; it’s in factories, warehouses, kitchens, and so on. Whoever collects it first may gain the same lead that OpenAI had over its competitors in 2020 with GPT-3. Thus, for European companies looking to integrate these systems, the question is not merely one of choosing hardware, but of determining which training paradigm we want to rely on—a decision that impacts both industrial strategy and digital sovereignty.
In any case, physical AI is everywhere at VivaTech, notes Philippe Lucas, Executive Vice President of Partnerships, Content, and Devices at Orange: “Here we see many robots, especially humanoids, because they are highly visible. However, we are also beginning to see a lot of small robots that could be used to assist elderly people who live alone, helping ensure they can continue living in their homes. We’re still in the early stages, much like we were nearly twenty years ago when the first apps appeared on mobile phones. Not all of those early apps were very useful, but today, it’s hard to imagine life without your mobile phone.” Specialized robots will be useful in the B2B context, and Orange intends to play a role in this regard: “Orange does not build robots, but we are planning to transform these products into services. That is where Orange can add value in this business.”
This text has been translated by an artificial intelligence.
Philippe Lucas







