Investment in Physical AI is accelerating, and the momentum shows little sign of slowing down. Some of the biggest names in the space have already attracted billions in funding. Figure has raised around $1.9 billion, while Skild AI and Applied Intuition have raised approximately $1.8 billion and $1.2 billion, respectively. NVIDIA CEO Jensen Huang has also repeatedly highlighted Physical AI as the next major wave of AI, following the rise of generative AI and agentic AI. The market opportunity is equally significant.
According to Strategy& analysis, Physical AI could unlock a global market of around EUR 430 billion by 2030, spanning industries such as manufacturing, logistics, autonomous driving, healthcare, robotics, aerospace, and more.
But what exactly is Physical AI? And what makes it different from the AI systems we use today?
From humanoid robots working alongside people to autonomous vehicles navigating real roads, Physical AI is moving from research labs into real-world environments.
In this article, we’ll explore what Physical AI is, how it works, the different types of Physical AI, and what it could mean for the future of AI and robotics.
Definition of Physical AI
Physical AI is artificial intelligence that senses, reasons about, and acts on the physical world through a body, whether that body is a robotic arm, a humanoid chassis, a drone, or the sensor suite of a self-driving car. It combines the same core ingredients used in modern AI, computer vision, language understanding, and machine learning, with actuators and sensors that let a system do something in three-dimensional space rather than only produce an answer.
Every physical AI system is built from the same three parts:
- Sensors that capture what is happening: cameras, LiDAR, tactile skins, microphones, inertial measurement units.
- A reasoning layer, usually a foundation model trained on large multimodal datasets, that interprets the scene and decides what to do.
- Actuators, motors, grippers, wheels, that carry out the decision and change the state of the world.
NVIDIA has framed this as the fourth stage in a broader arc of AI progress. Perception AI, dating back to AlexNet in 2012, taught systems to recognize and interpret images, sound, and sensor data. Generative AI taught systems to produce new text, images, and code from patterns learned in data. Agentic AI added planning and tool use, mostly within software. Physical AI is the stage where a system also has to reason about physics itself: friction, inertia, gravity, and cause and effect, then act on that reasoning in real time.
| Stage | What it does | Example |
| Perception AI | Recognizes and interprets multimodal signals | Computer vision for defect detection, lane recognition |
| Generative AI | Creates new content from learned patterns | Large language models, image generators |
| Agentic AI | Plans and executes multi-step tasks with tools | Software agents that browse, code, or orchestrate workflows |
| Physical AI | Perceives, reasons about physical laws, and acts in the real world | Humanoid robots, autonomous vehicles, robotic arms |
Physical AI is closely related to another term you’ll often see in robotics and AI research: embodied AI. The idea comes from embodied cognition, the view that physical experience plays an important role in how we learn and reason. A robot can’t fully understand what it means to lift a heavy box or grip an egg without breaking it just by reading text. It needs to interact with the physical world and learn from those interactions.
In practice, Physical AI and embodied AI are often used interchangeably. The main difference is where you’ll see the terms used. “Physical AI” is more common in industry and business discussions, while “embodied AI” appears more often in robotics research and academic papers.
This is also why Physical AI draws on more than computer science. Robotics teams often look at human movement, biomechanics, and how people interact with objects when designing systems that need to move, grasp, or manipulate things.
And Physical AI is not limited to humanoid robots. Autonomous vehicles, surgical robots, inspection drones, and smart manufacturing systems can all be considered Physical AI. What they have in common is simple: they sense what is happening in the physical world, make decisions, and take action based on what they perceive.
What Makes Physical AI Different from Traditional AI?
Traditional and generative AI systems are built to produce information: an answer, an image, a block of code. Physical AI is built to produce consequences. That distinction shows up in a handful of concrete ways.
| Aspect | Traditional/generative AI | Physical AI |
| Output | Text, image, audio, code | Physical motion: grasping, walking, steering |
| Feedback loop | Often one-shot: prompt in, response out | Continuous: sense, act, sense the result, adjust |
| Cost of error | A wrong answer can be regenerated | A dropped part, a collision, or a fall carries a real, sometimes irreversible cost |
| Data source | Text and images already available on the internet | Data that has to be physically captured: sensor readings, demonstrations, motion |
| Reasoning required | Statistical and semantic reasoning | Semantic reasoning plus physical reasoning: gravity, inertia, contact dynamics |
Physical reasoning is one of the biggest differences between Physical AI and systems built for text. A physical AI system needs to predict where a rolling ball will go, know how much force to use when picking up a wine glass, or recognize that a pedestrian may be hidden behind a parked car. These tasks require an understanding of how objects and environments behave, learned from real-world data or simulation.
The feedback loop is just as important. A language model generates an answer and stops. A physical AI system has to continuously sense, decide, and act as the environment changes. It may need to repeat this process dozens of times per second to keep a robot balanced, adjust its grip, or respond to a moving person or object. This makes Physical AI a different engineering challenge from AI systems that only work with text, images, or other digital inputs.
Safety is another major difference. A wrong answer from a language model can be a product or user experience issue. A warehouse robot that misjudges a pallet’s weight or a self-driving vehicle that misses a crosswalk can cause real-world harm. These systems also have to meet safety, regulatory, legal, and insurance requirements.
As a result, Physical AI requires more testing and validation before deployment. The goal is not just to build a model that works; it is to make sure the system can act reliably and safely in the real world.
How Does Physical AI Work?
Every physical AI system, whether it is a warehouse robot or a self-driving car, runs the same basic loop. Sensors capture the current state of the world: camera frames, LiDAR point clouds, tactile readings, sometimes audio or inertial data. A reasoning layer interprets that state alongside a goal or instruction and decides what to do next. Actuators carry out the decision, changing the world, and the sensors capture the new state so the loop runs again. In a mobile robot or humanoid, this cycle can repeat dozens of times per second.
Two families of models currently define how that reasoning layer actually works.
Vision-language-action models. Vision-language-action (VLA) models are the part of the stack that most directly resembles a robot’s brain. A VLA model takes in a camera image or video along with a plain-language instruction and outputs the next physical action: a joint angle, a gripper command, a steering adjustment. Rather than a hand-built pipeline where separate systems handle perception, planning, and control, a VLA model collapses that chain into a single model trained end to end.
World models for physical AI. Where a VLA model is reactive, looking at a scene and producing the next action, a world model is predictive. It is trained to simulate how an environment will change over time, including in response to actions that have not happened yet. NVIDIA frames its Cosmos platform this way: instead of predicting the next word in a sentence, Cosmos predicts what happens next in physical space. Drop a ball into a Cosmos-generated simulation, and the model predicts its trajectory, its bounce, and where it settles.
Explore now:
What Is a VLA Model? How Vision-Language-Action Models Enable Physical AI
What Is an AI World Model? A Complete Guide
What Are the Types of Physical AI?

Physical AI is not one product category. It spans several distinct classes of systems, each shaped by a different mix of hardware constraints, tasks, and data requirements.
Humanoid robots. Humanoid robots draw the most attention because a human-shaped body can, in theory, operate in spaces built for people: the same aisles, tools, and doorways, without redesigning a facility around a robot. The field moved quickly through 2026. Unitree’s G1 ships commercially from roughly $16,000, making it the accessible end of the market. Figure’s Figure 03 is running a manufacturing pilot with BMW. Agility Robotics’ Digit has documented warehouse testing with Amazon alongside a dedicated production facility in Oregon built for volume output. Boston Dynamics retired the hydraulic version of Atlas in favor of an all-electric model aimed at commercial deployment, with initial units heading to Hyundai’s Metaplant. Tesla’s Optimus remains in internal deployment at Tesla’s own factories, though the company has not published verified production or unit figures. None of this means humanoids are a solved category. Independent industry coverage through mid-2026 has pointed out that many widely circulated deployment numbers do not hold up against what companies have actually confirmed, and that the sim-to-real gap remains the dominant open question across the board.
Autonomous mobile robots and industrial arms. Away from the humanoid spotlight, autonomous mobile robots already move inventory through warehouses and distribution centers, and stationary robotic arms from manufacturers like Franka Robotics and Universal Robots handle repetitive assembly, packing, and inspection on factory lines. These systems are usually narrower in scope than a humanoid but further along in real deployment, since a wheeled base or a fixed arm is mechanically simpler to control than a two-legged body that also has to balance.
Autonomous vehicles. Self-driving cars and trucks apply physical AI to a single, high-stakes task: navigating roads safely. NVIDIA’s automotive stack, including its DRIVE platform and its newer Alpamayo models for driving policy, runs the same perception, reasoning, and action loop used in robotics, adapted to the physics of a vehicle moving at speed among other traffic. Autonomous driving teams draw on the same broad categories of data as humanoid robotics teams: camera and LiDAR footage and driving trajectories, plus simulation-generated edge cases such as rare weather or lighting conditions that are too dangerous or too infrequent to capture naturally.
Drones and UAV systems. Drones extend physical AI into open, often unstructured airspace. Infrastructure inspection, agricultural monitoring, and delivery are the three use cases seeing the most commercial traction. A drone’s version of the perception-action loop still has to account for wind and GPS-denied environments, and the cost of real-world failure is higher than for a ground robot, since a mid-flight crash is usually unrecoverable.
Specialized domains. A smaller but fast-growing set of applications puts the same architecture to work in more specialized settings: surgical and rehabilitation robots that assist rather than replace a clinician, construction robots that inspect facades or assist with repetitive on-site tasks, and smart manufacturing systems that catch defects or adjust a process in real time. NVIDIA has said medical robotics is one of the next areas it plans to expand its physical AI tooling into.
| Type | Primary environment | Example companies or platforms |
| Humanoid robots | Factories, warehouses, human-built spaces | Tesla Optimus, Figure 03, Unitree G1/H2, Boston Dynamics Atlas, Agility Digit, 1X NEO |
| Autonomous mobile robots and arms | Warehouses, factory floors | Franka Robotics, Universal Robots, warehouse AMR fleets |
| Autonomous vehicles | Public roads | NVIDIA DRIVE and Alpamayo-based stacks, automotive OEMs |
| Drones & UAV Systems | Open air, infrastructure sites, farmland | Inspection, agriculture, and delivery operators |
| Specialized robotics | Hospitals, construction sites, production lines | Surgical assist systems, construction robots, smart manufacturing |
What Data Does Physical AI Need?

Every system described above depends on a category of training data that looks nothing like the text and images that trained the first generation of large language models. Physical AI needs to learn what the world feels like, not just what it looks like, and that data has to be physically captured rather than scraped.
Multimodal data. Physical AI systems train on multimodal data: synchronized streams of vision, language, action, and often audio and depth captured together in time. A single training example might pair a camera frame, a spoken or typed instruction, and the exact joint movement a robot, or a human demonstrator, made in response. Because these streams have to line up precisely, in some tactile-focused datasets, vision runs at 20 to 30 frames per second while tactile sensors sample at 100Hz or higher, multimodal robotics data is far harder to assemble at scale than a scraped text corpus and far more sensitive to timestamp drift or calibration error between sensors.
Egocentric data. Egocentric data, first-person video of humans performing everyday tasks, has become one of the more efficient ways to pretrain a physical AI system before it ever touches a robot. Datasets like Ego4D, at roughly 3,670 hours, and newer efforts scaling past 20,000 hours give a model broad exposure to how people actually grasp, carry, and manipulate objects in ordinary settings. Filming someone tidying a kitchen is far cheaper than operating a robot arm through the same task thousands of times, and research has repeatedly shown that performance keeps scaling with the volume of egocentric video used in pretraining.
Teleoperation data for robotics. Teleoperation data narrows the gap between demonstration and deployment. A human operator directly controls a robot, often through a VR headset, haptic gloves, an exoskeleton, or a handheld interface modeled on the Universal Manipulation Interface, while the robot logs its own sensor readings and the exact actions taken. Because these demonstrations come from the robot’s own body rather than a human’s, teleoperation data tends to transfer more directly into a working policy than egocentric human video, at a higher collection cost per hour. Full-body teleoperation, capturing locomotion and manipulation together on a humanoid platform, has become one of the fastest-growing categories of robotics data collection through 2026.
Motion data. Motion data, captured through motion-capture suits, exoskeletons, and marker-based tracking systems, records how a body, human or robot, moves through space over time: joint angles, velocities, balance shifts. For humanoid robots in particular, motion data is what teaches a system to walk, recover from a stumble, or coordinate a whole-body reach instead of moving one limb at a time.
Tactile data. Vision alone cannot tell a robot how hard it is gripping an egg or whether a surface is slick. Tactile data, gathered through force-torque sensors and tactile skins such as DIGIT, GelSight, and AnySkin, fills that gap by recording pressure, shear force, and contact patterns directly from a robot’s fingers or a human operator’s gloves. Contact-rich tasks, folding fabric, inserting a connector, handling fragile objects, tend to fail without it, since a policy trained purely on vision has no way to know how much force it is actually applying.
Simulation and synthetic data. Real-world collection, teleoperation, egocentric filming, and motion capture do not scale infinitely, so simulation fills the remaining volume. Platforms like NVIDIA Omniverse and Cosmos Transfer take a small set of real demonstrations and generate large volumes of physically plausible variations, different lighting, textures, object placements, and room layouts, without a camera or a robot running for every version. This is also where world models double as data generators rather than only planning tools, narrowing the gap between how much data a model needs and how much a team can physically capture.
Pulling a training pipeline together across all of these formats, video, robot action logs, motion capture files, tactile readings, and simulation output, is closer to running a physical production line than a typical data labeling project. It calls for annotators who understand robot kinematics well enough to label a failed grasp correctly, quality checks that catch sensor drift a text reviewer would never notice, and infrastructure that keeps every stream time-aligned across an entire dataset. Quality control looks different here too: a mislabeled sentiment tag in a text dataset is a nuisance, while a mislabeled contact frame in a manipulation dataset can teach a policy the wrong grip force entirely, so review passes tend to focus on kinematic and contact accuracy rather than only labeling consistency.
This is the gap that dedicated physical AI companies and data partners exist to close, handling multimodal collection, egocentric and teleoperation data, motion capture, and tactile annotation so robotics and autonomous driving teams can spend engineering time on the model rather than the data pipeline. LTS GDS runs this kind of data operation for Physical AI and robotics clients across the US, China, and Europe, with quality acceptance rates of up to 97 percent across deliveries covering egocentric video collection, teleoperation, and tactile annotation.
Read more: Egocentric Data Collection: Building High-Quality AI Datasets for Physical AI and Robotics
FAQs About Physical AI
1. What is physical AI, in simple terms?
Physical AI is artificial intelligence built to perceive, reason about, and act on the physical world, usually through a robot, vehicle, or drone, rather than only producing text or images on a screen. If a system can look at its surroundings, understand an instruction, and then actually move, grasp, or steer in response, it counts as physical AI.
2. How can physical AI transform industries?
Physical AI changes what can be automated, not just how fast something already automated runs. In logistics, it allows warehouses to use robots that adapt to changing layouts instead of following fixed paths. In manufacturing, it supports lines that catch defects or adjust a process without a full retooling cycle. In healthcare, it assists surgeons and supports rehabilitation. In agriculture, drones and ground robots monitor crops and take on repetitive fieldwork. In construction, it handles inspection and repetitive assembly in environments that are often unsafe for people. In autonomous driving, it moves vehicles from following pre-mapped routes toward reasoning about unpredictable traffic in real time. Across every industry, the common thread is the same: physical AI extends automation into tasks that used to require a person’s judgment as well as their hands.
3. What is a physical AI company?
A physical AI company is any organization building the models, hardware, or data infrastructure that let machines sense and act in the real world. That covers infrastructure providers like NVIDIA, robot builders like Figure, Tesla, Unitree, Boston Dynamics, and Agility Robotics, foundation model labs like Physical Intelligence and Google DeepMind, and data partners like Scale AI, Appen, and LTS GDS that supply the multimodal, egocentric, teleoperation, and tactile datasets those models are trained on. Most physical AI products on the market today result from several of these companies working together rather than one company owning the entire stack.
4. What data does physical AI need to work reliably?
Physical AI depends on multimodal data that pairs vision, language, and action together, plus more specialized formats: egocentric video of humans performing tasks, teleoperation logs from human-controlled robots, motion capture of body movement, tactile readings from contact sensors, and simulation-generated data to cover scenarios that are too rare, expensive, or dangerous to collect in the real world.
5. How much data does it take to train a physical AI model?
It depends on the stage. Pretraining a generalist model from scratch typically draws on hundreds of thousands to millions of demonstrations, or thousands of hours of egocentric video, since the model is learning broad physical common sense from a wide spread of tasks and objects. Fine-tuning an already-pretrained model onto a specific robot and task usually needs far less, often a few dozen to a few thousand high-quality demonstrations, because the model is adapting existing knowledge rather than learning from zero. At the fine-tuning stage, data quality and consistency tend to matter more than raw volume.
Making Physical AI Work in the Real World
Physical AI marks the point where AI stops being confined to a chat window and starts sharing physical space with people: lifting, walking, steering, gripping, and occasionally failing in ways that carry real consequences. That a system needs to perceive its environment, reason about physical laws like friction and inertia, and act in real time is not different from what NVIDIA’s Jensen Huang described when framing it as the stage after perception, generative, and agentic AI. VLA models, world models, and humanoid hardware are approaching commercial price points and are mature enough to leave the research lab.
None of that progress happens without the underlying data. A VLA model or a world model is only as capable as the multimodal, egocentric, teleoperation, motion, and tactile data it was trained on, and that data has to be physically captured, carefully synchronized, and correctly annotated rather than pulled from the open internet. For founders, data leaders, and engineering teams building on humanoid platforms, autonomous vehicles, drones, or industrial automation, the data strategy is arguably as consequential as the choice of model architecture.
Organizations building the next generation of Physical AI can work with LTS GDS for end-to-end data solutions covering multimodal collection, egocentric and teleoperation data, motion capture, and tactile annotation across humanoid, autonomous driving, and drone use cases. Teams scoping a data partner for their next physical AI project can connect with the LTS GDS engineering team to discuss requirements.








