Billions of dollars are now chasing an idea that robotics researchers have worked on for decades: machines that can move and work like people, or in other words, humanoid AI.
Apptronik raised $520 million in February 2026; Figure AI has raised more than $1 billion in its Series C, and companies such as Agility Robotics and 1X are pushing humanoid robots into factories, warehouses, and other real-world environments.
What makes this moment different isn’t the shape of the robot. It’s the intelligence behind it.
A humanoid can now use cameras and sensors to understand its surroundings, follow a natural-language instruction, and figure out how to move its body to complete a task. That intelligence layer is what humanoid AI is about.
In this article, we’ll look at how humanoid AI works, the technologies behind it, where humanoid robots are being used, and what still stands between today’s prototypes and truly scalable deployment.
What is Humanoid AI?

Humanoid AI is the intelligence that allows a humanoid robot to understand and act in the physical world. A humanoid robot has a human-like body: a torso, two arms, two legs, hands, cameras, sensors, motors, and batteries. But having a body that looks like a person does not make a robot intelligent. The robot still needs to understand what it is seeing, figure out what it has been asked to do, decide what actions to take, and coordinate its movements to carry them out. Humanoid AI provides that layer of intelligence.
For example, imagine asking a humanoid robot to “pick up the box and put it on the shelf.” A traditional industrial robot may need the box, shelf, positions, and movement sequence to be carefully defined in advance. A humanoid robot powered by AI can instead use its cameras and sensors to identify the box, locate the shelf, estimate how to reach it, plan a sequence of movements, and adjust those movements as conditions change.
That process involves several capabilities working together:
- Perception: Understanding objects, people, space, and events through cameras and other sensors.
- Language understanding: Turning human instructions into tasks the robot can act on.
- Reasoning and planning: Breaking a goal into smaller actions and deciding what to do next.
- Motion and control: Translating those decisions into coordinated movements of the robot’s joints, hands, and body.
- Learning and adaptation: Using experience and feedback to handle new objects, environments, and variations in a task.
Read more: What Is Robotic Manipulation? How Robots Learn to Interact with the Physical World
How Does Humanoid AI Work?

Strip away the launch reels and every humanoid AI system, regardless of which company built it, runs the same basic loop.
Cameras and sensors capture the scene. A model interprets what it sees and what it has been asked to do. That model outputs a decision. The decision gets translated into motor commands that move the robot’s joints. The robot’s sensors capture the new state of the world, and the loop repeats, often dozens of times each second.
The differences between one system and another show up inside that loop: how perception is built, which model does the deciding, how many motors and degrees of freedom the body has to coordinate, and what data taught the model to make good decisions in the first place.
Perception: how a humanoid reads its surroundings
A humanoid uses cameras, depth sensors, IMUs, joint encoders, and force sensors to understand both its surroundings and its own body. These sensors help it recognize objects, track movement, maintain balance, and control how much force to use when gripping something.
The data is processed on the robot, where fast response matters. The perception system turns raw images, depth, and sensor readings into a real-time understanding of the environment.
It also needs to understand what it is seeing, not just where things are. Recognizing a cup, for example, means knowing what it is, how it can be handled, and where to grip it safely. Microphones can also let the robot receive spoken instructions, which becomes especially useful in service and home environments.
Cognition: the model deciding what happens next
This is the layer often described as the robot’s “brain.” Modern humanoids increasingly use vision-language-action (VLA) models that connect visual input and language instructions to physical actions, bringing perception, planning, and control closer together in a single system.
Learn more: What Is a VLA Model? for a deeper look at how these models work.
Several approaches are shaping this layer. NVIDIA’s Isaac GR00T, Google DeepMind’s Gemini Robotics, and Physical Intelligence’s pi-0 all use foundation-model techniques to connect what a robot sees and understands with what it should do. Some systems also separate high-level reasoning from fast motor control, allowing the robot to plan ahead while still reacting quickly to changes around it.
World models take a different approach. A world model is a learned simulation of the robot’s environment, a model trained to predict how a scene changes after a given action, rather than one trained only to copy demonstrations. With that kind of predictor, a robot can test a candidate action inside the simulation first and only carry out the one predicted to work, instead of learning purely by trial and error on real hardware.
Actuation and whole-body control
Deciding what to do is only half the challenge. The robot must turn that decision into coordinated movement across dozens of motors while staying balanced. Hands are especially important because many real-world tasks depend on precise manipulation.
Whole-body control keeps the robot stable as it walks, reaches, or carries uneven loads. Locomotion and balance are often trained through reinforcement learning in simulation before being transferred to the physical robot, while manipulation can be handled by the same VLA models used for perception and planning. Coordinating these systems in real time remains one of the harder challenges in humanoid robotics.
Power is another practical constraint. Most current humanoids operate for roughly a work shift before needing a recharge or battery swap. Speed, payload, and battery life also involve trade-offs, so these capabilities continue to change as hardware improves.
How Are Humanoid AI Systems Classified?
Humanoid AI systems can be classified in several ways, depending on their physical design, level of autonomy, and intended use. These categories are not mutually exclusive: a single robot can be bipedal, supervised-autonomous, and designed for general-purpose industrial tasks at the same time.
1. By form factor
Bipedal humanoids use two legs and a human-like upper body, allowing them to navigate stairs, uneven surfaces, and workspaces designed for people. Examples include Tesla Optimus, Figure, Apptronik Apollo, and Agility Digit.
Wheeled humanoids combine a human-like torso and arms with a wheeled base. They sacrifice some mobility for greater stability, simpler locomotion, and potentially lower cost, making them suitable for environments such as warehouses and homes.
2. By level of autonomy
Humanoids can also range from teleoperated systems, where a human directly controls the robot, to supervised-autonomous systems, where the robot handles tasks independently but a human can intervene, and eventually fully autonomous systems that operate without real-time human control.
Most commercial deployments today sit somewhere between teleoperation and full autonomy, particularly when robots encounter tasks or situations outside their trained capabilities.
3. By task scope
Another important distinction is between task-specific and general-purpose humanoids.
Task-specific systems are optimized for a defined job, such as moving totes or handling parts on a production line. General-purpose humanoids aim to perform a wider range of tasks and adapt to different environments, making them more flexible but also significantly harder to train and deploy reliably.
These categories help explain why two robots that look similar can have very different capabilities, training requirements, and deployment models.
Where Is Humanoid AI Being Used?
Manufacturing and automotive
Manufacturing is currently one of the strongest deployment grounds for humanoid AI. In mid-2024, Figure AI tested its Figure 02 robot at BMW’s Spartanburg plant, where it completed a multi-week trial operating 10-hour shifts to insert sheet metal parts into chassis fixtures, placing over 30,000 components into production vehicles.
Mercedes-Benz has partnered with Apptronik to test its Apollo humanoid for component delivery and quality inspections at facilities in Germany and Hungary. Meanwhile, Tesla continues to train Optimus on internal logistics and simple assembly tasks at its Fremont and Gigafactory sites as it prepares for broader internal deployment.
Warehousing and logistics
Warehousing remains a leading application because workflows involve repetitive material transport within structured, predictable environments. Agility Robotics’ Digit is deployed at a GXO-operated SPANX fulfillment facility in Georgia under a multi-year Robots-as-a-Service (RaaS) agreement. Agility also maintains active commercial pilots with enterprise partners including Amazon, Toyota Motor Manufacturing Canada, Mercado Libre, and Schaeffler.
Designed for tote-handling and logistics, Digit carries up to 35 pounds (16 kg) with a continuous battery runtime of around 1.5 to 2 hours per charge. To support long-term supply, Agility established its RoboFab facility in Salem, Oregon, with an intended production capacity of up to 10,000 units per year.
Healthcare and facility operations
Healthcare remains an early-stage market for humanoid AI. Near-term opportunities focus on non-clinical logistics such as delivering linens, moving pharmaceuticals, and facility maintenance, rather than direct patient care. Clinical applications like surgical support or patient handling demand levels of reliability, safety, and regulatory compliance that current systems cannot yet guarantee.
Retail and hospitality
Retail and hospitality are exploring humanoids for reception, concierge duties, and customer guidance. However, these environments present substantial operational challenges compared to structured warehouses: human traffic is unpredictable, layouts change frequently, and machines must interact safely with the public. Consequently, front-of-house hospitality applications remain strictly in pilot testing.
Home and domestic assistance
The domestic environment is arguably the most complex use case. Companies like 1X Technologies are preparing humanoid robots such as NEO for home environments. To bridge current technical gaps, these consumer-focused designs rely on a hybrid architecture: basic autonomous operation paired with on-demand remote human assistance (teleoperation) for complex or unfamiliar household tasks.
FAQs about Humanoid AI
1. What is humanoid AI?
Humanoid AI is the software and model layer, perception, foundation models, planning, and motor control that lets a human-shaped robot sense its surroundings, understand instructions given in natural language, and decide how to move its body to complete a task. It sits on top of the mechanical body rather than being the body itself, and it is what separates a humanoid robot that can only follow a fixed script from one that can generalize to a task it has not seen before.
2. How much does a humanoid robot cost?
Prices vary widely by capability and market. In 2026, list prices span from Unitree’s R1, launched under $6,000, up through research-grade platforms priced above $150,000. Commercial platforms cluster in the middle: 1X’s Neo lists at $20,000, or $499 a month on subscription, and several analysts point to a $25,000 to $60,000 range as the point where a robot’s payback period beats the annual cost of a human shift worker, which is why price compression toward that range is one of the more closely watched trends in the industry.
3. Is humanoid AI ready for full autonomy?
Not broadly, not yet. Most deployed systems in 2026 run on hybrid autonomy: the model handles known, well-defined tasks on its own, and a remote human operator takes over through teleoperation when the robot hits something it has not learned to do, as with 1X’s Expert Mode on Neo. Even Agility Robotics’ Digit, one of the more mature deployments, operates autonomously within a defined, narrow work cell rather than across an open set of tasks. Full autonomy across unpredictable environments remains a research goal more than a shipped product.
4. What data trains humanoid AI models?
Three main sources feed most humanoid AI training pipelines: teleoperated demonstrations, where a human remotely controls a real robot to complete a task; egocentric human video and motion capture, which record people performing everyday tasks because filming humans is far cheaper than operating a robot; and simulation, where platforms like NVIDIA’s Omniverse and Isaac Sim generate large volumes of synthetic trajectories from a smaller set of real demonstrations. Most production models blend all three, since each source covers a gap the others leave.
5. Is humanoid AI the same thing as a VLA model?
Not quite. A vision-language-action model is the architecture most commonly used to build the cognition layer described earlier in this guide, the part that turns camera images and a language instruction into a robot action. Humanoid AI is the broader system that layer sits inside, including perception, whole-body control, and the training data pipeline behind all of it. A useful way to hold the two apart: a VLA model is one component that can, in principle, run on several different robot bodies, while humanoid AI describes the full stack purpose-built for a specific, human-shaped one.
The Road Ahead for Humanoid AI
Humanoid AI is moving from controlled demonstrations into early real-world use, driven by growing investment and increasing interest from governments and industries looking to develop the technology. The technology is already being tested and deployed in real environments, from warehouse cells and factory floors to a small number of early home applications. But it is still a long way from the general-purpose robot assistant often shown in demos and keynote presentations. In most deployments today, tasks remain relatively narrow, while teleoperation or human supervision still plays an important role. Looking at what humanoids can reliably do today, rather than what they may eventually be capable of, gives a more realistic picture of where the industry stands.
The data behind these systems is becoming just as important as the models themselves. Many leading humanoid AI systems combine vision-language-action models, simulation, teleoperation, and other learning approaches. Training them effectively requires large volumes of high-quality data, including teleoperation sessions, motion capture, and egocentric footage. Capturing, synchronizing, annotating, and validating this data at scale is a complex challenge in its own right.
For teams building humanoid and other Physical AI systems, LTS GDS provides data solutions covering teleoperation, motion capture, egocentric data collection, annotation, and quality assurance. These capabilities support AI training across humanoid robotics, warehouse automation, and autonomous systems. To learn more, explore the Physical AI & Robotics Data Services or contact the LTS GDS team to discuss your project requirements.









