Vision-language-action models are the AI systems behind a new wave of general-purpose robots, and that wave is starting to show up in shipment data: roughly 18,000 humanoid robots shipped worldwide in 2025, up about 508% year over year (Source: IDC Worldwide Humanoid Robotics Market Analysis 2026).
Most of those machines still do narrow jobs. The models steering the most capable ones do not.
A VLA model looks at a scene through a camera, reads an instruction such as “put the red mug in the dishwasher,” and outputs the motor commands to do it. One model, many tasks, taught with words and demonstrations instead of hand-written code. That’s the pitch. The reality is messier, and the gap between the two matters if you are evaluating robotics vendors, planning automation budgets, or just trying to tell a real demo from a staged one.
Key Takeaways
- VLA models map camera images and text instructions directly to robot actions.
- Most start as a pretrained vision-language model, then learn from robot data.
- Google DeepMind, Physical Intelligence, NVIDIA, and Figure lead current development.
- Robot training data, not model size, is the main bottleneck.
- Generalization is improving fast but is not yet reliable unsupervised.
Table of contents
What Is a Vision-Language-Action Model?
A vision-language-action model is a single neural network that takes visual input and a natural-language instruction and produces physical actions for a robot. Vision and language come in. Movement goes out.
The term entered wide use in July 2023, when Google DeepMind released RT-2. Its researchers took a vision-language model trained on web images and text, then fine-tuned it on robot demonstrations, writing each arm movement as a short string of numbers the model could predict the same way it predicts words. The result could handle objects and instructions it had never seen in robot training. Google reported that RT-2 roughly doubled performance on unseen scenarios compared with its predecessor, RT-1, moving from 32% to 62%.
That was the breakthrough. Before VLAs, most industrial robots ran on explicit programs: move to these coordinates, close the gripper, move here. Change the part or shift the bin six inches and someone rewrites the code. A VLA inherits general knowledge about the world from internet-scale pretraining, so “pick up the thing you’d use to hammer a nail” can map to a rock on the table without anyone labeling rocks as hammers.
If you already know multimodal AI systems, the easiest way to place VLAs is this: a vision-language model describes the world, and a VLA acts on it.
How VLA Models Work
Under the hood, almost every current VLA shares the same three-part structure.

Seeing and reading
Camera frames pass through a vision encoder that turns pixels into numerical tokens. The instruction gets tokenized like any chatbot prompt. Many models also take in the robot’s own state, meaning joint angles and gripper position, so they know where the arm is before deciding where it goes.
All of this feeds into a pretrained vision-language backbone where the “common sense” lives: what a mug is, what “left of” means, that glass breaks. OpenVLA, an open model released by Stanford-led researchers in June 2024, uses a 7-billion-parameter backbone. Physical Intelligence’s π0 started from Google’s PaliGemma.
Turning thoughts into motion
The action head is where designs split. Early models like RT-2 and OpenVLA predict actions as discrete tokens, chopping each joint movement into one of 256 bins. It works, but it’s jerky and slow for fine manipulation.
Newer models generate continuous “action chunks,” a short sequence of smooth movements covering the next fraction of a second. π0 does this with a technique called flow matching and can output commands at up to 50 times per second. That’s the difference between a robot that pokes at a shirt and one that folds it.
Fast and slow thinking
A large model is too slow to react in real time so several teams split the job in two.
Figure’s Helix, announced in February 2025, pairs a 7-billion-parameter model that reasons about the scene at about 7 to 9 updates per second with an 80-million-parameter controller that adjusts the robot’s upper body 200 times per second. The slow system decides what to do while the fast one keeps the hand from knocking the cup over while doing it.
Google DeepMind’s Gemini Robotics line takes a similar route from a different angle, pairing an embodied-reasoning model that plans multi-step tasks with a VLA that executes them.
The Models That Defined the Category
The field moved from a single research paper to a crowded commercial race in about three years.

A few releases shaped where things stand now:
| Model | Developer | Released | Why it matters |
|---|---|---|---|
| RT-2 | Google DeepMind | July 2023 | Proved web-pretrained models can control robots |
| OpenVLA | Stanford-led consortium | June 2024 | First widely used open-weight VLA |
| π0 | Physical Intelligence | October 2024 | Smooth, high-frequency control via flow matching; weights opened in 2025 |
| Helix | Figure | February 2025 | Dual-system design for full humanoid upper-body control |
| GR00T N1 | NVIDIA | March 2025 | Open humanoid foundation model tied to NVIDIA’s simulation stack |
| SmolVLA | Hugging Face | June 2025 | Roughly 450 million parameters, small enough for consumer hardware |
| π0.7 | Physical Intelligence | April 2026 | Early evidence of combining skills to solve untrained tasks |
| Gemini Robotics 2 | Google DeepMind | July 2026 | Full-body humanoid control, plus an on-device variant |
Open data made much of this possible. The Open X-Embodiment dataset, pooled in 2023 by researchers from 21 institutions, combined demonstrations from 22 robot types covering 527 distinct skills. OpenVLA was trained on roughly 970,000 real-world robot episodes drawn from it, and its authors report it beat Google’s far larger RT-2-X by 16.5 percentage points in task success. Without shared data, only the largest labs could have competed.
The most recent entries push hardest on generalization. When Physical Intelligence released π0.7, it showed the model operating an air fryer after seeing only two loosely related training clips. Gemini Robotics 2 extends control from hands to the whole body, managing balance so a humanoid can reach and carry without tipping over.
Why VLA Models Matter for Business
Robots are already everywhere in factories. The International Federation of Robotics counted 5 million industrial robots in operation worldwide in 2025, with more than 600,000 new installations that year (Source: IFR World Robotics 2026). The catch is that nearly all of them repeat one programmed motion in a controlled cell.
VLAs target the work those robots cannot do: tasks with variation. Consider a few realistic starting points.
Warehouse picking and packing. Items change shape, size, and packaging weekly. A model that understands “the blue box behind the bottle” handles variety that fixed programs choke on. The same logic runs through broader AI in logistics work.
Light manufacturing and kitting. Short production runs make reprogramming expensive. Teaching a new task by demonstration and instruction cuts changeover time, which is the problem manufacturers hit when moving AI from pilot to production.
Machine tending and inspection. Loading parts into machines and flagging defects combine perception with simple manipulation, a good fit for current capability.
Unstructured sites. Construction and field work sit further out, but the direction is clear, as seen in how AI is changing construction sites already.
Household robots get the headlines. Commercial settings will likely see useful deployments first, because the environment is more predictable and a human supervisor is nearby.
Where VLA Models Still Fall Short
Here’s the honest part.
Data is scarce
Language models learn from trillions of words scraped from the web. There is no web of robot actions. Every training episode has to be collected by a physical robot, usually driven by a human operator, which makes data slow and expensive to gather. This is the core constraint on the whole field, and it is why companies are racing to collect demonstrations at scale and generate synthetic data in simulation.
Reliability is not there yet
A demo that works 80% of the time looks impressive on video. On a production line, a 20% failure rate is a disaster. Physical Intelligence’s own researchers described π0.7’s results as early signs of generalization, not deployment-ready technology. In one test, success jumped from 5% to 95% after the team reworded the task instructions. Useful, yes. Also a reminder that these systems are sensitive to phrasing in ways a factory cannot tolerate.
Nobody agrees on the scorecard
There is no standard benchmark equivalent to what exists for language models. Most results come from the developer’s own lab, on the developer’s own robot. Arena-style comparisons such as RoboArena are emerging, but vendor claims remain hard to compare directly. Ask for task-level success rates in conditions that match yours.
Safety and compute
A model that misreads an instruction in a chat window produces bad text. One that misreads it with a 30-kilogram arm can injure someone. Most deployments still keep humans close or confine robots to fenced areas. Running large models on the robot itself also demands serious onboard computing, which is why on-device versions such as Gemini Robotics On-Device and SmolVLA exist.
A rival architecture is emerging
VLAs start from language understanding. A newer approach, called world-action models, starts from video models that already predict how scenes change over time. NVIDIA has said its next model, GR00T N2, will use this architecture. Whether that replaces VLAs or merges with them is one of the open questions for 2027.
How to Get Started With VLA Models
Your path depends on which side of the table you sit on.
If you are an engineer, start small and open. Hugging Face’s LeRobot library and SmolVLA run on modest hardware, and Physical Intelligence publishes π0-family code and weights through its openpi repository. OpenVLA remains a well-documented baseline. A low-cost robot arm and a few dozen of your own demonstrations will teach you more about the data problem than any paper.
If you are an executive evaluating vendors, skip the demo reel and ask:
- What is the task success rate in an environment like ours, measured over hundreds of attempts?
- How many demonstrations does a new task require, and who collects them?
- What happens when the robot fails, and how does it signal uncertainty?
- Does the model run on the robot or depend on a network connection?
- Who owns the data your operations generate?
Vendors with good answers to the second and third questions are the ones worth a pilot.
Conclusion
Vision-language-action models change how robots are taught. Instead of programming every motion, you show and tell, and the model draws on general knowledge learned from the internet to fill the gaps. In three years, the field went from RT-2’s proof of concept to full-body humanoid control and early signs of real skill transfer.
For you, the practical read is patience with direction. Expect the first reliable business value in semi-structured settings like warehouses and light manufacturing, with a human nearby. Judge vendors on measured success rates and data requirements, not videos, and watch whether world-action models reshape the architecture before you commit to a long-term platform.
Read Next
For more on how AI is moving from software into physical operations, start here:
- AI-Powered Tools That Are Revolutionizing Manufacturing Efficiency
- Engineering Enterprise Platforms for Intelligent Automation
- Utilizing AI Hyperautomation Frameworks
Frequently Asked Questions
Vision-language-action models are AI models that take camera images and a natural-language instruction as input and output robot actions. They are usually built by fine-tuning a pretrained vision-language model on robot demonstration data. This lets a single model handle many tasks without task-specific programming.
VLA models add an action output to a vision-language model. A vision-language model can describe an image or answer questions about it, while a VLA model uses that understanding to produce motor commands. In practice, most VLAs are a vision-language model with an action head attached and trained on robot data.
Vision-language-action models are built by Google DeepMind (Gemini Robotics), Physical Intelligence (the π models), NVIDIA (Isaac GR00T), and Figure (Helix), among others. Open-weight options include OpenVLA, SmolVLA from Hugging Face, and Physical Intelligence’s openpi releases.
VLA models are ready for supervised pilots in semi-structured settings, but not for fully unsupervised deployment. Reliability, sensitivity to how instructions are phrased, and the lack of standard benchmarks remain real limits. Most current deployments keep a human nearby.
The biggest challenge for vision-language-action models is training data. Unlike text, robot action data has to be collected on physical hardware, which is slow and costly. Simulation, shared datasets, and large-scale teleoperation programs are the main ways labs are trying to close the gap.











