Knowledge
Robot Foundation Models — the 2026 landscape
Since 2024 a race has been on for the “brain” of humanoid robots. This overview maps the most important robot foundation models with their current status (mid-2026). Vendor performance superlatives are often self-reported and not yet independently verified — treated cautiously here.
At a glance
| Provider | Latest model | Status | Type | License |
|---|---|---|---|---|
| Google DeepMind | Gemini Robotics 1.5 / ER 1.6 | 2025–2026 | VLA + reasoning VLM | proprietary |
| NVIDIA | Isaac GR00T N1.7 (N2 preview) | 2026 | VLA / world-action | open |
| NVIDIA | Cosmos (Predict/Transfer/Reason) | 2025–2026 | World model | open |
| Physical Intelligence | π0.7 | 2026 | VLA | partly open (π0) |
| Figure AI | Helix | since 2025 | VLA (System 1/2) | proprietary |
| Tesla | Optimus control net (unnamed) | — | end-to-end | proprietary |
| Skild AI | Skild Brain | 2026 | “omni-bodied” foundation | proprietary |
| 1X Technologies | Redwood + World Model | since 2025 | VLA + world model | proprietary |
| Hugging Face | SmolVLA / LeRobot | since 2025 | VLA / ecosystem | open |
| ByteDance | GR-3 | 2025 | VLA (Mixture-of-Transformers) | research |
| AgiBot | GO-1 (ViLLA) | 2025 | embodied foundation | dataset open |
| World Labs | Marble | 2026 | large world model | proprietary |
The big picture
Proprietary & integrated (Google, Figure, Tesla, 1X): Companies that build their own robots or own a platform mostly keep the model closed and tune it tightly to their hardware. Gemini Robotics adds a separate reasoning model (ER) that “thinks before acting”.
Open & cross-platform (NVIDIA, Hugging Face): NVIDIA positions GR00T and Cosmos as open building blocks for the whole industry — flanked by its chip and simulation business. Hugging Face aims at democratisation with LeRobot/SmolVLA: small, open VLAs that run on consumer hardware.
Pure model startups (Physical Intelligence, Skild AI): They don't sell a robot body but the “brain” — a model that controls as many robot bodies as possible. Physical Intelligence's π0 is a widely noted, partly open VLA; Skild AI pursues an “omni-bodied” model.
The China bloc (AgiBot, ByteDance, Fourier, Unitree …): Chinese players combine aggressive hardware mass-production with their own embodied models and partly open datasets (AgiBot World). ByteDance's GR-3 and AgiBot's GO-1 are among the most visible.
The shared bottleneck
They all share one problem: too little robot data. Answers range from massive teleoperation and shared datasets (Open X-Embodiment) to synthetic data from world models (NVIDIA's “GR00T-Dreams”/Cosmos). This is where it will be decided whose model becomes reliable enough for everyday use first.
A more detailed, interlinked version (incl. individual model pages) is currently available in German: Foundation-Model-Landschaft.