Illustration by tuput
English
Google DeepMind published success rates for Gemini Robotics 2 on 30 July 2026, from 32% on a dustpan task to 92% on unscrewing a bulb, without trial counts. Of the models tuput checked, NVIDIA, Xiaomi and Unitree have put weights online, and the headline numbers are the labs' own or come without a stated tester.
Google DeepMind’s 30 July 2026 release of Gemini Robotics 2 came with charts of success rates for a humanoid robot: 92% for unscrewing a light bulb, 36% for screwing one in, 32% for using a dustpan. DeepMind’s page does not say how many trials sit behind any of them.
A robot foundation model is one trained network that takes camera images and a written instruction and outputs motor commands, for many tasks and many robot bodies. Eight labs have put out such a model or an update to one in 2026. For every headline number tuput could trace, the lab ran its own tests, or the page does not say who did. Which weights, meaning the trained numbers a developer can download, are public varies a lot.
Google’s humanoid numbers, and who can use them
Gemini Robotics 2 is a vision-language-action model, or VLA: it reads images and language and outputs actions. Carolina Parada’s DeepMind post says it can control full humanoids “from feet to fingertips” as well as two-armed robots. Two companions shipped with it. ER 2 is a planner that splits a job into steps and calls the VLA. On-Device 2 is a smaller version that runs on the robot’s own computer.
On Apptronik’s Apollo 2 humanoid with Inspire hands, DeepMind reports picking up an object at 68.4% from a table, 76.3% from a shelf and 45.7% from the floor. With Sharpa hands it reports 92% for unscrewing a bulb, 44% for tying a trash bag, 40% for a ziplock bag, 36% for screwing a bulb in and 32% for the dustpan. On Franka Duo arms with grippers the scores are 74.2% for pick and place, 78.9% for tool kitting and 89.6% for precise insertion. Google says multi-finger dexterity remains challenging and that its robots still need to move faster. The hand problem behind those low scores is explained in our piece on how a robot hand learns to grasp.
Access is narrow. The VLA and On-Device 2 go to early-access partners, and the model page cites more than 100 trusted testers. ER 2 can be tried in Google AI Studio, where it is labelled a preview.
ER 2 has its own figures from Google’s evaluations: 57.4% accuracy at progress classification, which sorts a video frame into five bands of how far a task has got, and 91.3% at moment finding, which picks the frame where a key event happens, with a mean error of 0.96 seconds. Google also published a safety benchmark called ASIMOV-Agentic, and says in tests ER 2 halted a humanoid when a person came close.
The On-Device 2 model card is blunter than the blog. It lists limited generalisation to tasks outside the training distribution and limited ability to control robots with many joints. Its safety testing covered standing two-armed work only, so whole-body risks are outside its scope. DeepMind says a new two-armed robot can be adapted with fewer than 200 examples collected over a few hours.
Physical Intelligence: pi0.7 and the trouble with “unseen”
Physical Intelligence posted its pi0.7 paper to arXiv on 16 April 2026. The model is about 5 billion parameters: a 4-billion-parameter language-and-vision backbone started from Google’s Gemma 3 plus an 860-million-parameter action module.
The headline test is shirt folding on a UR5e arm for which no folding data had been collected. The authors report 80% success and 85.6% task progress for pi0.7, against 80.6% success and 90.9% progress for ten experienced human teleoperators. The authors ran the study. Across unseen tasks or new task-and-robot pairings they report roughly 60% to 80% success, while tasks already seen often topped 90%. Laundry folding is a common stress test, and why it stays hard is covered in our piece on robots that ace chess and flunk laundry.
The authors write that it is hard to determine what counts as unseen given the size of their dataset. The Decoder reported one example: a robot loading a sweet potato into an air fryer needed step-by-step coaching, and the training data held only two episodes of a robot closing an air fryer.
The paper does not say whether weights will be released. The openpi repository page tuput read lists pi0, pi0-FAST and pi0.5 under an Apache 2.0 licence, and no pi0.7.
Skild’s S1: 66% on unseen tasks, 86% for plain training
Skild AI’s blog introduces S1 as a model that watches one video of a task and then does it, with no change to its weights. Tasks run up to ten minutes. On unseen tasks after 100,000 hours of pretraining, Skild reports 66% success against 9% for a language-prompted VLA trained on the same data, architecture and compute. The metric is per-step success averaged over long tasks. Skild says two internal benchmark suites were used and does not give the number of tasks or trials. Human operators could step in to recover failed rollouts, mainly for the baseline, which otherwise scored zero on long unseen tasks.
Skild’s own figures also show the limit. A VLA post-trained on 2,000 demonstrations reached 86% on the new task Skild tested, above the 66% from one video. Skild estimates one video is worth about 380 demonstrations and calls that figure interpolated. The post offers an early-access sign-up, not a download.
The Robot Report quotes CEO Deepak Pathak as saying S1 has not yet been scaled to humanoids and is not ready for homes. The post does not name the robots it ran on.
Open weights from NVIDIA, Xiaomi and Unitree
NVIDIA’s GR00T N1.7 has about 3 billion parameters. Its 16 March press release called it available in early access with commercial licensing. A 7 July technical post says it is released under Apache 2.0, with weights on Hugging Face and training code on GitHub, after pretraining on about 32,000 hours of real and first-person human data and 8,000 hours of simulation. Its gains over N1.6 are given only as relative: up 61% on one DROID test and 5% on SimplerEnv Bridge. The Hugging Face model card for the same model names the NVIDIA Open Model License, so the two NVIDIA pages disagree on the licence.
NVIDIA’s March release also previewed GR00T N2, which is based on DreamZero research. NVIDIA’s engineers describe DreamZero as a world-action model, one network that generates predicted video and actions together. NVIDIA says it succeeds at new tasks in new environments more than twice as often as leading VLAs and ranks first on the MolmoSpaces and RoboArena leaderboards. The release gives no data behind that claim, and N2 is slated for the end of 2026.
Xiaomi’s XR-1 publishes its tests alongside its weights. Its README, under Apache 2.0, describes pretraining on over 100,000 hours of demonstration data recorded without a robot body, then post-training on over 10,000 hours across robot types. Checkpoints, post-training code and evaluation code are public, and the technical report went onto arXiv on 16 July. Xiaomi reports 57.4% on the RoboCasa365 simulation benchmark against 46.6% for the second-best model. On four real-robot tasks with under 10 hours of data per task it reports 75% overall against 40% for pi0.5, and 85% against 53% with under 40 hours. Laundry loading reached 100% against 50% in the 40-hour setting. The real-robot comparison is Xiaomi’s own, and the README says XR-1 ranks first on the RoboCasa365 and RoboDojo leaderboards as of 15 July.
Unitree announced on 10 September that it was open-sourcing UnifoLM-WLA-1.0, a model it says drives both tabletop and whole-body work. Its Hugging Face page lists a base checkpoint under Apache 2.0, but the model card was empty when tuput read it, so no task figures were available. Unitree’s humanoid hardware, and what it shares with robot vacuums, is covered in our comparison of the two.
Figure and AgiBot
Figure’s Helix 02, introduced on 27 January 2026, runs a four-minute dishwasher job with 61 actions and no resets, according to Figure. Its balance layer is a 10-million-parameter network trained on over 1,000 hours of human motion data. The page gives no success rates and ties the finger-level tasks to Figure 03 hardware, with fingertip sensors that detect forces as small as three grams. Figure calls the results early. What humanoids of this kind have done in paid pilots is in our report on humanoids clocking in.
AgiBot’s GO-2, announced on 9 April, is credited by The Robot Report with 98.5% on the LIBERO simulation suite, 86.6% on LIBERO-Plus and 82.9% in real-world transfer after simulation-only training. All three are AgiBot’s figures, and the article does not say whether GO-2’s weights are public.
Why the percentages cannot be lined up
A 98.5% on LIBERO is a score on a fixed simulation suite. A 32% on a dustpan is a real humanoid with 22-joint hands. A 66% from Skild averages per-step success across long tasks. No two of these share a task, a robot or a scoring rule, and the labs that compare themselves name pi0.5 or earlier versions of their own models as the opponent.
NVIDIA’s March release says GR00T N2 is slated to be available by the end of 2026.
Sources & further reading
- Google DeepMind: Gemini Robotics 2 brings whole body intelligence to robots (30 July 2026)
- Google DeepMind: Gemini Robotics 2 overview and trusted tester programme
- Google DeepMind: Gemini Robotics On-Device 2 model card
- Google: Introducing Gemini Robotics ER 2 (30 July 2026)
- Physical Intelligence: pi 0.7, a Steerable Generalist Robotic Foundation Model with Emergent Capabilities (arXiv 2604.15483, 16 April 2026)
- The Decoder: Physical Intelligence shows robot model with LLM-like generalization, flaws included
- Physical Intelligence: openpi repository (GitHub)
- Skild AI: Introducing S1, In-Context Learning for Robotics (August 2026)
- The Robot Report: Skild AI unveils S1 flagship robot foundation model (31 August 2026)
- NVIDIA: NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (16 March 2026)
- NVIDIA Technical Blog: Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T (7 July 2026)
- NVIDIA Technical Blog: Pretrained to Imagine, Fine-Tuned to Act, the Rise of World-Action Models (15 June 2026)
- Hugging Face: NVIDIA GR00T-N1.7-3B model card
- GitHub: Xiaomi-Robotics-1 (XR-1) README
- Hugging Face: Unitree Robotics models, UnifoLM-WLA-1.0-Base
- PANews: Unitree Robotics open-sources UnifoLM-WLA-1.0 embodied foundation model
- Figure: Introducing Helix 02, Full-Body Autonomy (27 January 2026)
- The Robot Report: AGIBOT releases GO-2 foundation model for embodied AI (9 April 2026)
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.