Gemini Robotics 2
Whole-body humanoid control from vision and language.
Gemini Robotics 2 is Google DeepMind’s vision-language-action model family for robots from bi-arm systems to full humanoids, adding whole-body control (walking, crouching, manipulation) and multi-robot coordination. The companion Gemini Robotics ER 2 embodied-reasoning model is in preview in Google AI Studio; the VLA and On-Device 2 models go to early-access partners.
- Provider
- Type
- Robotics (VLA)
- Released
- Jul 30, 2026
- Input
- text, image, video
- Output
- action, text
- Open weights
- No
Best for
- Agents
- Vision
- Multimodal
Strengths
- Whole-body humanoid control
- ER 2 reasoning model in AI Studio
- On-device variant
More from Google
- Gemini 4 Argon — Google. Google’s next-generation frontier model, currently restricted to vetted cyber defenders.
- Gemini 3.8 Flash — Google. Google’s most intelligent Flash model — fully multimodal, fast and affordable.
- Gemini 3.5 Flash Lite — Google. High-efficiency multimodal model for focused subagent tasks.
- Nano Banana 2.1 — Google. Google’s latest image generation and editing model.
- Gemma 4 31B — Google. Open-weight dense multimodal model from Google DeepMind.
- Gemma 4 26B A4B — Google. Open MoE Gemma: ~31B-class quality with only 3.8B active parameters.
- Gemini Embedding 2 — Google. Google’s first multimodal embedding model.
- Gemini Omni Flash 1.1 — Google. Google’s fast, low-cost video generation model.