Isaac GR00T N1.7

Open 3B cross-embodiment VLA for humanoids and manipulators.

GR00T N1.7 (early access) is NVIDIA’s open vision-language-action model. A Cosmos-Reason2 VLM backbone reads camera frames and a language instruction; a flow-matching diffusion-transformer action head outputs continuous motor commands for the robot’s embodiment. Post-train it on real or synthetic data for a specific robot.

Provider
NVIDIA
Type
Robotics (VLA)
Released
Apr 17, 2026
Price
Open weights — self-host
Input
text, image
Output
action
Open weights
Yes
License
NVIDIA Open Model License

Best for

Strengths

More from NVIDIA

Tools that use it

Sources