Gemma 4 31B
Open-weight dense multimodal model from Google DeepMind.
Gemma 4 31B Instruct is Google DeepMind’s 30.7B dense multimodal model supporting text, image and video input with a 256K-token context window, configurable thinking and native function calling.
- Provider
- Type
- Language model
- Released
- Apr 2, 2026
- Context window
- 262K tokens
- Max output
- 16K tokens
- Price
- $0.090 input / $0.34 output per 1M tokens
- Input
- text, image, video
- Output
- text
- Open weights
- Yes
- License
- Gemma Terms of Use
- Knowledge cutoff
- Jan 1, 2025
Best for
- Local & on-device
- Vision
- Chat & assistants
- Extraction & classification
Strengths
- Single-GPU deployable
- Multimodal open model
- Free tiers widely available
More from Google
- Gemini 4 Argon — Google. Google’s next-generation frontier model, currently restricted to vetted cyber defenders.
- Gemini 3.8 Flash — Google. Google’s most intelligent Flash model — fully multimodal, fast and affordable.
- Gemini 3.5 Flash Lite — Google. High-efficiency multimodal model for focused subagent tasks.
- Nano Banana 2.1 — Google. Google’s latest image generation and editing model.
- Gemma 4 26B A4B — Google. Open MoE Gemma: ~31B-class quality with only 3.8B active parameters.
- Gemini Embedding 2 — Google. Google’s first multimodal embedding model.
- Gemini Omni Flash 1.1 — Google. Google’s fast, low-cost video generation model.
- Veo 3.1 — Google. Google’s Veo video model with native audio; powers Flow.
Tools that use it
- Ollama — Ollama. Run open models locally with one command.