Gemma 4 26B A4B
Open MoE Gemma: ~31B-class quality with only 3.8B active parameters.
Gemma 4 26B A4B is an instruction-tuned mixture-of-experts model from Google DeepMind. Of its 25.2B total parameters only 3.8B activate per token, giving near-31B quality at much lower inference cost. Accepts text, images and video.
- Provider
- Type
- Language model
- Released
- Apr 3, 2026
- Context window
- 262K tokens
- Max output
- 236K tokens
- Price
- $0.068 input / $0.23 output per 1M tokens
- Input
- text, image, video
- Output
- text
- Open weights
- Yes
- License
- Gemma Terms of Use
- Knowledge cutoff
- Jan 1, 2025
Best for
- Local & on-device
- Vision
- Chat & assistants
- High volume / low cost
Strengths
- 3.8B active parameters
- Open weights
- Very low hosted price
More from Google
- Gemini 4 Argon — Google. Google’s next-generation frontier model, currently restricted to vetted cyber defenders.
- Gemini 3.8 Flash — Google. Google’s most intelligent Flash model — fully multimodal, fast and affordable.
- Gemini 3.5 Flash Lite — Google. High-efficiency multimodal model for focused subagent tasks.
- Nano Banana 2.1 — Google. Google’s latest image generation and editing model.
- Gemma 4 31B — Google. Open-weight dense multimodal model from Google DeepMind.
- Gemini Embedding 2 — Google. Google’s first multimodal embedding model.
- Gemini Omni Flash 1.1 — Google. Google’s fast, low-cost video generation model.
- Veo 3.1 — Google. Google’s Veo video model with native audio; powers Flow.