Gemini 3.5 Transcribe
Google’s dedicated transcription model.
Gemini 3.5 Transcribe is Google’s dedicated speech-to-text model in the Gemini family.
- Provider
- Type
- Speech-to-text
- Released
- Aug 26, 2026
- Price
- ≈$5 per 1,000 minutes
- Input
- audio
- Output
- text
- Open weights
- No
Best for
- Extraction & classification
- Multimodal
Strengths
- ≈2.6% WER
More from Google
- Gemini 4 Argon — Google. Google’s next-generation frontier model, currently restricted to vetted cyber defenders.
- Gemini 3.8 Flash — Google. Google’s most intelligent Flash model — fully multimodal, fast and affordable.
- Gemini 3.5 Flash Lite — Google. High-efficiency multimodal model for focused subagent tasks.
- Nano Banana 2.1 — Google. Google’s latest image generation and editing model.
- Gemma 4 31B — Google. Open-weight dense multimodal model from Google DeepMind.
- Gemma 4 26B A4B — Google. Open MoE Gemma: ~31B-class quality with only 3.8B active parameters.
- Gemini Embedding 2 — Google. Google’s first multimodal embedding model.
- Gemini Omni Flash 1.1 — Google. Google’s fast, low-cost video generation model.