Mercury 2.5
Diffusion LLM generating ~750 tokens/second — the fastest in the index.
Mercury 2.5 is a diffusion LLM (dLLM). Instead of generating tokens sequentially, it produces and refines multiple tokens in parallel, achieving extremely high throughput.
- Provider
- Inception
- Type
- Language model
- Released
- Sep 8, 2026
- Context window
- 260K tokens
- Max output
- 66K tokens
- Price
- $0.040 input / $0.15 output per 1M tokens
- Input
- text
- Output
- text
- Open weights
- No
Best for
- Real-time / low latency
- High volume / low cost
- Coding
Strengths
- ~750 tok/s output
- Very cheap