Mercury 2.5

Diffusion LLM generating ~750 tokens/second — the fastest in the index.

Mercury 2.5 is a diffusion LLM (dLLM). Instead of generating tokens sequentially, it produces and refines multiple tokens in parallel, achieving extremely high throughput.

Provider
Inception
Type
Language model
Released
Sep 8, 2026
Context window
260K tokens
Max output
66K tokens
Price
$0.040 input / $0.15 output per 1M tokens
Input
text
Output
text
Open weights
No

Best for

Strengths

Sources