AI models
Every major AI model in one sortable table: pricing per 1M tokens, context window, benchmarks, speed, modalities and open weights.
- Claude Opus 5.5 — Anthropic. Anthropic’s flagship for demanding reasoning, coding and long-horizon agents.
- Claude Sonnet 5.5 — Anthropic. Near-flagship intelligence at a mid-tier price; the everyday workhorse for builders.
- Claude Haiku 5.5 — Anthropic. Small, fast and inexpensive — built for subagents, summarization and browser use.
- Claude Fable 5.1 — Anthropic. Premium model for long-running agentic workflows, refactors and visual front-end work.
- Claude Opus 5 — Anthropic. Previous-generation Opus; still a strong reasoning and coding model.
- GPT-6 Astra — OpenAI. OpenAI’s flagship for deep research, science and end-to-end engineering.
- GPT-6.1 Sol — OpenAI. Mid-tier GPT-6 model with excellent cost per task for agentic coding.
- GPT-6 Luna — OpenAI. Fast, cost-efficient GPT-6 tier for chat, classification and light agents.
- GPT-5.6 Terra — OpenAI. Balanced GPT-5.6 model for everyday coding and reasoning.
- gpt-oss-120b — OpenAI. OpenAI’s Apache-2.0 open-weight reasoning model; runs on a single 80GB GPU.
- gpt-oss-20b — OpenAI. Compact open-weight reasoning model that runs on consumer hardware.
- GPT-5.4 Image 2 — OpenAI. OpenAI’s native image generation and editing model.
- Gemini 4 Argon — Google. Google’s next-generation frontier model, currently restricted to vetted cyber defenders.
- Gemini 3.8 Flash — Google. Google’s most intelligent Flash model — fully multimodal, fast and affordable.
- Gemini 3.5 Flash Lite — Google. High-efficiency multimodal model for focused subagent tasks.
- Nano Banana 2.1 — Google. Google’s latest image generation and editing model.
- Gemma 4 31B — Google. Open-weight dense multimodal model from Google DeepMind.
- Grok 4.7 — xAI (SpaceXAI). xAI’s flagship for long-running software engineering and self-verification.
- Grok 4.20 — xAI (SpaceXAI). Fast reasoning model with a 2M-token context window and strict prompt adherence.
- Llama 4 Maverick — Meta. Meta’s 128-expert open multimodal MoE model.
- Llama 4 Scout — Meta. Efficient open model with an industry-leading 10M-token native context.
- Qwen3.8 Max — Alibaba (Qwen). Alibaba’s 2.4T-parameter proprietary flagship with video understanding.
- Qwen3.8 2.4T A95B — Alibaba (Qwen). The open-weight variant of Qwen3.8 Max — the largest open model available.
- Qwen3.8 Flash — Alibaba (Qwen). Open multimodal reasoning model for coding, desktop interaction and long video.
- Qwen3.6 35B A3B — Alibaba (Qwen). Tiny-active-parameter open MoE — a local-inference favorite.
- DeepSeek V4.1 Flash — DeepSeek. Open-weight, very fast MoE built on DeepSeek’s new encoder-decoder architecture.
- DeepSeek V4 Pro — DeepSeek. DeepSeek’s large open MoE, GA release (0813).
- Kimi K3 — Moonshot AI. A 2.8T-parameter open multimodal reasoning model — joint-top open model on GPQA.
- GLM-5.3 — Z.ai. Open reasoning model for complex software engineering and long-horizon agents.
- GLM-5.3 Flash — Z.ai. Native multimodal open model with hybrid sparse/linear attention for long context.
- MiMo-V2.6-Pro — Xiaomi. Xiaomi’s 1T+ open flagship — the best intelligence per dollar in the index.
- MiMo-V2.6-Flash — Xiaomi. Open 309B MoE with omni-modal input at near-zero cost.
- MiniMax M3 — MiniMax. Open multimodal model built for long-horizon agentic work.
- Step 5 Preview — StepFun. StepFun’s new agentic flagship — 600B MoE with video input.
- Mistral Large 4 — Mistral AI. Mistral’s 1T-parameter multimodal MoE flagship (preview; weights promised).
- Mistral Medium 3.5 — Mistral AI. Fast enterprise-grade multimodal model.
- Mistral Small 4 — Mistral AI. Apache-licensed open model unifying reasoning, vision and coding.
- Command A+ — Cohere. Cohere’s enterprise agent model with strict tool schemas and very low latency.
- Mercury 2.5 — Inception. Diffusion LLM generating ~750 tokens/second — the fastest in the index.
- Nemotron 3 Super — NVIDIA. Open hybrid Mamba-Transformer MoE for multi-agent applications.
- Nemotron 3.5 Lightning — NVIDIA. Tiny open 30B-A3B model with the lowest first-token latency in the index.
- Granite 4.2 8B — IBM. Apache-licensed small reasoning model for governed enterprise use.
- Nova 2 Lite — Amazon. Amazon’s low-cost multimodal model, native to AWS Bedrock.
- Muse Spark 1.3 — Meta. Meta’s multimodal reasoning model for long-running agents and coding.
- Gemma 4 26B A4B — Google. Open MoE Gemma: ~31B-class quality with only 3.8B active parameters.
- Nemotron 3 Ultra — NVIDIA. NVIDIA’s largest open model: hybrid Mamba-Transformer MoE, 55B active.
- Granite 4.0 H Micro — IBM. Tiny hybrid Mamba-2/Transformer model for cheap, long-context enterprise tasks.
- LFM2.5-2.6B — Liquid AI. Compact on-device reasoning model on Liquid’s hybrid architecture.
- Ling 3.1 Flash — inclusionAI (Ant Group). Fast hybrid-reasoning MoE (25B active / 560B) scoring 41 on the Intelligence Index.
- Seed 2.1 Turbo — ByteDance Seed. ByteDance’s multimodal model for coding and long-horizon agents.
- LongCat 2.0 — Meituan (LongCat). Open 1.6T-parameter MoE with a 1M-token context for repo-scale coding.
- Jev — TypeSafe AI. The first System One model: unstructured state in, typed probabilistic decisions out — in milliseconds.
- voyage-4-large — Voyage AI (MongoDB). Voyage’s highest-quality general and multilingual retrieval embedding.
- voyage-code-4 — Voyage AI (MongoDB). Code-retrieval embedding built for coding agents.
- voyage-multimodal-3.5 — Voyage AI (MongoDB). Embeds text, images and interleaved content into one space.
- Gemini Embedding 2 — Google. Google’s first multimodal embedding model.
- text-embedding-3-large — OpenAI. OpenAI’s most capable embedding; the default in countless RAG stacks.
- text-embedding-3-small — OpenAI. Cheap, fast OpenAI embedding for high-volume indexing.
- Qwen3-Embedding-8B — Alibaba (Qwen). Leading open-weight multilingual embedding model.
- Nemotron 3 Embed 1B — NVIDIA. Small open embedding model for high-throughput enterprise retrieval.
- pplx-embed-v1 4B — Perplexity. Perplexity’s web-scale retrieval embedding.
- LFM2.5-Embedding-350M — Liquid AI. Tiny open embedding model for on-device semantic search.
- GPT Image 2.5 — OpenAI. Top-ranked model in the Artificial Analysis image arena.
- Grok Imagine Image 2.0 — xAI (SpaceXAI). SpaceXAI’s image model, top-5 in the image arena.
- MAI-Image-2.6 — Microsoft AI. Microsoft AI’s in-house image model.
- FLUX 3 Image — Black Forest Labs. Black Forest Labs’ newest FLUX image model.
- Muse Image — Meta. Meta’s low-cost image model.
- Seedream 5.0 Pro — ByteDance Seed. ByteDance Seed’s flagship image model.
- Qwen-Image-3.0-Pro — Alibaba (Qwen). Alibaba’s flagship image model with strong text rendering.
- Qwen-Image-2.1 — Alibaba (Qwen). Highest-ranked open-weight image model in the arena.
- Wan 3.0 — Alibaba (Qwen). #1 in the Artificial Analysis text-to-video arena.
- Dreamina Seedance 2.5 — ByteDance Seed. ByteDance’s top video model.
- MiniMax H3 — MiniMax. Highest-ranked open-weight video model — and one of the cheapest.
- Gemini Omni Flash 1.1 — Google. Google’s fast, low-cost video generation model.
- Veo 3.1 — Google. Google’s Veo video model with native audio; powers Flow.
- Kling 3.0 — Kuaishou (Kling). Kuaishou’s 1080p video model, popular with creators.
- LTX-2.5 — Lightricks. Open-weight video model you can run locally.
- Eleven v4 — ElevenLabs. The most preferred voice in the Artificial Analysis speech arena.
- Qwen-Audio-3.1-TTS-Plus — Alibaba (Qwen). Top-3 voice quality at a quarter of the leader’s price.
- Sonic 3.6 — Cartesia. Low-latency streaming TTS for voice agents.
- Gemini 3.8 Flash TTS — Google. Google’s controllable TTS with a cheaper Flash-Lite tier.
- Inworld Realtime TTS-2 — Inworld. Real-time TTS for games and interactive characters.
- StepAudio 3 ASR — StepFun. Lowest word error rate on the Artificial Analysis transcription benchmark.
- MAI-Transcribe-2 — Microsoft AI. Top-3 accuracy at one of the lowest per-minute prices.
- Scribe v2 — ElevenLabs. ElevenLabs’ highly accurate transcription model.
- Grok Voice Transcribe 2.0 — xAI (SpaceXAI). SpaceXAI’s low-cost, high-accuracy transcription.
- Gemini 3.5 Transcribe — Google. Google’s dedicated transcription model.
- Universal-3 Pro — AssemblyAI. AssemblyAI’s flagship speech-to-text model.
- Voxtral Small — Mistral AI. Best open-weight transcription model — also understands audio.
- Lyria 3 Pro — Google. Google’s full-length song generation model.
- V-JEPA 2 — Meta. Meta’s open JEPA world model for video understanding, prediction and robot planning.
- Genie 3 — Google. Real-time interactive world model that generates explorable 3D environments.
- Cosmos 3 — NVIDIA. NVIDIA’s open omni-modal world foundation model for physical AI.
- Gemini Robotics 2 — Google. Whole-body humanoid control from vision and language.
- Isaac GR00T N1.7 — NVIDIA. Open 3B cross-embodiment VLA for humanoids and manipulators.
- π0.5 — Physical Intelligence. Open VLA from Physical Intelligence with strong open-world generalisation.
- Llama Guard 4 12B — Meta. Open multimodal safety classifier for prompts and responses.
- gpt-oss-safeguard-20b — OpenAI. Open safety-reasoning model that applies your own written policy.
- Nemotron 3.5 Content Safety — NVIDIA. Compact 4B multimodal guardrail model for LLM and VLM traffic.