NVIDIA NIM
Prebuilt, optimised inference microservices for NVIDIA GPUs.
Containerised model servers with TensorRT-LLM optimisations and OpenAI-compatible APIs, runnable on any NVIDIA GPU or tried for free on build.nvidia.com.
- Vendor
- NVIDIA
- Category
- Inference & model hosting
- Pricing
- Enterprise
- Open source
- No
- Platforms
- Cloud, On-prem, Kubernetes
- Launched
- 2024
- Website
- build.nvidia.com
Features
- Optimised containers
- OpenAI-compatible API
- Free hosted trials
Best for
- Enterprise
- Local & on-device
Works with
- Nemotron 3 Ultra — NVIDIA. NVIDIA’s largest open model: hybrid Mamba-Transformer MoE, 55B active.
- Cosmos 3 — NVIDIA. NVIDIA’s open omni-modal world foundation model for physical AI.
More inference & model hosting
- Ollama — Ollama. Run open models locally with one command.
- OpenRouter — OpenRouter. One API for hundreds of models.
- vLLM — vLLM project. High-throughput LLM serving engine.
- llama.cpp — ggml. LLM inference in C/C++ on any hardware.
- Amazon Bedrock — AWS. Foundation models on AWS.
- Vertex AI — Google Cloud. Google Cloud’s AI platform.
- LM Studio — Element Labs. Desktop app to discover and run local LLMs.
- Azure AI Foundry — Microsoft. Build and run AI apps on Azure.