TRL
Post-training library: SFT, DPO, GRPO and RL for LLMs.
Hugging Face’s Transformer Reinforcement Learning library with trainers for supervised fine-tuning, preference optimisation and RL with verifiable rewards — the toolkit behind many open reasoning models.
- Vendor
- Hugging Face
- Category
- Data & ML platforms
- Pricing
- Open source
- Open source
- Yes
- License
- Apache 2.0
- Platforms
- Python
- Website
- huggingface.co/docs/trl
Features
- SFT / DPO / GRPO trainers
- PEFT integration
- vLLM generation
Best for
- Research
- Reasoning
More data & ml platforms
- Hugging Face — Hugging Face. The home of open models and datasets.
- Weights & Biases — Weights & Biases. The AI developer platform.
- Unsloth — Unsloth AI. Fine-tune LLMs faster with less memory.
- Scale AI — Scale AI. Data labeling and evaluation for frontier AI.
- Label Studio — HumanSignal. Open-source data labelling for text, images, audio and LLM outputs.
- Axolotl — Axolotl AI. Config-driven fine-tuning for open LLMs.