DeepEval

Pytest-style unit testing for LLM outputs.

Open-source framework with 30+ LLM-as-judge metrics (G-Eval, hallucination, RAG, agent metrics) that run like unit tests in CI, plus the Confident AI platform.

Vendor
Confident AI
Category
Observability & evals
Pricing
Open source
Open source
Yes
License
Apache 2.0
Platforms
Python, CI
Launched
2023
Website
deepeval.com

Features

Best for

More observability & evals

Links