DeepEval
Pytest-style unit testing for LLM outputs.
Open-source framework with 30+ LLM-as-judge metrics (G-Eval, hallucination, RAG, agent metrics) that run like unit tests in CI, plus the Confident AI platform.
- Vendor
- Confident AI
- Category
- Observability & evals
- Pricing
- Open source
- Open source
- Yes
- License
- Apache 2.0
- Platforms
- Python, CI
- Launched
- 2023
- Website
- deepeval.com
Features
- Pytest integration
- G-Eval & RAG metrics
- Red-teaming module
Best for
- Coding
- Enterprise
More observability & evals
- LangSmith — LangChain. Tracing, evals and monitoring for LLM apps.
- Langfuse — Langfuse. Open-source LLM engineering platform.
- Braintrust — Braintrust. The evals platform for AI products.
- promptfoo — promptfoo. Test and red-team your LLM apps.
- Arize Phoenix — Arize AI. Open-source AI observability and evaluation.
- W&B Weave — Weights & Biases. Track and evaluate LLM applications.
- Helicone — Helicone. LLM observability via a one-line proxy.
- Ragas — Exploding Gradients. Evaluation metrics for RAG pipelines and agents.