promptfoo
Test and red-team your LLM apps.
CLI and library for prompt testing, model comparison and automated red-teaming in CI.
- Vendor
- promptfoo
- Category
- Observability & evals
- Pricing
- Open source
- Open source
- Yes
- License
- MIT
- Platforms
- CLI, CI
- Launched
- 2023
- Website
- promptfoo.dev
Features
- Declarative tests
- Red teaming
- Model comparison
Best for
- Enterprise
More observability & evals
- LangSmith — LangChain. Tracing, evals and monitoring for LLM apps.
- Langfuse — Langfuse. Open-source LLM engineering platform.
- Braintrust — Braintrust. The evals platform for AI products.
- Arize Phoenix — Arize AI. Open-source AI observability and evaluation.
- W&B Weave — Weights & Biases. Track and evaluate LLM applications.
- Helicone — Helicone. LLM observability via a one-line proxy.
- Ragas — Exploding Gradients. Evaluation metrics for RAG pipelines and agents.
- DeepEval — Confident AI. Pytest-style unit testing for LLM outputs.