# DeepEval

> Apache-2.0 pytest-style LLM/agent evals; Confident AI is the paid cloud

**Category:** AI Agents  
**Pricing:** open-source  
**URL:** https://deepeval.com  
**GitHub:** https://github.com/confident-ai/deepeval  
**Added:** 2026-09-04

---

## The Hook

DeepEval is pytest for LLM apps when Promptfoo is YAML red-team and Langfuse is the tracing cousin.

## What's Good

Apache-2.0. 50+ metrics (G-Eval, hallucination, relevancy, task completion, multi-turn). Local-first unit/regression tests, synthetic data, CI. No paywalled metrics in the OSS package.

## The Catch

OSS is the runner — collaboration/dashboards/monitoring are Confident AI (Free: 2 seats/1 project/5 runs/wk; Starter $200/mo; Team $2k/mo). LLM-as-judge burns provider tokens. Vs Ragas: broader agent/chat coverage, not RAG-only.

## Verdict

Apache-2.0 LLM unit tests. Free locally; Confident AI starts free then $200.
