DeepEval
Apache-2.0 pytest-style LLM/agent evals; Confident AI is the paid cloud
AI Agents / Memory & Evals · Apache-2.0 · 18.1k
What's Good
Apache-2.0. 50+ metrics (G-Eval, hallucination, relevancy, task completion, multi-turn). Local-first unit/regression tests, synthetic data, CI. No paywalled metrics in the OSS package.
The Catch
OSS is the runner — collaboration/dashboards/monitoring are Confident AI (Free: 2 seats/1 project/5 runs/wk; Starter $200/mo; Team $2k/mo). LLM-as-judge burns provider tokens. Vs Ragas: broader agent/chat coverage, not RAG-only.
Verdict
Apache-2.0 LLM unit tests. Free locally; Confident AI starts free then $200.
Embed
[](https://stackgems.com/gems/deepeval)