Stack Gems
Menu
← Catalog

DeepEval

Apache-2.0 pytest-style LLM/agent evals; Confident AI is the paid cloud

AI Agents / Memory & Evals · Apache-2.0 · 18.1k

What's Good

Apache-2.0. 50+ metrics (G-Eval, hallucination, relevancy, task completion, multi-turn). Local-first unit/regression tests, synthetic data, CI. No paywalled metrics in the OSS package.

The Catch

OSS is the runner — collaboration/dashboards/monitoring are Confident AI (Free: 2 seats/1 project/5 runs/wk; Starter $200/mo; Team $2k/mo). LLM-as-judge burns provider tokens. Vs Ragas: broader agent/chat coverage, not RAG-only.

Verdict

Apache-2.0 LLM unit tests. Free locally; Confident AI starts free then $200.

Embed

Reviewed on Stack Gems
[![Reviewed on Stack Gems](https://stackgems.com/badge/deepeval.svg)](https://stackgems.com/gems/deepeval)