{"data":{"slug":"deepeval","name":"DeepEval","tagline":"Apache-2.0 pytest-style LLM/agent evals; Confident AI is the paid cloud","category":"agents","categoryName":"AI Agents","url":"https://deepeval.com","github":"https://github.com/confident-ai/deepeval","pricing":"open-source","hook":"DeepEval is pytest for LLM apps when Promptfoo is YAML red-team and Langfuse is the tracing cousin.","whatsGood":"Apache-2.0. 50+ metrics (G-Eval, hallucination, relevancy, task completion, multi-turn). Local-first unit/regression tests, synthetic data, CI. No paywalled metrics in the OSS package.","theCatch":"OSS is the runner — collaboration/dashboards/monitoring are Confident AI (Free: 2 seats/1 project/5 runs/wk; Starter $200/mo; Team $2k/mo). LLM-as-judge burns provider tokens. Vs Ragas: broader agent/chat coverage, not RAG-only.","verdict":"Apache-2.0 LLM unit tests. Free locally; Confident AI starts free then $200.","addedAt":"2026-09-04","featured":false,"domain":"AI Agents","subSpecialty":"Memory & Evals","capabilities":["evals","pytest","agents"],"surfaces":["Self-host"],"ecosystem":["Python"],"licenseModel":"Apache-2.0","githubStars":18102},"related":[{"slug":"claude-code","name":"Claude Code","tagline":"Terminal-first coding agent with deep repository context"},{"slug":"cursor","name":"Cursor","tagline":"The AI-first code editor that understands your codebase"},{"slug":"devin","name":"Devin","tagline":"Autonomous AI engineer for defined engineering tasks"},{"slug":"stagehand","name":"Stagehand","tagline":"AI-native browser automation that mixes code and natural language"},{"slug":"browser-use","name":"Browser Use","tagline":"Open-source framework for LLM-driven browser agents"}],"links":{"self":"/api/gems/deepeval","html":"/gems/deepeval","markdown":"/gems/deepeval.md","category":"/api/gems?category=agents","categoryPage":"/stacks/agents"}}