← Catalog
Overlap
Memory & Evals for Python
Memory & Evals tools that run on Python. The Catch for each, not a winner list.
2 tools · the Catch, not a ranking
- DeepEvalApache-2.0 pytest-style LLM/agent evals; Confident AI is the paid cloud
OSS is the runner — collaboration/dashboards/monitoring are Confident AI (Free: 2 seats/1 project/5 runs/wk; Starter $200/mo; Team $2k/mo). LLM-as-judge burns provider tokens. Vs Ragas: broader agent/chat coverage, not RAG-only.
- GraphitiApache-2.0 temporal knowledge graphs for agent memory; Zep's OSS engine
Framework only — you run a graph DB + ops. Zep Community Edition deprecated; full Zep SaaS is separate (Free credits then Flex ~$104/mo billed annually). Not a drop-in chat history store. Schema/ops cost vs Mem0 Docker.