{"data":{"slug":"promptfoo","name":"Promptfoo","tagline":"MIT LLM evals & red-teaming CLI; Community free, 10k probes/mo","category":"agents","categoryName":"AI Agents","url":"https://www.promptfoo.dev","github":"https://github.com/promptfoo/promptfoo","pricing":"freemium","hook":"Promptfoo is declarative LLM evals and red-teaming when Braintrust is the hosted eval platform and you want YAML-in-CI first.","whatsGood":"MIT. Compare prompts/models, score outputs, CI gates, red-team/OWASP-style scans. Runs locally — prompts stay on your machine. Works with OpenAI, Anthropic, Azure, Bedrock, Ollama, custom providers. Now part of OpenAI with stated MIT permanence.","theCatch":"Vs DeepEval/Braintrust: YAML/CLI-first, not pytest-native Python or Braintrust SaaS. LOUD: Community caps red-team probes (~10k/mo); Enterprise/On-Prem are contact-sales (SSO, dashboards, dedicated runner). LLM API costs are yours. Acquisition governance is a multi-year bet.","verdict":"MIT eval/red-team CLI. Free core; probe cap + Enterprise sales wall.","addedAt":"2026-09-04","featured":false,"domain":"AI Agents","subSpecialty":"Memory & Evals","capabilities":["evals","red-team","ci"],"surfaces":["CLI","Hosted"],"ecosystem":["Node.js","Multi-platform"],"licenseModel":"Freemium","githubStars":24815},"related":[{"slug":"claude-code","name":"Claude Code","tagline":"Terminal-first coding agent with deep repository context"},{"slug":"cursor","name":"Cursor","tagline":"The AI-first code editor that understands your codebase"},{"slug":"devin","name":"Devin","tagline":"Autonomous AI engineer for defined engineering tasks"},{"slug":"stagehand","name":"Stagehand","tagline":"AI-native browser automation that mixes code and natural language"},{"slug":"browser-use","name":"Browser Use","tagline":"Open-source framework for LLM-driven browser agents"}],"links":{"self":"/api/gems/promptfoo","html":"/gems/promptfoo","markdown":"/gems/promptfoo.md","category":"/api/gems?category=agents","categoryPage":"/stacks/agents"}}