{"data":{"slug":"braintrust","name":"Braintrust","tagline":"Eval-first LLM platform that blocks bad prompts in CI","category":"agents","categoryName":"AI Agents","url":"https://braintrust.dev","github":null,"pricing":"freemium","hook":"Braintrust blocks your merge when the prompt makes the model worse.","whatsGood":"Braintrust is built around one loop: change a prompt, measure the result, ship the better one. The Eval() primitive runs experiments with real scorers—exact match, embedding similarity, LLM-as-judge factuality—and the CI integration blocks merges on statistically significant regressions.","theCatch":"Eval-first means observability is secondary—if your primary loop is debugging traces, look elsewhere. Learning curve to build eval datasets that actually represent production.","verdict":"The platform for teams whose deployment loop is change-measure-ship.","addedAt":"2026-08-19","featured":true,"domain":"AI Agents","subSpecialty":"Memory & Evals","capabilities":["evals","ci"],"surfaces":["Hosted"],"ecosystem":["Multi-platform"],"licenseModel":"Freemium","githubStars":null},"related":[{"slug":"claude-code","name":"Claude Code","tagline":"Terminal-first coding agent with deep repository context"},{"slug":"cursor","name":"Cursor","tagline":"The AI-first code editor that understands your codebase"},{"slug":"devin","name":"Devin","tagline":"Autonomous AI engineer for defined engineering tasks"},{"slug":"stagehand","name":"Stagehand","tagline":"AI-native browser automation that mixes code and natural language"},{"slug":"browser-use","name":"Browser Use","tagline":"Open-source framework for LLM-driven browser agents"}],"links":{"self":"/api/gems/braintrust","html":"/gems/braintrust","markdown":"/gems/braintrust.md","category":"/api/gems?category=agents","categoryPage":"/stacks/agents"}}