Evaluating Agents, Prompts and CLAUDE.md Changes: promptfoo, Inspect SWE, Langfuse and Braintrust
Evaluating a coding agent means running a fixed set of real tasks against two configurations and comparing outcomes that code can check. promptfoo runs Claude Code and Codex through their SDKs, Inspect SWE runs them inside Docker sandboxes, and Langfuse or Braintrust trace each session so that a failed task can be explained.
Someone on your team opens a pull request that cuts CLAUDE.md from 400 lines to 120 and adds two new rules. The review thread is six opinions and no data: one person says the agent “felt sharper” after a day with the new file, another says it stopped running the tests. You are about to merge a change that every agent session in the repository will read, and nobody can say whether it helps. This page is for the developer who builds the evaluation and the tech lead who decides whether the rules change ships.
What you get from an eval setup for agent rules
Section titled “What you get from an eval setup for agent rules”- A promptfoo config that A/B tests two
CLAUDE.mdvariants on the same tasks, the same commit and the same pinned model, with deterministic checks (tests pass, diff stays in scope, tests were run) before any model-graded score. - The same experiment for Codex and
AGENTS.md, and the route for Cursor, where promptfoo has no provider. - An Inspect SWE task for when you want every run in a disposable Docker sandbox with the model key kept on the host.
- A decision rule for reading the results, so a rules change is merged on evidence rather than on a feeling.
- Where trace platforms fit: Langfuse, Braintrust, Opik and Phoenix explain why a task failed; they do not replace the fixed task set.
- Three copy-paste prompts and a failure-modes table for the ways these experiments quietly measure nothing.
Every command and config key below was checked on 2026-09-26 against promptfoo 0.123.1, inspect-ai 0.3.269 with inspect-swe 0.2.71, Claude Code 2.1.283 and Codex CLI 0.157.1. For the general method of building an eval set to choose a model, see reading benchmarks and building your own evals; this page is about evaluating changes to your harness: rules files, prompts, skills and plugins.
Which eval tool answers which question?
Section titled “Which eval tool answers which question?”Start from the question. The first three rows run the agent; the rest observe it or test something you ship.
| Question | Tool | Runs the agent? | Where it runs |
|---|---|---|---|
Is the new CLAUDE.md or AGENTS.md better on our tasks? | promptfoo (anthropic:claude-agent-sdk, openai:codex-sdk) | Yes, through the vendor SDK | Your machine or a CI runner |
| Same, but every run isolated in a container, several agents in one harness | Inspect AI + Inspect SWE (claude_code(), codex_cli()) | Yes, inside a sandbox; model calls proxied through Inspect | Docker or Kubernetes |
| Does this plugin or skill help compared with no plugin? | claude plugin eval (Claude Code 2.1.283) | Yes, with a no-plugin baseline arm | Your machine |
| Why did this session fail, and what did it cost? | Langfuse plugin, Braintrust bt trace, Opik | No, traces real sessions | SaaS or self-hosted |
| Local trace UI for OpenTelemetry data, no account | Arize Phoenix | No | Your machine |
| Does the LLM feature we ship regress or leak? | DeepEval, garak, OpenLLMetry | No, tests your app | CI, your app |
Adoption as of 2026-09-26, from GitHub and the package registries: promptfoo 25.5k stars (npm 0.123.1; its README says it “is now part of OpenAI” and remains MIT-licensed), Langfuse 35.1k (the Claude Code plugin repository itself: 24), Opik 22.2k, DeepEval 18.4k, Arize Phoenix 11.6k, garak 9.4k, OpenLLMetry 7.4k, Inspect AI 2.9k and Inspect SWE 33. Braintrust’s coding-agent plugin monorepo had 2 stars (npm braintrust 3.35.0). The platforms are established; their coding-agent plugins are months old.
A/B test a CLAUDE.md change with promptfoo
Section titled “A/B test a CLAUDE.md change with promptfoo”The experiment has one variable. Both arms get the same repository commit, the same tasks, the same model and the same tool permissions; only the rules file differs. promptfoo runs each task through each arm, runs your checks, and shows the pass rate, cost and latency per arm side by side.
-
Build a fixed task set of 10 to 20 real tasks. Take them from merged pull requests that came with tests: a bug fix, a small feature, a refactor, a test-writing task. Each task needs a prompt a developer would really write and a check that code can run. Freeze the list; changing tasks between runs breaks the comparison.
-
Create two worktrees outside the repository, both at the same commit, and commit the proposed rules into the second one. Claude Code reads
CLAUDE.mdfiles from parent directories too, so a worktree inside your repository would also load the current file into the “proposed” arm. In a terminal, from the repository root:Terminal window BASE=$(git rev-parse origin/main)mkdir -p ../rules-abgit worktree add --detach ../rules-ab/current "$BASE"git worktree add --detach ../rules-ab/proposed "$BASE"git show origin/trim-claude-md:CLAUDE.md > ../rules-ab/proposed/CLAUDE.mdgit -C ../rules-ab/proposed add CLAUDE.md && git -C ../rules-ab/proposed commit -qm "eval: proposed CLAUDE.md"(cd ../rules-ab/current && npm ci) && (cd ../rules-ab/proposed && npm ci)cd ../rules-ab && npm install --save-dev promptfoo@0.123.1 @anthropic-ai/claude-agent-sdktrim-claude-mdstands for the branch of the pull request under review. Committing the proposed file lets the reset hook return each arm to a clean state withgit reset --hardwithout losing the rules under test. -
Write the config. Save it as
../rules-ab/promptfooconfig.yaml. The two providers are identical except forlabelandworking_dir:rules-ab/promptfooconfig.yaml description: CLAUDE.md A/B on a fixed task setprompts:- '{{task}}'providers:- id: anthropic:claude-agent-sdklabel: currentconfig:working_dir: ./currentsetting_sources: ['project'] # load CLAUDE.md from working_dir; no user-level filesmodel: claude-opus-5-5 # pin it: a default that moves mid-experiment ruins the A/Btools: ['Read', 'Grep', 'Glob', 'Write', 'Edit', 'Bash'] # what the agent can seepermission_mode: dontAsk # anything not pre-approved below is denied, never promptedappend_allowed_tools: ['Write', 'Edit', 'Bash(npm test:*)', 'Bash(npx tsc:*)']disallowed_tools: ['WebFetch', 'WebSearch']max_turns: 40max_budget_usd: 3- id: anthropic:claude-agent-sdklabel: proposedconfig:working_dir: ./proposedsetting_sources: ['project']model: claude-opus-5-5tools: ['Read', 'Grep', 'Glob', 'Write', 'Edit', 'Bash'] # what the agent can seepermission_mode: dontAsk # anything not pre-approved below is denied, never promptedappend_allowed_tools: ['Write', 'Edit', 'Bash(npm test:*)', 'Bash(npx tsc:*)']disallowed_tools: ['WebFetch', 'WebSearch']max_turns: 40max_budget_usd: 3extensions:- file://reset.js:extensionHookdefaultTest:options:provider: openai:responses:gpt-6-sol # the judge comes from another vendor than the agentassert:- type: javascriptvalue: file://checks.js:testsPassmetric: tests_pass- type: javascriptmetric: ran_testsvalue: |const calls = context.providerResponse?.metadata?.toolCalls || [];return calls.some(t => t.name === 'Bash' && t.input?.command?.includes('npm test'));- type: llm-rubricmetric: conventionsvalue: >-The final message states what changed and why, and names the tests it ran.You see only this message, not the diff.tests:- description: invoice rounding bug (from PR 1841)vars:task: >-Invoice totals are one cent off when a line has quantity 3 and unit price 0.10.Fix it in src/billing and add a regression test.assert:- type: javascriptmetric: in_scopevalue: file://checks.js:diffInScopeconfig: { allowedPrefixes: ['src/billing/', 'test/billing/'] }# ...9 to 19 more tasks in the same shapesetting_sourcesis off by default in the promptfoo provider, so without it neither arm reads anyCLAUDE.mdand the two arms are the same experiment twice. Settoolsexplicitly: with aworking_dirand notools, promptfoo 0.123.1 derives the tool list from the allowed-tools entries, and a pattern such asBash(npm test:*)is a permission rule, not a tool name.toolsdecides what the agent can call;append_allowed_toolsdecides what runs without a prompt;dontAskrefuses the rest. Take the current model ID from the model comparison hub when you run this. -
Add the checks and the reset hook next to the config. The checks look at the arm’s working directory after the agent finishes; the hook puts both arms back to their commit after every test.
rules-ab/checks.js const { execFileSync } = require('node:child_process');const path = require('node:path');const dirOf = (context) => path.join(__dirname, context.provider.label);const run = (cmd, args, cwd) => {try {return { ok: true, out: execFileSync(cmd, args, { cwd, encoding: 'utf8', stdio: 'pipe', timeout: 600000 }) };} catch (err) {return { ok: false, out: `${err.stdout || ''}${err.stderr || ''}` };}};module.exports.testsPass = (output, context) => {const r = run('npm', ['test', '--silent'], dirOf(context));return { pass: r.ok, score: r.ok ? 1 : 0, reason: r.ok ? 'npm test passed' : r.out.slice(-800) };};module.exports.diffInScope = (output, context) => {const cwd = dirOf(context);const changed = run('git', ['diff', '--name-only', 'HEAD'], cwd).out;const added = run('git', ['ls-files', '--others', '--exclude-standard'], cwd).out;const files = `${changed}\n${added}`.split('\n').filter(Boolean);const outside = files.filter((f) => !context.config.allowedPrefixes.some((p) => f.startsWith(p)));return { pass: outside.length === 0, score: outside.length ? 0 : 1, reason: outside.length ? `outside scope: ${outside.join(', ')}` : `${files.length} files, all in scope` };};rules-ab/reset.js const { execFileSync } = require('node:child_process');const path = require('node:path');function reset() {for (const arm of ['current', 'proposed']) {const cwd = path.join(__dirname, arm);execFileSync('git', ['reset', '--hard', '--quiet', 'HEAD'], { cwd });execFileSync('git', ['clean', '-fdq'], { cwd }); // keeps ignored files such as node_modules}}module.exports.extensionHook = async (hookName) => {if (hookName === 'beforeAll' || hookName === 'afterEach') reset();}; -
Run it serially, three times per task, with the cache off. Serial, because both arms share the reset hook; three repeats, because one agent run per task is too noisy to compare; no cache, because a cached answer is not a new run.
Terminal window cd ../rules-ab# Load ANTHROPIC_API_KEY (the agent) and OPENAI_API_KEY (the judge) into the environment# from your secret manager; never type a key on the command line. To use your local# Claude Code login instead of an API key, add apiKeyRequired: false to both providers.npx promptfoo eval -j 1 --repeat 3 --no-cache -o results.jsonnpx promptfoo viewpromptfoo evalexits with code100when at least one test fails, so the same command works as a CI gate later.promptfoo viewopens the matrix in a browser: one column per arm, one row per task, with the named metrics (tests_pass,ran_tests,conventions,in_scope) and cost per cell.
Run the same experiment in Claude Code, Codex and Cursor
Section titled “Run the same experiment in Claude Code, Codex and Cursor”The task set, the checks and the decision rule do not change. What changes is the provider and the rules file each agent reads.
Use the config above. The provider IDs are anthropic:claude-agent-sdk and its alias anthropic:claude-code; anthropic:claude-code-sdk does not exist. promptfoo 0.123.1 installs @anthropic-ai/claude-agent-sdk 0.3.263 as an optional dependency; installing it explicitly, as in step 2, gets the current release (npm 0.3.283 on 2026-09-26) and covers installs that skip optional dependencies.
To A/B a skill or a plugin rather than the rules file, Claude Code 2.1.283 has a built-in runner. It reads eval cases from the plugin’s evals/ directory and adds a no-plugin baseline arm automatically:
claude plugin eval ./plugins/release-notes --runs 3 --threshold 0.8 --no-publishThe first run in a new plugin directory asks you to confirm trust; in CI pass --trust-plugin, and only for plugins you would run yourself. --threshold makes the command exit 1 when any case scores below it, and --max-cost-usd sets a hard ceiling for the run. It runs the plugin’s code as you, so evaluate only plugins you trust. The case format is covered on building and distributing a plugin.
Codex reads AGENTS.md, so the variants differ in that file. Swap the two providers for the Codex SDK provider; openai:codex is an alias. promptfoo 0.123.1 declares @openai/codex-sdk ^0.153.2 as an optional dependency (npm 0.157.1 on 2026-09-26).
providers: - id: openai:codex-sdk label: current config: working_dir: ./current model: gpt-6-astra sandbox_mode: workspace-write # writes stay inside working_dir approval_policy: never network_access_enabled: false - id: openai:codex-sdk label: proposed config: working_dir: ./proposed model: gpt-6-astra sandbox_mode: workspace-write approval_policy: never network_access_enabled: falseDrop the ran_tests assertion: it reads Claude-style tool calls. The testsPass and diffInScope checks work unchanged. Since Codex 0.150.0 untrusted projects supply no project AGENTS.md; mark both worktree paths as trusted, and run the canary rule on the Codex arms before you trust a tie.
promptfoo 0.123.1 has no Cursor provider (checked against its provider index at tag 0.123.1). Keep the task set and the checks, and run them through Harbor’s cursor-cli agent (Harbor 0.23.0, 2026-09-12), which installs the Cursor CLI in each task’s container and drives it in print mode with auto-approval. Build two task directories that differ only in the rules files baked into the task image, then run both. Load CURSOR_API_KEY into the environment from your secret manager; Harbor passes it into the task container.
harbor run -p evals/tasks-current -a cursor-cli -m "$CURSOR_MODEL" -k 3 -n 4harbor run -p evals/tasks-proposed -a cursor-cli -m "$CURSOR_MODEL" -k 3 -n 4CURSOR_MODEL takes Harbor’s provider/model form; Harbor passes the part after the slash to Cursor, so use a model name your Cursor plan offers. -k 3 runs three attempts per task and -n 4 runs four at once. Because Harbor auto-approves every command, the task container is the only boundary: keep credentials other than the Cursor key out of it. Task layout, oracle and nop sanity runs and the other flags are on the benchmarks page.
Cursor’s model list and rules loading could not be verified on 2026-09-26 (cursor.com was unreachable from the writing environment). Before the first run, add a canary rule to both arms, such as “end every final message with the word CANARY”, and confirm it appears. A rules file that is not loaded produces a perfect tie.
Run the A/B in a sandbox with Inspect SWE
Section titled “Run the A/B in a sandbox with Inspect SWE”Choose Inspect SWE when you want every run in a fresh container, or one harness for several agents: its claude_code() and codex_cli() agents run inside the sample’s sandbox, and their model calls are proxied back through Inspect, so token limits, time limits and transcripts work as for any Inspect eval and the API key never enters the container.
from pathlib import Path
from inspect_ai import Task, taskfrom inspect_ai.dataset import json_datasetfrom inspect_ai.scorer import CORRECT, INCORRECT, Score, Target, accuracy, scorer, stderrfrom inspect_ai.solver import Generate, TaskState, solverfrom inspect_ai.util import sandboxfrom inspect_swe import claude_code
@solverdef install_rules(rules_file: str): async def solve(state: TaskState, generate: Generate) -> TaskState: await sandbox().write_file("/repo/CLAUDE.md", Path(rules_file).read_text()) return state return solve
@scorer(metrics=[accuracy(), stderr()])def tests_pass(): async def score(state: TaskState, target: Target) -> Score: result = await sandbox().exec(["npm", "test", "--silent"], cwd="/repo") return Score(value=CORRECT if result.success else INCORRECT, explanation=result.stderr[-800:]) return score
@taskdef rules_ab(rules: str = "rules/current.md") -> Task: return Task( dataset=json_dataset("tasks.jsonl"), # one {"id", "input", "target"} per line setup=install_rules(rules), solver=claude_code(cwd="/repo", disallowed_tools=["WebSearch", "WebFetch"]), scorer=tests_pass(), sandbox=("docker", "compose.yaml"), # image with the repository at /repo, deps installed, no secrets epochs=3, )Run each arm, then compare them in the log viewer:
pip install inspect-ai==0.3.269 inspect-swe==0.2.71inspect eval rules_ab.py --model anthropic/claude-opus-5-5 -T rules=rules/current.mdinspect eval rules_ab.py --model anthropic/claude-opus-5-5 -T rules=rules/proposed.mdinspect viewFor Codex, import codex_cli and write AGENTS.md instead. In Inspect SWE 0.2.71, codex_cli() defaults to sandbox_mode="danger-full-access" (Codex’s own sandbox off) and web_search="live". The first is acceptable only because the Inspect container is the boundary, so keep that container free of credentials. Pass web_search="disabled" to both arms, the equivalent of disallowed_tools=["WebSearch", "WebFetch"] on the Claude Code side, so neither arm can look up the answer. The PyPI package installs with pip install inspect-swe; the README’s own install line is the git development build.
How do you decide from the results?
Section titled “How do you decide from the results?”Write the rule before you run, so the result cannot be argued into the answer someone wanted.
| Result | Decision |
|---|---|
tests_pass for proposed is at least current’s, in_scope and ran_tests no worse, cost per task within 20% | Merge the rules change, and paste the per-arm table into the pull request |
| Proposed wins on some tasks and loses on others | Read the tasks where the arms disagree in all three repeats; those are signal, one-off flips are noise. Fix the rule that caused the loss and rerun |
| Pass rates equal, proposed cheaper or faster | Merge: a shorter file that keeps behaviour saves tokens on every session |
Proposed worse on tests_pass in any task across all three repeats | Do not merge. Find the rule that the old file had and the new one lost |
Only the conventions judge score differs | Treat it as a hint, not a verdict, until the judge is calibrated against your own labels (model-graded checks) |
Row 4 overrides row 1: a consistent loss on any task blocks the merge even when the overall pass rate is equal or higher.
Deterministic checks outrank the judge. A judge that likes longer final messages will reward a verbose rules file, and that is how a regression gets merged with a better score.
Where trace platforms fit in the loop
Section titled “Where trace platforms fit in the loop”A fixed task set tells you whether a change helps. A trace tells you why one task failed: which file the agent read first, which rule it quoted, where it gave up. Add tracing when the eval shows a loss you cannot explain from the final message.
The official plugin traces prompts, turns, model generations with token usage and cost, tool calls, skills and subagents for the claude CLI and the desktop app in Code mode. It needs uv on PATH.
claude plugin marketplace add langfuse/Claude-Observability-Pluginclaude plugin install langfuse-observability@langfuse-observabilityThen configure the keys from inside a session with /plugin configure langfuse-observability@langfuse-observability; the secret key goes to the OS keychain, not a file. Restart Claude Code. If no trace appears, read ~/.claude/state/langfuse_hook.log. Langfuse can then run datasets and LLM-as-judge evaluations over the traced sessions.
The bt CLI installs a plugin and a local tracing daemon for Claude Code, Codex, OpenCode, Pi, Grok and Antigravity. Each session becomes one trace with turn, model-call, tool and subagent spans.
curl -fsSL https://bt.dev/cli/install.sh | bash # Windows and mise installs: README of braintrustdata/btbt loginbt trace enable claude --project rules-abbt trace doctor claudebt trace run --project rules-ab --additional-metadata '{"arm":"proposed"}' claude -- -p "Fix the invoice rounding bug in src/billing"The plugin does not redact prompts, tool input and output, or system prompts locally. Configure Braintrust’s redaction before you enable tracing on a real repository.
Opik (Comet, Apache-2.0, self-hostable) connects to Claude Code, Codex, Cursor and others over MCP, so you can ask the agent to score its own recent traces:
pip install opik && opik configureuvx opik mcp configureArize Phoenix gives a local OpenTelemetry trace and eval UI with no account: pip install arize-phoenix && phoenix serve. Wiring Claude Code’s own sessions into Phoenix could not be verified on 2026-09-26; use it for traces your application emits.
Context cost: promptfoo, Inspect and bt trace run run outside your interactive sessions and add nothing to their context. The Langfuse and Braintrust plugins work through hooks; run claude plugin details <plugin> after installing to see the projected token cost of any component they add. Opik’s MCP server adds tool definitions to every session that loads it, so measure it with /context before and after.
For your own application’s LLM features rather than the coding agent, the tools differ: DeepEval (deepeval test run test_chatbot.py) gates a feature in CI with pytest-style assertions, garak (pip install garak) probes a model endpoint for jailbreaks, prompt injection and leakage, and OpenLLMetry (pip install traceloop-sdk, then Traceloop.init()) instruments the app with OpenTelemetry. None of the three traces a coding agent. The legacy openai/evals repository now points to evals in the OpenAI dashboard; its PyPI package evals last shipped in May 2024.
Copy-paste prompts for agent evals
Section titled “Copy-paste prompts for agent evals”How do you know the eval itself works?
Section titled “How do you know the eval itself works?”An eval that cannot fail is a tautology with a cost line. Before you trust a result, prove the harness can detect a difference:
- Canary rule. Add a harmless rule to one arm only (“end your final message with CANARY-B”) and run one task. If the marker does not appear, the rules file is not being loaded.
- Known-bad variant. Run once with a proposed file that tells the agent not to run tests.
ran_testsand usuallytests_passmust drop. If they do not, the checks are too weak; see oracle strength. - Judge from another vendor. The agent arm never grades itself: the config above sends
llm-rubricto a different vendor’s model. Calibrate it on 20 hand-labelled outputs before it can block a merge. - Rerun on every model change. A rules file tuned on one model can regress on the next. The new-model playbook reruns this task set as its first step.
Ownership: the developer who proposes the rules change runs the A/B and attaches the per-arm table to the pull request, as part of its evidence bundle. The tech lead owns the task set, reviews the table rather than the prose of the new file, and signs off the merge.
What breaks when you evaluate agent rules?
Section titled “What breaks when you evaluate agent rules?”| Symptom | Cause | Recovery |
|---|---|---|
| Both arms score identically on every task | setting_sources missing, so neither arm reads CLAUDE.md | Add setting_sources: ['project']; confirm with the canary rule |
| The proposed arm still follows a rule you deleted | Worktree inside the repository: Claude Code also loads the parent directory’s CLAUDE.md | Put both worktrees outside the repository, as in step 2 |
| A personal preference changes results on one laptop | user in setting_sources loads ~/.claude/CLAUDE.md | Use ['project'] only |
| The second run finishes in seconds with the same scores | promptfoo served cached responses | Always pass --no-cache for agent evals |
| Tests pass in one arm because of the other arm’s edits | Parallel runs or no reset between tests | -j 1 and the reset.js hook |
| Scores flip between runs of the same task | One run per task, or fewer than 10 tasks | --repeat 3 and more tasks; decide on consistent disagreements only |
The Codex arm ignores AGENTS.md | Worktree path not trusted (Codex 0.150.0 and later) | Mark both eval paths trusted; confirm with the canary rule |
| One task burns the budget | A loop of failing tool calls | max_turns and max_budget_usd per provider; --max-cost-usd for claude plugin eval |
| Results shift between Monday and Friday | The default model changed | Pin model in both arms, and rerun both arms together |
| Traces show API keys and file contents in the vendor UI | Braintrust and Langfuse capture prompts and tool output | Configure redaction first; trace a sanitised fixture repository |
Where to go next with agent evals
Section titled “Where to go next with agent evals”To run this task set on every harness change rather than once per PR, see continuous evals.
Before you change the rules file, read how to write and trim CLAUDE.md; this page tells you whether the trimmed version still works. For the production side of the same data, see observing the agents. The section hub, measuring agentic engineering, puts evals beside telemetry, cost tracking, review bots and security gates.