Skip to content

Evaluating Agents, Prompts and CLAUDE.md Changes: promptfoo, Inspect SWE, Langfuse and Braintrust

Evaluating a coding agent means running a fixed set of real tasks against two configurations and comparing outcomes that code can check. promptfoo runs Claude Code and Codex through their SDKs, Inspect SWE runs them inside Docker sandboxes, and Langfuse or Braintrust trace each session so that a failed task can be explained.

Someone on your team opens a pull request that cuts CLAUDE.md from 400 lines to 120 and adds two new rules. The review thread is six opinions and no data: one person says the agent “felt sharper” after a day with the new file, another says it stopped running the tests. You are about to merge a change that every agent session in the repository will read, and nobody can say whether it helps. This page is for the developer who builds the evaluation and the tech lead who decides whether the rules change ships.

What you get from an eval setup for agent rules

Section titled “What you get from an eval setup for agent rules”
  • A promptfoo config that A/B tests two CLAUDE.md variants on the same tasks, the same commit and the same pinned model, with deterministic checks (tests pass, diff stays in scope, tests were run) before any model-graded score.
  • The same experiment for Codex and AGENTS.md, and the route for Cursor, where promptfoo has no provider.
  • An Inspect SWE task for when you want every run in a disposable Docker sandbox with the model key kept on the host.
  • A decision rule for reading the results, so a rules change is merged on evidence rather than on a feeling.
  • Where trace platforms fit: Langfuse, Braintrust, Opik and Phoenix explain why a task failed; they do not replace the fixed task set.
  • Three copy-paste prompts and a failure-modes table for the ways these experiments quietly measure nothing.

Every command and config key below was checked on 2026-09-26 against promptfoo 0.123.1, inspect-ai 0.3.269 with inspect-swe 0.2.71, Claude Code 2.1.283 and Codex CLI 0.157.1. For the general method of building an eval set to choose a model, see reading benchmarks and building your own evals; this page is about evaluating changes to your harness: rules files, prompts, skills and plugins.

Start from the question. The first three rows run the agent; the rest observe it or test something you ship.

QuestionToolRuns the agent?Where it runs
Is the new CLAUDE.md or AGENTS.md better on our tasks?promptfoo (anthropic:claude-agent-sdk, openai:codex-sdk)Yes, through the vendor SDKYour machine or a CI runner
Same, but every run isolated in a container, several agents in one harnessInspect AI + Inspect SWE (claude_code(), codex_cli())Yes, inside a sandbox; model calls proxied through InspectDocker or Kubernetes
Does this plugin or skill help compared with no plugin?claude plugin eval (Claude Code 2.1.283)Yes, with a no-plugin baseline armYour machine
Why did this session fail, and what did it cost?Langfuse plugin, Braintrust bt trace, OpikNo, traces real sessionsSaaS or self-hosted
Local trace UI for OpenTelemetry data, no accountArize PhoenixNoYour machine
Does the LLM feature we ship regress or leak?DeepEval, garak, OpenLLMetryNo, tests your appCI, your app

Adoption as of 2026-09-26, from GitHub and the package registries: promptfoo 25.5k stars (npm 0.123.1; its README says it “is now part of OpenAI” and remains MIT-licensed), Langfuse 35.1k (the Claude Code plugin repository itself: 24), Opik 22.2k, DeepEval 18.4k, Arize Phoenix 11.6k, garak 9.4k, OpenLLMetry 7.4k, Inspect AI 2.9k and Inspect SWE 33. Braintrust’s coding-agent plugin monorepo had 2 stars (npm braintrust 3.35.0). The platforms are established; their coding-agent plugins are months old.

A/B test a CLAUDE.md change with promptfoo

Section titled “A/B test a CLAUDE.md change with promptfoo”

The experiment has one variable. Both arms get the same repository commit, the same tasks, the same model and the same tool permissions; only the rules file differs. promptfoo runs each task through each arm, runs your checks, and shows the pass rate, cost and latency per arm side by side.

  1. Build a fixed task set of 10 to 20 real tasks. Take them from merged pull requests that came with tests: a bug fix, a small feature, a refactor, a test-writing task. Each task needs a prompt a developer would really write and a check that code can run. Freeze the list; changing tasks between runs breaks the comparison.

  2. Create two worktrees outside the repository, both at the same commit, and commit the proposed rules into the second one. Claude Code reads CLAUDE.md files from parent directories too, so a worktree inside your repository would also load the current file into the “proposed” arm. In a terminal, from the repository root:

    Terminal window
    BASE=$(git rev-parse origin/main)
    mkdir -p ../rules-ab
    git worktree add --detach ../rules-ab/current "$BASE"
    git worktree add --detach ../rules-ab/proposed "$BASE"
    git show origin/trim-claude-md:CLAUDE.md > ../rules-ab/proposed/CLAUDE.md
    git -C ../rules-ab/proposed add CLAUDE.md && git -C ../rules-ab/proposed commit -qm "eval: proposed CLAUDE.md"
    (cd ../rules-ab/current && npm ci) && (cd ../rules-ab/proposed && npm ci)
    cd ../rules-ab && npm install --save-dev promptfoo@0.123.1 @anthropic-ai/claude-agent-sdk

    trim-claude-md stands for the branch of the pull request under review. Committing the proposed file lets the reset hook return each arm to a clean state with git reset --hard without losing the rules under test.

  3. Write the config. Save it as ../rules-ab/promptfooconfig.yaml. The two providers are identical except for label and working_dir:

    rules-ab/promptfooconfig.yaml
    description: CLAUDE.md A/B on a fixed task set
    prompts:
    - '{{task}}'
    providers:
    - id: anthropic:claude-agent-sdk
    label: current
    config:
    working_dir: ./current
    setting_sources: ['project'] # load CLAUDE.md from working_dir; no user-level files
    model: claude-opus-5-5 # pin it: a default that moves mid-experiment ruins the A/B
    tools: ['Read', 'Grep', 'Glob', 'Write', 'Edit', 'Bash'] # what the agent can see
    permission_mode: dontAsk # anything not pre-approved below is denied, never prompted
    append_allowed_tools: ['Write', 'Edit', 'Bash(npm test:*)', 'Bash(npx tsc:*)']
    disallowed_tools: ['WebFetch', 'WebSearch']
    max_turns: 40
    max_budget_usd: 3
    - id: anthropic:claude-agent-sdk
    label: proposed
    config:
    working_dir: ./proposed
    setting_sources: ['project']
    model: claude-opus-5-5
    tools: ['Read', 'Grep', 'Glob', 'Write', 'Edit', 'Bash'] # what the agent can see
    permission_mode: dontAsk # anything not pre-approved below is denied, never prompted
    append_allowed_tools: ['Write', 'Edit', 'Bash(npm test:*)', 'Bash(npx tsc:*)']
    disallowed_tools: ['WebFetch', 'WebSearch']
    max_turns: 40
    max_budget_usd: 3
    extensions:
    - file://reset.js:extensionHook
    defaultTest:
    options:
    provider: openai:responses:gpt-6-sol # the judge comes from another vendor than the agent
    assert:
    - type: javascript
    value: file://checks.js:testsPass
    metric: tests_pass
    - type: javascript
    metric: ran_tests
    value: |
    const calls = context.providerResponse?.metadata?.toolCalls || [];
    return calls.some(t => t.name === 'Bash' && t.input?.command?.includes('npm test'));
    - type: llm-rubric
    metric: conventions
    value: >-
    The final message states what changed and why, and names the tests it ran.
    You see only this message, not the diff.
    tests:
    - description: invoice rounding bug (from PR 1841)
    vars:
    task: >-
    Invoice totals are one cent off when a line has quantity 3 and unit price 0.10.
    Fix it in src/billing and add a regression test.
    assert:
    - type: javascript
    metric: in_scope
    value: file://checks.js:diffInScope
    config: { allowedPrefixes: ['src/billing/', 'test/billing/'] }
    # ...9 to 19 more tasks in the same shape

    setting_sources is off by default in the promptfoo provider, so without it neither arm reads any CLAUDE.md and the two arms are the same experiment twice. Set tools explicitly: with a working_dir and no tools, promptfoo 0.123.1 derives the tool list from the allowed-tools entries, and a pattern such as Bash(npm test:*) is a permission rule, not a tool name. tools decides what the agent can call; append_allowed_tools decides what runs without a prompt; dontAsk refuses the rest. Take the current model ID from the model comparison hub when you run this.

  4. Add the checks and the reset hook next to the config. The checks look at the arm’s working directory after the agent finishes; the hook puts both arms back to their commit after every test.

    rules-ab/checks.js
    const { execFileSync } = require('node:child_process');
    const path = require('node:path');
    const dirOf = (context) => path.join(__dirname, context.provider.label);
    const run = (cmd, args, cwd) => {
    try {
    return { ok: true, out: execFileSync(cmd, args, { cwd, encoding: 'utf8', stdio: 'pipe', timeout: 600000 }) };
    } catch (err) {
    return { ok: false, out: `${err.stdout || ''}${err.stderr || ''}` };
    }
    };
    module.exports.testsPass = (output, context) => {
    const r = run('npm', ['test', '--silent'], dirOf(context));
    return { pass: r.ok, score: r.ok ? 1 : 0, reason: r.ok ? 'npm test passed' : r.out.slice(-800) };
    };
    module.exports.diffInScope = (output, context) => {
    const cwd = dirOf(context);
    const changed = run('git', ['diff', '--name-only', 'HEAD'], cwd).out;
    const added = run('git', ['ls-files', '--others', '--exclude-standard'], cwd).out;
    const files = `${changed}\n${added}`.split('\n').filter(Boolean);
    const outside = files.filter((f) => !context.config.allowedPrefixes.some((p) => f.startsWith(p)));
    return { pass: outside.length === 0, score: outside.length ? 0 : 1, reason: outside.length ? `outside scope: ${outside.join(', ')}` : `${files.length} files, all in scope` };
    };
    rules-ab/reset.js
    const { execFileSync } = require('node:child_process');
    const path = require('node:path');
    function reset() {
    for (const arm of ['current', 'proposed']) {
    const cwd = path.join(__dirname, arm);
    execFileSync('git', ['reset', '--hard', '--quiet', 'HEAD'], { cwd });
    execFileSync('git', ['clean', '-fdq'], { cwd }); // keeps ignored files such as node_modules
    }
    }
    module.exports.extensionHook = async (hookName) => {
    if (hookName === 'beforeAll' || hookName === 'afterEach') reset();
    };
  5. Run it serially, three times per task, with the cache off. Serial, because both arms share the reset hook; three repeats, because one agent run per task is too noisy to compare; no cache, because a cached answer is not a new run.

    Terminal window
    cd ../rules-ab
    # Load ANTHROPIC_API_KEY (the agent) and OPENAI_API_KEY (the judge) into the environment
    # from your secret manager; never type a key on the command line. To use your local
    # Claude Code login instead of an API key, add apiKeyRequired: false to both providers.
    npx promptfoo eval -j 1 --repeat 3 --no-cache -o results.json
    npx promptfoo view

    promptfoo eval exits with code 100 when at least one test fails, so the same command works as a CI gate later. promptfoo view opens the matrix in a browser: one column per arm, one row per task, with the named metrics (tests_pass, ran_tests, conventions, in_scope) and cost per cell.

Run the same experiment in Claude Code, Codex and Cursor

Section titled “Run the same experiment in Claude Code, Codex and Cursor”

The task set, the checks and the decision rule do not change. What changes is the provider and the rules file each agent reads.

Use the config above. The provider IDs are anthropic:claude-agent-sdk and its alias anthropic:claude-code; anthropic:claude-code-sdk does not exist. promptfoo 0.123.1 installs @anthropic-ai/claude-agent-sdk 0.3.263 as an optional dependency; installing it explicitly, as in step 2, gets the current release (npm 0.3.283 on 2026-09-26) and covers installs that skip optional dependencies.

To A/B a skill or a plugin rather than the rules file, Claude Code 2.1.283 has a built-in runner. It reads eval cases from the plugin’s evals/ directory and adds a no-plugin baseline arm automatically:

Terminal window
claude plugin eval ./plugins/release-notes --runs 3 --threshold 0.8 --no-publish

The first run in a new plugin directory asks you to confirm trust; in CI pass --trust-plugin, and only for plugins you would run yourself. --threshold makes the command exit 1 when any case scores below it, and --max-cost-usd sets a hard ceiling for the run. It runs the plugin’s code as you, so evaluate only plugins you trust. The case format is covered on building and distributing a plugin.

Choose Inspect SWE when you want every run in a fresh container, or one harness for several agents: its claude_code() and codex_cli() agents run inside the sample’s sandbox, and their model calls are proxied back through Inspect, so token limits, time limits and transcripts work as for any Inspect eval and the API key never enters the container.

rules_ab.py
from pathlib import Path
from inspect_ai import Task, task
from inspect_ai.dataset import json_dataset
from inspect_ai.scorer import CORRECT, INCORRECT, Score, Target, accuracy, scorer, stderr
from inspect_ai.solver import Generate, TaskState, solver
from inspect_ai.util import sandbox
from inspect_swe import claude_code
@solver
def install_rules(rules_file: str):
async def solve(state: TaskState, generate: Generate) -> TaskState:
await sandbox().write_file("/repo/CLAUDE.md", Path(rules_file).read_text())
return state
return solve
@scorer(metrics=[accuracy(), stderr()])
def tests_pass():
async def score(state: TaskState, target: Target) -> Score:
result = await sandbox().exec(["npm", "test", "--silent"], cwd="/repo")
return Score(value=CORRECT if result.success else INCORRECT, explanation=result.stderr[-800:])
return score
@task
def rules_ab(rules: str = "rules/current.md") -> Task:
return Task(
dataset=json_dataset("tasks.jsonl"), # one {"id", "input", "target"} per line
setup=install_rules(rules),
solver=claude_code(cwd="/repo", disallowed_tools=["WebSearch", "WebFetch"]),
scorer=tests_pass(),
sandbox=("docker", "compose.yaml"), # image with the repository at /repo, deps installed, no secrets
epochs=3,
)

Run each arm, then compare them in the log viewer:

Terminal window
pip install inspect-ai==0.3.269 inspect-swe==0.2.71
inspect eval rules_ab.py --model anthropic/claude-opus-5-5 -T rules=rules/current.md
inspect eval rules_ab.py --model anthropic/claude-opus-5-5 -T rules=rules/proposed.md
inspect view

For Codex, import codex_cli and write AGENTS.md instead. In Inspect SWE 0.2.71, codex_cli() defaults to sandbox_mode="danger-full-access" (Codex’s own sandbox off) and web_search="live". The first is acceptable only because the Inspect container is the boundary, so keep that container free of credentials. Pass web_search="disabled" to both arms, the equivalent of disallowed_tools=["WebSearch", "WebFetch"] on the Claude Code side, so neither arm can look up the answer. The PyPI package installs with pip install inspect-swe; the README’s own install line is the git development build.

Write the rule before you run, so the result cannot be argued into the answer someone wanted.

ResultDecision
tests_pass for proposed is at least current’s, in_scope and ran_tests no worse, cost per task within 20%Merge the rules change, and paste the per-arm table into the pull request
Proposed wins on some tasks and loses on othersRead the tasks where the arms disagree in all three repeats; those are signal, one-off flips are noise. Fix the rule that caused the loss and rerun
Pass rates equal, proposed cheaper or fasterMerge: a shorter file that keeps behaviour saves tokens on every session
Proposed worse on tests_pass in any task across all three repeatsDo not merge. Find the rule that the old file had and the new one lost
Only the conventions judge score differsTreat it as a hint, not a verdict, until the judge is calibrated against your own labels (model-graded checks)

Row 4 overrides row 1: a consistent loss on any task blocks the merge even when the overall pass rate is equal or higher.

Deterministic checks outrank the judge. A judge that likes longer final messages will reward a verbose rules file, and that is how a regression gets merged with a better score.

A fixed task set tells you whether a change helps. A trace tells you why one task failed: which file the agent read first, which rule it quoted, where it gave up. Add tracing when the eval shows a loss you cannot explain from the final message.

The official plugin traces prompts, turns, model generations with token usage and cost, tool calls, skills and subagents for the claude CLI and the desktop app in Code mode. It needs uv on PATH.

Terminal window
claude plugin marketplace add langfuse/Claude-Observability-Plugin
claude plugin install langfuse-observability@langfuse-observability

Then configure the keys from inside a session with /plugin configure langfuse-observability@langfuse-observability; the secret key goes to the OS keychain, not a file. Restart Claude Code. If no trace appears, read ~/.claude/state/langfuse_hook.log. Langfuse can then run datasets and LLM-as-judge evaluations over the traced sessions.

Context cost: promptfoo, Inspect and bt trace run run outside your interactive sessions and add nothing to their context. The Langfuse and Braintrust plugins work through hooks; run claude plugin details <plugin> after installing to see the projected token cost of any component they add. Opik’s MCP server adds tool definitions to every session that loads it, so measure it with /context before and after.

For your own application’s LLM features rather than the coding agent, the tools differ: DeepEval (deepeval test run test_chatbot.py) gates a feature in CI with pytest-style assertions, garak (pip install garak) probes a model endpoint for jailbreaks, prompt injection and leakage, and OpenLLMetry (pip install traceloop-sdk, then Traceloop.init()) instruments the app with OpenTelemetry. None of the three traces a coding agent. The legacy openai/evals repository now points to evals in the OpenAI dashboard; its PyPI package evals last shipped in May 2024.

An eval that cannot fail is a tautology with a cost line. Before you trust a result, prove the harness can detect a difference:

  • Canary rule. Add a harmless rule to one arm only (“end your final message with CANARY-B”) and run one task. If the marker does not appear, the rules file is not being loaded.
  • Known-bad variant. Run once with a proposed file that tells the agent not to run tests. ran_tests and usually tests_pass must drop. If they do not, the checks are too weak; see oracle strength.
  • Judge from another vendor. The agent arm never grades itself: the config above sends llm-rubric to a different vendor’s model. Calibrate it on 20 hand-labelled outputs before it can block a merge.
  • Rerun on every model change. A rules file tuned on one model can regress on the next. The new-model playbook reruns this task set as its first step.

Ownership: the developer who proposes the rules change runs the A/B and attaches the per-arm table to the pull request, as part of its evidence bundle. The tech lead owns the task set, reviews the table rather than the prose of the new file, and signs off the merge.

What breaks when you evaluate agent rules?

Section titled “What breaks when you evaluate agent rules?”
SymptomCauseRecovery
Both arms score identically on every tasksetting_sources missing, so neither arm reads CLAUDE.mdAdd setting_sources: ['project']; confirm with the canary rule
The proposed arm still follows a rule you deletedWorktree inside the repository: Claude Code also loads the parent directory’s CLAUDE.mdPut both worktrees outside the repository, as in step 2
A personal preference changes results on one laptopuser in setting_sources loads ~/.claude/CLAUDE.mdUse ['project'] only
The second run finishes in seconds with the same scorespromptfoo served cached responsesAlways pass --no-cache for agent evals
Tests pass in one arm because of the other arm’s editsParallel runs or no reset between tests-j 1 and the reset.js hook
Scores flip between runs of the same taskOne run per task, or fewer than 10 tasks--repeat 3 and more tasks; decide on consistent disagreements only
The Codex arm ignores AGENTS.mdWorktree path not trusted (Codex 0.150.0 and later)Mark both eval paths trusted; confirm with the canary rule
One task burns the budgetA loop of failing tool callsmax_turns and max_budget_usd per provider; --max-cost-usd for claude plugin eval
Results shift between Monday and FridayThe default model changedPin model in both arms, and rerun both arms together
Traces show API keys and file contents in the vendor UIBraintrust and Langfuse capture prompts and tool outputConfigure redaction first; trace a sanitised fixture repository

To run this task set on every harness change rather than once per PR, see continuous evals.

Before you change the rules file, read how to write and trim CLAUDE.md; this page tells you whether the trimmed version still works. For the production side of the same data, see observing the agents. The section hub, measuring agentic engineering, puts evals beside telemetry, cost tracking, review bots and security gates.