Skip to content

Skills That Make Agents Test and Debug Properly

Testing and debugging skills are Agent Skills that make a coding agent reproduce a bug, write the failing test first and show command output before it claims “fixed”: Matt Pocock’s tdd and diagnosing-bugs, Superpowers’ test-driven-development, systematic-debugging and verification-before-completion, and Anthropic’s webapp-testing. The skills shape behaviour but enforce nothing; CI and a protected test suite do.

The pattern they fix: a customer reports a wrong date, the agent edits three lines and says “Fixed”, and nothing shows the bug reproduced or the fix proven. A week later the bug is back in another time zone. This page is for developers who want the agent to debug like a careful colleague, and for tech leads who want every agent fix to arrive with evidence a reviewer can check without rereading the diff.

What you get from testing and debugging skills

Section titled “What you get from testing and debugging skills”
  • A shortlist of six skills, what each one makes the agent do, and which to combine so they do not fight
  • Install commands for Claude Code, Codex and Cursor, checked against the skills CLI 1.7.0 on 2026-09-26
  • A worked systematic-debugging session, from a customer bug report to a regression test that failed before it passed
  • A red-green-refactor workflow that ends in pasted verification evidence, not a claim
  • The gates outside the agent that prove the skills worked, and the failure modes to watch for

Which testing and debugging skills are worth installing?

Section titled “Which testing and debugging skills are worth installing?”

Six skills cover testing and debugging. They come from three repositories: mattpocock/skills, obra/superpowers (the Superpowers methodology) and anthropics/skills.

SkillRepositoryWhat it makes the agent doFires
tddmattpocock/skillsAgrees the test seams with you first, then one failing test and the minimum code per slice. Refactoring moves to the review stageWhen you ask for test-first work, red-green-refactor or integration tests
diagnosing-bugsmattpocock/skillsBuilds a fast pass/fail feedback loop that goes red on this bug before any theory, then 3–5 ranked, falsifiable hypothesesWhen you say “diagnose” or “debug this”, or report something broken or slow
test-driven-developmentobra/superpowersStrict red-green-refactor. Code written before its test is deleted and rewrittenBefore any feature or bugfix code
systematic-debuggingobra/superpowersFour phases: root-cause investigation, pattern analysis, one hypothesis at a time, then a failing test and a single fix. Stops after three failed fixes to question the designOn any bug, test failure or unexpected behaviour
verification-before-completionobra/superpowersNo success claim without fresh output from the command that proves itBefore the agent says done, fixed or passing, or commits
webapp-testinganthropics/skillsWrites Python Playwright scripts against a local web app, with a server-lifecycle helper, screenshots and console logsWhen you ask it to test or debug a local web UI

Popularity, as of 2026-09-26: all-time installs on skills.sh were tdd 966,524, diagnosing-bugs 665,090, systematic-debugging 271,653, test-driven-development 236,900, verification-before-completion 221,579 and webapp-testing 164,051. These counts come from the third-party LinklyAI/best-skills scrape dated 2026-09-26 (a secondary source; check the live numbers on skills.sh). GitHub stars on the same date: obra/superpowers 291.7k, mattpocock/skills 269.8k, anthropics/skills 178.4k.

Pocock or Superpowers: which set should you pick?

Section titled “Pocock or Superpowers: which set should you pick?”

Pick one family for tests and one for debugging. Two skills that fire on the same trigger give the agent two conflicting procedures.

If your team…InstallWhy
wants to agree what gets tested before any codetdd + diagnosing-bugs (Pocock)tdd refuses to test at a seam you have not confirmed; diagnosing-bugs refuses to theorise before a red-capable command exists
already runs the Superpowers workflow (brainstorm, plan, subagents)test-driven-development + systematic-debugging + verification-before-completion, through the Superpowers pluginThe plugin’s session-start bootstrap makes the skills fire without being asked; the three cross-reference each other
wants lightweight skills plus a hard “show me the output” ruletdd + diagnosing-bugs + verification-before-completionPocock’s set had no dedicated completion gate on 2026-09-26; the Superpowers skill fills it and installs on its own
builds a web UI and wants browser checksadd webapp-testing, or agent-browserUnit tests do not prove what the user sees

The two debugging skills differ in their first move. diagnosing-bugs spends most of its effort on one command that goes red on the reported symptom in seconds, then ranks several hypotheses and shows you the list. systematic-debugging starts from evidence (the full error, recent changes, logs at each component boundary), then tests one hypothesis with the smallest possible change. Pick diagnosing-bugs for bugs you can script, systematic-debugging for failures spanning CI, services and configuration.

How do you install the test-and-debug set in Claude Code, Codex and Cursor?

Section titled “How do you install the test-and-debug set in Claude Code, Codex and Cursor?”

The portable route is one skills CLI command (npm skills 1.7.0). Codex and Cursor read the copy in .agents/skills/; Claude Code reads .claude/skills/. The plugin route adds updates and, for Superpowers, the bootstrap that makes the method fire automatically.

  1. Install the skills at project scope. Run this in the repository root.

    Terminal window
    # Option A: plugins from the official marketplace (update themselves)
    claude plugin install mattpocock-skills --scope project
    claude plugin install superpowers@claude-plugins-official --scope project
    # Option B: portable copies of the recommended mix, no plugins
    npx skills add mattpocock/skills --skill tdd codebase-design diagnosing-bugs -a claude-code -y
    npx skills add obra/superpowers --skill verification-before-completion -a claude-code -y

    Plugin skills are namespaced: /mattpocock-skills:tdd, /superpowers:systematic-debugging. Portable copies keep the bare name: /tdd, /diagnosing-bugs. --scope project records the plugin in .claude/settings.json, which you commit. If claude plugin install reports that the marketplace is unknown (for example, in a fresh CI container), run claude plugin marketplace add anthropics/claude-plugins-official first. The official marketplace pins Superpowers 6.4.1; the author’s superpowers-marketplace carries 6.4.2 (both checked 2026-09-26).

    If the team uses more than one agent, install once for all of them: npx skills add mattpocock/skills --skill tdd codebase-design diagnosing-bugs -a claude-code -a codex -a cursor -y. In a test run with skills 1.7.0 on 2026-09-26, that wrote one copy to .agents/skills/ and symlinked it into .claude/skills/; a Claude Code-only install copies straight into .claude/skills/. The Pocock README also asks for setup-matt-pocock-skills; neither tdd nor diagnosing-bugs calls it, so add it only if you use the rest of the set.

  2. Add webapp-testing if you test a web UI. npx skills add anthropics/skills --skill webapp-testing -a claude-code -a codex -a cursor -y, or in Claude Code /plugin marketplace add anthropics/skills then /plugin install example-skills@anthropic-agent-skills (a bundle of 12 skills). The skill writes Python, so the machine needs Python with the playwright package (1.63.0 on PyPI on 2026-09-26) and a browser from python -m playwright install chromium.

  3. Check what landed. npx skills list shows the installed skills. codebase-design must sit next to tdd: tdd calls it when the shape of an interface is in question. In Claude Code, claude plugin details mattpocock-skills prints the plugin’s inventory and projected token cost.

  4. Read what you installed, then commit it. Skills run with the agent’s permissions. Commit .agents/skills/, .claude/skills/ and skills-lock.json (or .claude/settings.json for plugins), so a pull request, not a chat message, changes the team’s set.

What does a systematic-debugging session look like on a real bug report?

Section titled “What does a systematic-debugging session look like on a real bug report?”

The bug report below is the input. The transcript is an abridged illustration of the steps the systematic-debugging skill prescribes. It is not a recorded run, and the repository, commit and numbers are invented for the example.

Issue #318 Renewal date shown one day early for NZ customers
Customers in Auckland see "Renews on 14 March 2027" on the account page.
The card is charged on 15 March. Started after release 3.8.0.

What the agent does, phase by phase:

Phase 1: root cause investigation
Symptom is the displayed date; the charge date is correct.
$ TZ=Pacific/Auckland npx vitest run src/billing/renewal.test.ts
6 passed <- existing tests do not cover the symptom
$ TZ=Pacific/Auckland npx tsx scripts/print-renewal.ts --sub sub_1042
renewsOn: 2027-03-14 (charge job uses 2027-03-15)
$ git log --oneline v3.7.2..v3.8.0 -- src/billing
a41c9e2 replace date-fns addMonths with native Date arithmetic
Phase 2: pattern analysis
Working: src/invoices/period.ts builds dates with Date.UTC(...)
Broken: src/billing/renewal.ts builds new Date(year, month + 1, day) in local time,
then serialises with toISOString()
Phase 3: hypothesis and minimal test
Hypothesis: local midnight in UTC+13 serialises to the previous day in UTC.
Prediction: the same script with TZ=UTC prints 2027-03-15.
$ TZ=UTC npx tsx scripts/print-renewal.ts --sub sub_1042
renewsOn: 2027-03-15 prediction holds
Phase 4: implementation
Failing test first: "renews on the same calendar day in UTC+13"
$ TZ=Pacific/Auckland npx vitest run src/billing/renewal.test.ts
1 failed, 6 passed expected '2027-03-15', received '2027-03-14'
Single fix: addBillingMonths() builds the date with Date.UTC
$ TZ=Pacific/Auckland npx vitest run src/billing/renewal.test.ts
7 passed

That transcript shows what you pay the skill for: the agent proved the existing tests were blind to the symptom, named the commit that introduced the bug, and made the regression test fail with the customer’s exact symptom before the fix turned it green. Had the fix failed, the skill would send the agent back to phase 1; after three failed fixes it stops and asks whether the design is wrong.

For the same bug with diagnosing-bugs, expect a different shape: the agent first builds the TZ=Pacific/Auckland script as its one red-capable command, shrinks the scenario to the smallest input that still fails, then shows you three to five ranked hypotheses before it tests any of them. It tags temporary logs with a prefix such as [DEBUG-a4f2] and removes them with one grep at the end.

How do you run red-green-refactor with verification evidence before “done”?

Section titled “How do you run red-green-refactor with verification evidence before “done”?”

The skills slot into one loop. The loop is the same in all three agents; only the invocation differs. It starts from a bug report or an acceptance criterion and ends with evidence a reviewer can check without reading every line.

  1. Record the baseline. Before any change, the agent runs the full suite, the type check and the linter and saves the result. A test that was already red is recorded as known, so it cannot later hide or pose as a regression.

  2. Agree the seam (Pocock tdd). The agent names the public interface it will test through and waits for your confirmation. For a bug, the seam is where the reported symptom is observable. A seam too shallow to express the bug gives false confidence; diagnosing-bugs treats “no correct seam exists” as a finding in its own right.

  3. Red. One failing test for one behaviour. The agent runs it and shows the failure message, which must be the expected failure (the wrong date), not a typo or a missing import. Superpowers’ skill states the reason: if you did not watch the test fail, you do not know it tests the right thing.

  4. Green. The minimum code that passes. No speculative options, no neighbouring clean-up.

  5. Refactor while green. Superpowers refactors inside the loop and re-runs the tests after each change. Pocock’s tdd moves refactoring to the review stage and its code-review skill. Either way, the tests stay green through every structural change.

  6. Prove the test can catch the bug. Revert the fix, run the new test and watch it fail; restore the fix and watch it pass. verification-before-completion names this red-green cycle as the proof that a regression test works; a test that passes once proves nothing.

  7. Show the evidence, then claim done. The agent runs the full suite, the type check, the linter and, for UI changes, a browser check. It pastes the commands with their exit codes and failure counts, then compares them with the baseline from step 1.

  8. Hand off to CI and review. The pull request carries the evidence; CI re-runs the same commands on clean infrastructure. A red CI run overrides any “done” in the chat.

With plugins, name the skill when you want it for certain: /superpowers:systematic-debugging, /mattpocock-skills:tdd. Add /verify after step 7 when the change is visible in the running app; it builds and drives the app rather than relying on tests. Put the step 1 baseline and step 6 red-green rule in CLAUDE.md, so they hold even when a skill does not fire.

How do you check a web UI with webapp-testing?

Section titled “How do you check a web UI with webapp-testing?”

webapp-testing covers bugs that only show in the browser. It teaches a “reconnaissance-then-action” pattern: wait for the page to settle, take a screenshot or read the DOM, find selectors from what rendered, then act. Its scripts/with_server.py helper starts one or more dev servers, waits for their ports and runs your Playwright script against them.

Once the bug reproduces in the browser, the script becomes the red-capable command for the loop above. When the fix lands, promote it into your end-to-end suite so it runs on every change. If your agents already use agent-browser, stay with it: same CLI in all three agents, no Python. The agent-browser page compares both with the Playwright and Chrome DevTools MCP servers.

How do you prove the skills did their job?

Section titled “How do you prove the skills did their job?”

A skill changes what the agent tends to do. It does not guarantee it: a skill can fail to fire, and an agent can misread its own output. The proof therefore sits outside the agent.

  • CI re-runs the evidence. The commands the agent pasted run again in CI on the pull request. The pasted output is a convenience for the reviewer; the CI result is the gate.
  • The test is protected. A debugging agent under pressure can “fix” a failing test by editing the assertion. Put tests under the deny rules and CODEOWNERS described in protecting the oracle, and fail any pull request that deletes or skips a test without a named approver.
  • The red-green check is in the pull request. The reviewer checks one thing: the new test failed with the reported symptom before the fix and passes after. Attach it to the evidence bundle with the root-cause commit the agent found.
  • The skill set is tested too. Before you standardise a skill for the team, check that it triggers on your prompts. In Claude Code 2.1.283, claude plugin eval runs a plugin’s eval cases against a no-plugin baseline; Anthropic’s skill-creator skill runs test prompts with and without a skill.

Who signs off. The developer who opens the pull request owns the evidence. The reviewer approves on the red-green check, the root cause and a green CI run, and reads code line by line only for high-risk paths (auth, money, schema migrations). The tech lead owns the committed skill set and changes it through review, like any other code.

What do testing and debugging skills cost in context?

Section titled “What do testing and debugging skills cost in context?”

Each installed skill puts its name and description into every session; the body loads only when the skill fires. Measured on 2026-09-26:

What loadsWhenSize
tdd / diagnosing-bugs SKILL.mdWhen the skill fires3,549 / 8,529 bytes
test-driven-development / systematic-debugging SKILL.mdWhen the skill fires9,578 / 9,465 bytes
verification-before-completion / webapp-testing SKILL.mdWhen the skill fires3,646 / 3,913 bytes
superpowers@claude-plugins-official plugin (15 skills)Every sessionabout 838 tokens, claude plugin details in Claude Code 2.1.283
mattpocock-skills@claude-plugins-official plugin (25 skills)Every sessionabout 1,609 tokens, same measurement

A fired debugging skill costs about 2,000–2,500 tokens (roughly four bytes per token). The larger cost is duplicates: the Superpowers plugin plus a portable copy of the same skills lists each one twice. In Claude Code, compare /context before and after an install.

What breaks when agents follow testing and debugging skills?

Section titled “What breaks when agents follow testing and debugging skills?”

Two skills fire on the same bug. With diagnosing-bugs and systematic-debugging both installed, the agent starts one procedure and drifts into the other. Recovery: keep one debugging skill per repository and npx skills remove the other.

The agent writes a test that passes at once. The test was written against the current behaviour, or its assertion recomputes the expected value the way the code does. Pocock’s tdd names that second case “tautological”. Recovery: insist on a red run with the reported symptom in the failure message, and on the revert-the-fix check before any “done”.

The skill does not fire on a terse request. “fix the date thing” may not match a skill’s description. Recovery: name the skill in the prompt, install Superpowers as a plugin so the bootstrap runs, or add a line to CLAUDE.md or AGENTS.md: “For any bug report, use the systematic-debugging skill before proposing a fix.”

The feedback loop is flaky. A test that fails one run in three turns every hypothesis into noise. Recovery: pin the time, seed randomness and isolate the filesystem, as diagnosing-bugs prescribes. For intermittent tests, replace fixed sleeps with condition polling (Superpowers’ condition-based-waiting.md). When a test leaves stray files, use find-polluter.sh from the same skill folder to bisect which test creates them.

The agent loops through fixes. Each fix reveals a new symptom somewhere else. Recovery: stop after the third failed fix, which systematic-debugging requires, and discuss the design with a human before any fourth attempt.

Verification stalls on tests that were red before you started. Without a baseline, the agent cannot tell an old failure from a new regression and keeps retrying. Recovery: record the baseline first (step 1 of the loop) and report old failures separately; never let the agent skip a test file without a written team policy.

Where to go next with testing and debugging skills

Section titled “Where to go next with testing and debugging skills”

For tool-specific debugging workflows, see debugging in the Claude Code CLI, systematic debugging in Codex and AI-assisted debugging in Cursor.