Skip to content

Keeping a team current when the tools ship weekly

Keeping a team current with AI coding agents takes three routines: a rotating weekly release watch that sorts each change into break, default shift, candidate or noise; a graduation path that admits a new feature to the shared harness only after an evaluated trial; and team docs that stamp every tool fact with a version, channel and checked date.

Claude Code published 52 releases between 4 August and 25 September 2026, and Codex published 21 between 7 August and 26 September (npm registry, read 26 September 2026). In the same weeks, Codex removed codex exec --full-auto, Claude Code changed its default model, and Claude Code extended auto mode as the starting permission mode to every plan and provider on its latest channel. Your team’s AGENTS.md still names a removed flag, two engineers are on a different channel from CI, and the one person who noticed forwarded a changelog link. This page is for the tech lead who has to turn that into a routine that takes an hour a week.

What a release-watch routine gives your team

Section titled “What a release-watch routine gives your team”
  • A rota with a 45-minute time box, a fixed source list per tool and a digest file the whole team reads.
  • A four-way triage table that says what to do with each change and by when.
  • A graduation path with a decision record that keeps untested features out of the shared harness.
  • Version pinning per tool, so developers and CI run the version you tested.
  • A date-stamping convention for team docs, plus two CI checks that fail when a fact is stale or a retired flag is still in use.
  • Three copy-paste prompts: triage the week’s releases, design a trial, and audit the harness for drift.

The five working days from 21 to 25 September 2026 show the pattern:

DateTool and versionChangeWhat it meant for a team
2026-09-22Claude Code v2.1.280 (latest)Claude Opus 5.5 became the default model on every paid planCost and behaviour of every default session changed; evals run on the old default no longer describe the tool
2026-09-23Claude Code v2.1.281 (latest)AGENTS.md read when a project has no CLAUDE.md, now also on Bedrock, Google Cloud, Foundry and gatewaysClaude Code sessions on those providers started reading instructions from a repository that has only AGENTS.md
2026-09-23Codex CLI 0.156.1GPT-6 Sol and GPT-6 Luna in the model pickerA cheaper option to trial for routine loops; no change for anyone who does nothing
2026-09-25Claude Code v2.1.283 (latest)Auto mode became the starting permission mode for interactive terminal and VS Code sessions on every plan and providerEnterprise, API and cloud-provider sessions (Pro/Max/Team already since 14 August) started in auto mode unless settings disable it; claude -p and unsupported models still start in Manual

On 26 September 2026 the Claude Code stable channel was still v2.1.274 (npm dist-tags), so none of the three Claude Code rows applied to a team on stable. That split is the first thing the rota has to control. The models hub has the current model lineup and prices; this page links there rather than repeating them.

The rota puts one named person on the watch each week, rotating through the team. Rotation spreads the knowledge and survives holidays.

  1. Name the watcher and the slot. One engineer per week, in a fixed 45-minute slot on Monday morning, listed in the sprint plan as a ticket. Pair a junior with a senior for their first two turns.
  2. Fix the source list. The watcher reads only the sources below, from the team’s pinned version to the newest release.
  3. Triage every change into one of the four classes in the next section. Anything unclear goes in as a candidate with a question, not as noise.
  4. Write the digest to docs/agents/release-watch/<year>-W<week>.md in the repository and open a pull request. The digest is short: breaks, default shifts, candidates, and a one-line count of noise.
  5. Hand off. Breaks get an owner and a fix date before the slot ends. Candidates go to the tech lead, who picks at most one per fortnight for a trial.
SourceWhat to readHow to read it
Claude CodeChangelog and the channel versionsnpm view @anthropic-ai/claude-code dist-tags in a terminal; /release-notes inside a session
CodexGitHub releases for openai/codex and the bundled model catalognpm view @openai/codex version; codex features list; codex debug models --bundled
CursorCursor changelogBrowser; record the date you read up to, not a version
ModelsClaude model deprecations and the Codex catalog aboveAny retirement date inside the next 90 days is a break
StandardsMCP blog (2026-07-28 spec release post)Only when a spec version lands; see what changed in MCP 2026-07-28

Diff what each tool says about itself instead of trusting a summary: snapshot its surface when you pin a version, then diff the new version against it.

Terminal window
# Terminal: which versions are on each channel today
npm view @anthropic-ai/claude-code dist-tags
# Snapshot the flags and subcommands of the pinned version, then diff after an upgrade
claude --version > docs/agents/snapshots/claude-version.txt
claude --help > docs/agents/snapshots/claude-help.txt
# after upgrading on the watcher's machine only:
claude --help | diff docs/agents/snapshots/claude-help.txt - || true

Inside a session, /release-notes shows the notes for recent versions. Read every entry between the team’s pinned version and the target, not only the newest one — Claude Code ships roughly one release per day.

Triage each change into one of four classes

Section titled “Triage each change into one of four classes”

The triage class decides the deadline. Without it, every change is equally urgent, which in practice means none of them is.

ClassTestActionDeadline and owner
BreakRemoves or renames a flag, command, setting, model or file our harness usesFix pull request to the harness; add the old term to the retired-terms listBefore anyone upgrades; the watcher, reviewed by the tech lead
Default shiftChanges a default the team relies on: model, effort, permission mode, instruction file, sandboxRun the team’s eval set on the new version before changing the pin; tell the team what changes for themBefore the pin moves; the tech lead decides
CandidateNew capability that could replace manual work or a home-grown scriptLog it; the tech lead picks at most one per fortnight for a trialNext trial slot; a volunteer
NoiseFixes and features that touch nothing the team usesCount it in the digestNone

Teams miss default shifts most, because nothing fails. When Codex made GPT-6 Astra its bundled default on 4 September 2026 (CLI 0.153.4), a team whose config.toml set no model got a different model the day they upgraded. Tests passed; cost and behaviour moved. Treat a default shift as a change the team makes on purpose: run the evals, then move the pin. The new-model playbook covers the model case in detail.

Graduate a feature into the shared harness

Section titled “Graduate a feature into the shared harness”

The shared harness is everything the team’s agents run inside: the instruction files, settings, hooks, skills, plugins and MCP servers that live in the repository or in managed policy. The harness overview describes the layers. A feature enters the harness only through the path below, because a change to the harness changes every engineer’s agent at once.

  1. Observed. The watcher logs the feature as a candidate in the digest, with its release note and the minimum version it needs.
  2. Trial. One volunteer uses the feature on one loop for at most two weeks, in a worktree or branch, with the team’s normal permission mode. Before starting, they write down the hypothesis and the eval: which recent tickets or katas they will re-run, and what result counts as better. The eval result decides, not the volunteer’s impression.
  3. Harness pull request. If the trial passes, the volunteer opens one pull request that adds the feature to the shared harness: a line in CLAUDE.md or AGENTS.md, a key in .claude/settings.json or .codex/config.toml, a skill, or a plugin in the team marketplace. The pull request carries the decision record below, a version gate (“needs Claude Code v2.1.283 or later, latest channel”) and a one-line rollback. CODEOWNERS on the harness paths routes it to the tech lead.
  4. Teach. The volunteer spends ten minutes on it in the next review-to-learn session from upskilling a team for agentic engineering. A feature that nobody else knows how to use did not graduate.
  5. Review at 90 days. The decision record has a review date. On that date the owner keeps, changes or retires the feature, and says which in the digest.

Keep the decision record in the pull request description and in docs/agents/decisions/:

docs/agents/decisions/2026-10-codex-gpt-6-sol-for-dependency-bumps.yaml
feature: "GPT-6 Sol for the dependency-bump loop in Codex"
observed_in: "docs/agents/release-watch/2026-W39.md"
needs: "Codex CLI 0.156.1 or later"
hypothesis: "Sol passes the same dependency-bump evals as the default model at lower usage"
trial:
owner: "m.wisniewska"
loop: "weekly dependency bumps, services/billing"
window: "2026-09-28 to 2026-10-09"
eval: "re-run the last 8 merged dependency-bump tickets; hidden tests must pass on all 8"
result: "8 of 8 passed; review findings unchanged" # filled in at the end of the trial
decision: adopt # adopt | reject | extend-trial
harness_change: ".github/workflows/deps.yml: codex exec -m gpt-6-sol"
rollback: "remove -m gpt-6-sol from the deps job"
approved_by: "tech-lead-billing"
review_on: 2027-01-09

A feature the team cannot turn off in one line does not enter the harness at all. The tooling roadmap is where the tech lead’s rejected and extended trials feed the organisation’s capability bets.

Pin versions so the team and CI run what you tested

Section titled “Pin versions so the team and CI run what you tested”

Graduation means little if half the team runs a different version. Pin the version the eval set passed on, everywhere, and move the pin as a deliberate change.

Claude Code has two release channels. latest gets every release; stable is “typically about a week behind” and skips releases with major regressions (Claude Code setup docs, checked 26 September 2026). Choose one for the whole team.

{
"autoUpdatesChannel": "stable",
"requiredMinimumVersion": "2.1.274"
}

Put this in managed settings, not in each person’s user settings, so it applies to everyone; enforcing one policy across every coding agent covers delivery. The settings schema in Claude Code 2.1.283 describes requiredMinimumVersion and requiredMaximumVersion as enforced only from managed settings: a version outside the range exits at startup with instructions to install an approved one. requiredMinimumVersion only stops machines that are behind; add requiredMaximumVersion as well if you want auto-updates to stop at the version your evals passed, which makes every pin move a managed-settings change.

On a single machine, claude install stable or claude install 2.1.274 installs a channel or an exact version. In CI, install the exact version the team pinned, for example npm install -g @anthropic-ai/claude-code@2.1.274 for the stable pin above, and pin the model with the model setting if your evals depend on it.

Date-stamp the team’s agent docs the way this site does

Section titled “Date-stamp the team’s agent docs the way this site does”

This site keeps every model name, price, default and flag in one facts sheet, each row with its source and checked date. Four of its rules transfer directly to a team’s own agent docs:

  1. Every fact carries a checked date and a source. A fact is true on its checked date only.
  2. Every version-gated claim names its version and channel. Write “from v2.1.283 (latest channel)”, never “Claude Code now…”.
  3. Every negative claim carries the version it was checked against. Write “Codex CLI 0.157.1 has no --full-auto flag (checked 26 September 2026)”, never “Codex has no --full-auto”.
  4. Retired terms go on a never-write list with what replaces them, so a stale instruction is caught by a machine, not by a confused new hire.

In a team repository, rules 1 to 3 become one table and rule 4 becomes one text file. Keep the table in docs/agents/tool-facts.md:

| Fact | Tool and version | Source | Checked | Re-check by |
| --- | --- | --- | --- | --- |
| Default model is Claude Opus 5.5 on paid plans and the Anthropic API (not Foundry) | Claude Code v2.1.280+ (`latest`) | https://code.claude.com/docs/en/model-config.md | 2026-09-26 | 2026-10-26 |
| `-a` accepts only `on-request` and `never` | Codex CLI 0.157.1 | `codex --help` | 2026-09-26 | 2026-10-26 |
| Team pin: Claude Code `stable`, Codex 0.157.1 | Team decision | docs/agents/decisions/ | 2026-09-26 | 2026-10-26 |

Keep the never-write list in docs/agents/retired-terms.txt, one extended regular expression per line, with a comment saying what replaces it:

# One extended regex per line. The comment above each says what replaces it,
# since which version, and when the team checked it.
# Codex: `codex exec --full-auto` removed in 0.147.0 (2026-08-07); use a permission profile.
--full-auto
# Codex: cannot run as an MCP server since 0.154.0 (2026-09-09); use codex exec or the Codex SDK.
codex mcp-server
# Codex: -a accepts only on-request and never (checked on 0.157.1, 2026-09-26).
(-a|--ask-for-approval)[ =](on-failure|untrusted)
# Claude Code: /ultraplan left the command reference in v2.1.222; use /plan.
/ultraplan
# Cursor: Background Agents are now Cloud Agents (name checked 2026-08-28).
Background Agents?

Two scripts turn both files into gates. The first fails when a shared agent file still uses a retired term:

scripts/check-harness-drift.sh
#!/usr/bin/env bash
# Fails when a shared agent file still uses a term the team has retired.
set -euo pipefail
terms=docs/agents/retired-terms.txt
mapfile -t files < <(git ls-files -- '*CLAUDE.md' '*AGENTS.md' '.claude/*' '.codex/*' '.cursor/*' '.github/workflows/*')
[ "${#files[@]}" -eq 0 ] && exit 0
if grep -nE -f <(grep -vE '^[[:space:]]*(#|$)' "$terms") -- "${files[@]}"; then
echo "Retired agent terms found. See $terms for what replaces each one." >&2
exit 1
fi
echo "No retired agent terms in ${#files[@]} shared agent files."

The second fails when a fact is past its re-check date:

scripts/check-tool-facts.mjs
// Fails when a row in docs/agents/tool-facts.md is past its re-check date.
import { readFileSync } from 'node:fs';
const today = new Date().toISOString().slice(0, 10);
const rows = readFileSync('docs/agents/tool-facts.md', 'utf8')
.split('\n')
.filter((line) => /^\|.*\|\s*\d{4}-\d{2}-\d{2}\s*\|\s*$/.test(line));
const stale = rows.filter((line) => {
const cells = line.split('|').map((c) => c.trim()).filter(Boolean);
return cells.at(-1) < today; // last column: Re-check by (ISO dates compare as strings)
});
for (const line of stale) console.error(`STALE: ${line}`);
console.log(`${rows.length} dated facts, ${stale.length} past their re-check date.`);
process.exit(stale.length ? 1 : 0);

Run both on every pull request and once a week before the watch, with a read-only token:

.github/workflows/harness-freshness.yml
name: harness-freshness
on:
pull_request:
schedule:
- cron: '17 6 * * 1' # Mondays 06:17 UTC, before the release watch
permissions:
contents: read
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- run: bash scripts/check-harness-drift.sh
- run: node scripts/check-tool-facts.mjs

The scheduled run matters most: a fact expires on a date, not on a commit, so the Monday failure lands in the watcher’s slot.

How do you know the release watch is working?

Section titled “How do you know the release watch is working?”

These measures show whether the routine changes anything; each has an owner.

MeasureDefinitionTarget and owner
Break lead timeDays from a break’s release date to the merged harness fix, for releases newer than the team’s pinFixed before the pin moves, every time; the watcher of that week
Unplanned breaksBreaks the team discovered from a failure instead of from the digestZero per quarter; reviewed by the tech lead in the retrospective
Drift gate failuresFailures of check-harness-drift.sh and check-tool-facts.mjs on the Monday runFixed within the watch slot; the watcher
Trial yieldShare of trials in the quarter that ended in a decision record with an eval resultEvery trial; one that ends without a result is a failed trial, not a pass
Pin skewPeople or CI jobs on a Claude Code channel or Codex version other than the team pin, from claude --version and codex --version output collected in the watchZero; the tech lead

The sign-offs are fixed. The watcher signs the digest. The tech lead approves every harness pull request, every pin move and every default-shift decision. The platform owner or CTO approves anything that changes managed policy for more than one team; who owns the harness, the gates and the evals sets that boundary.

How does this answer the Tech Lead Scorecard’s question on keeping current?

Section titled “How does this answer the Tech Lead Scorecard’s question on keeping current?”

Question 21 of the Tech Lead Scorecard asks how you keep the team current as tools change, and its top answer is “owned radar + sandbox + rollout decisions”. On this page the rota is the owned radar, the two-week trial is the sandbox, and the triage classes, pins and decision records are the rollout decisions.

What breaks when a team tries to keep current

Section titled “What breaks when a team tries to keep current”

The watch becomes changelog forwarding. The watcher pastes links into the team channel, and nobody acts on them. To recover, require the digest file in a pull request and require an owner and a date for every break before the slot ends.

Auto-update breaks CI on a Tuesday. A CI job installs the newest CLI, a flag it uses was removed, and every agent job fails. To recover, pin an exact version in CI, add the flag to retired-terms.txt, and move the CI pin only in the same pull request that moves the team pin.

Half the team is on a different channel. Some engineers installed from the latest channel and some through a package manager that tracks stable. On 26 September 2026 that meant one group had Opus 5.5 as the default and read AGENTS.md, while the other ran Opus 5 or Sonnet 5 and ignored AGENTS.md; on Enterprise, API and cloud-provider seats only the latest group started in auto mode. Review findings stop being comparable. To recover, set autoUpdatesChannel and requiredMinimumVersion in managed settings and check pin skew weekly.

A default shift goes unnoticed. Nothing fails, but cost or behaviour changes after an upgrade or a server-side model default change. To recover, re-run the eval set on every pin move, and set the model explicitly for any loop whose evals depend on it.

An enthusiast merges a feature straight into the shared harness. A new hook or skill lands in .claude/ because it worked once for one person. To recover, add CODEOWNERS on the harness paths, revert to the last approved state, and send the feature through a trial. Shared hooks governance and shared skills cover the review bar for each type.

The rota lapses. Two holidays and a deadline later, nobody has read a changelog for a month. To recover, keep the watch as a sprint ticket with a named owner, let the Monday check-tool-facts.mjs run fail loudly, and catch up by reading from the pinned version forward rather than skimming the last week.

Frequently asked questions

How does a team keep up with AI coding tools that release every week?

Give one person a rotating, time-boxed weekly release watch. They read each tool's changelog since the team's pinned version, sort every change into break, default shift, candidate or noise, and write a short digest. Breaks are fixed that week; candidates go through a trial before they reach the shared harness.

How does a new agent feature get into the team's shared setup?

Through five stages: observed in the digest, trialled by one volunteer on one loop against the team's own eval set, merged into the shared harness in a reviewed pull request with a version gate and a rollback line, taught in the next review session, and reviewed again after 90 days.

Should the team run the stable or latest channel of Claude Code?

Pick one and pin it for everyone, including CI. On 26 September 2026 the latest channel was v2.1.283 and stable was v2.1.274, and several behaviours (Opus 5.5 as the default, reading AGENTS.md, and auto mode as the starting mode on Enterprise, API and cloud-provider seats) existed only on latest. A mixed team runs two different agents.

How do you stop team agent docs from going stale?

Stamp every tool fact with the version, channel and date it was checked, give each row a re-check date, and run two CI checks: one fails when a fact is past its re-check date, the other fails when a shared agent file still uses a flag or command the team has retired.