On this page you find the problem you are looking at, learn why it happens, and fix it. Each entry has the same three parts: what you see, why, and what to do.
Start with ds machine ready. It proves the machine works, one line per proof, and every failing line names its fix. A full run takes a few minutes, because it starts apps and runs a landing check; ds machine ready --quick starts no app and calls no model. A line marked held (another session or check was using what that proof needed), you (a step only you can take) or note never stops the verdict being READY.
ds verify run is goingds pool or ds work says you do not hold the checkoutWhat you see
is expired or missing, with the command to run.ds command that needs AWS opens the sign-in in your browser when you run it at a terminal. From an agent's shell, where nobody can finish a sign-in, it stops instead and prints the command to run, aws sso login --profile <profile>.ds machine ready fails a line in its aws group, ending sign in: aws sso login --profile <profile>.ds secrets pull warns that it could not refresh from AWS, and writes from its local copy instead.Why
On a person's machine, setup writes a self-renewing AWS sign-in (the sso-session form of an AWS profile). The AWS command line renews it by itself for as long as your AWS IAM Identity Center session lasts. When that session ends, you sign in once more.
Agent machines never sign in: they use certificates instead, and a sign-in can never fix a refused certificate. If the message talks about a machine identity, go to AWS refuses the agent machine identity.
What to do
aws sso login --profile <profile>
devstride-dev unless you already had one for the dev account, and setup's block in your shell startup file exports it as DEVSTRIDE_DEV_PROFILE, so aws sso login --profile "$DEVSTRIDE_DEV_PROFILE" also works. Inside a Claude Code session, type it with a leading ! (! aws sso login --profile <profile>) to run it from the prompt.Always use the command spelled out in full. Some machines have a shell helper called sso; it exists only where someone pasted it into their shell, and it is never the fix to give anyone.
A lapsed sign-in blocks only what reaches AWS. The local Docker app, ds pool, ds work and ds verify need no sign-in, and ds secrets pull keeps writing from its local copy until you sign in again.
Two cases a sign-in will not fix:
is not configured at all. Run ./setup again: on a person's machine it writes the sign-in session and the dev profile to ~/.aws/config. On an agent machine, run ./setup --agent <name>, which enrols the machine and writes its profiles../setup names the one-line change; the repository's README, "AWS SSO setup", shows both forms.What you see
ds pool status ends with Docker is not reachable — containers not surveyed.ds machine ready fails a checkout's database line, ending is Docker running?ds verify cannot reach the test databases. ds verify full says the test Postgres is not accepting connections.Docker has <n> CPUs and <x> GB; the backend tests need at least 4 CPUs and 8 GB — raise them in Docker Desktop: Settings, Resources.Why
Every checkout's app, its database and its test databases run in Docker. On a person's Mac, setup leaves Docker Desktop's own start-up settings as they are, so after a restart Docker may simply not be running. (On an agent machine, setup makes Docker Desktop start whenever someone signs in to the Mac.) The backend tests need at least 4 CPUs and 8 GB of memory given to Docker, which is why setup checks both.
What to do
ds worktree up main starts main's app. A wt1 or wt2 app starts when you take that checkout with ds pool checkout../setup again; it re-checks Docker's CPUs and memory.What you see
ds seed or ds neon says SEED_TOKEN or NEON_API_KEY is not in the agent credentials file, and names the file.ds machine ready fails a line in its credentials group: a checkout's .env is missing keys the bundles deliver, or the agent credentials file lacks a DevStride key. Each line ends with what to run.aws group, the secrets line says nothing has been pulled yet, or that a bundle has a newer version: run: ds secrets pull../setup, under When you can, a note says SEED_TOKEN and NEON_API_KEY are not on this machine yet.Why
Keys are never copied by hand. They live in three bundles in the dev AWS account's Secrets Manager: a shared bundle every machine gets, one machine bundle with a single machine's own stage settings, and an agent bundle with the agents' own service keys. ds secrets pull writes the shared and machine bundles into every checkout's .env, and the agent bundle only into ~/.config/devstride/agent.env, never into a .env.
A key is missing for one of two reasons: this machine has not pulled since the key was added, or nobody has stored it yet.
Today's decision (owner, 2026-10-05): the keys are shared, each with full access, to get everything working. Narrower access per person or per machine comes later.
What to do
ds secrets pull
.env completely, so never edit a .env by hand: the next pull overwrites it.ds machine ready (or ds machine ready --quick). Each credentials line names whatever is still missing.pbpaste | ds secrets set agent <KEY>
ds secrets pull.The Seed and Neon keys. SEED_TOKEN (for ds seed deploy, which redeploys a chosen commit) and NEON_API_KEY (for ds neon, which runs the command-line tool of Neon, our Postgres host) are created by an admin, in Seed's and Neon's organisation settings (For admins). Until they exist, those two commands refuse. Your machine's AWS stage also waits on NEON_API_KEY, so setup lists the stage under Needs you until it arrives; nothing else waits on either key. Seed needs no key for ordinary deploys: a merge to develop or master deploys by itself.
Every tool's credential (where it comes from, which bundle holds it, who creates and rotates it, how ds machine ready proves it) is in the repository's credentials map, .claude/skills/ds-credentials/SKILL.md. Credentials and access walks through it.
What you see
ds pool status says a checkout's data is not recorded (its data could be anything)../setup, under When you can:
demo-data record for wt1 (or wt2) line beginning not recorded: its database already exists, so the hourly check leaves its data alone; ordemo data in wt1 (or wt2) line saying why setup left it as it is: another session holds it, it is on a branch, it has uncommitted changes, or it holds organisations besides the demo data.ds machine ready fails a checkout's demo sign-in, saying acme.devstride could not sign in.ds fleet status marks the checkout unrecorded.Why
Setup builds the demo data once, keeps it as the machine's snapshot, and copies it into each checkout. It only replaces a database it knows is safe to replace: in a checkout nobody is using, holding nothing or only the demo data.
A checkout whose database already existed before setup could record it is not recorded. Its data could be anything, a customer's organisation copied in by hand for example, so nothing resets it on its own: not setup, not the hourly check (ds machine doctor, which keeps idle checkouts current on an agent machine). It waits until a person has looked.
What to do, for wt1 or wt2
export DEVSTRIDE_SESSION="<your name>"
ds pool checkout wt1 --purpose "demo data"
# in a plain terminal, now run the export DEVSTRIDE_LEASE_TOKEN=… line it printed
ds pool reset wt1
ds pool checkin wt1
./setup again, which fills an empty checkout from the snapshot.If the reset says there is no golden template yet, the machine has no snapshot. Run ./setup again: it builds one in a free wt1 or wt2, which takes a few minutes.
For main. Setup fills main's database only while it holds nothing else and nobody holds main, and it never moves main's branch or stops its app. When main needs the demo organisation back, rebuild just that organisation in place:
ds worktree cli main golden build # rebuilds only the Acme demo organisation; other data is untouched
ds worktree cli main golden cognito # or: only re-create the demo logins
This rebuild is what setup, ds machine ready and the hourly check print for main. A reservation would not do: taking main needs --i-know-main-is-free, and handing it back stops its app.
What you see
ds stops with the same message instead of opening a sign-in../setup --agent <name> blocks with AWS refused <name>'s identity (revoked, expired, or its key is gone) — setup never replaces an identity by itself; re-enrol on purpose: ds machine enroll <name> --force, then run setup again.ds machine ready fails its aws lines: the machine's certificate identity was refused or has expired.Why
An agent machine (one that runs sessions with nobody at the keyboard) cannot sign in through a browser, so it has its own certificates: one for the dev AWS account and one for production. AWS exchanges them for short sessions named after the machine. They last a year.
AWS refuses them when the machine was revoked (ds machine revoke), when the certificate has expired, or when the machine's key file is gone. No sign-in can fix that, so nothing offers one.
What to do
ds machine enroll <name> --force
./setup --agent <name>
aws sts get-caller-identity --profile devstride-machine-<name>-dev
Arn it prints (the AWS name of the identity in use) ends in assumed-role/devstride-machine-identity/<name>.When the approver cannot sign: the administrator route. The one approval must come from someone whose AWS sign-in can sign with the management account (where the certificate authority lives). If yours cannot, setup says so, and names this route instead:
dev.csr.pem and prod.csr.pem, in ~/.devstride/machine-identity/<name>/ on the agent machine. They hold no secrets: copy them to the administrator. The machine's private keys (dev.key, prod.key) never leave it.ds machine sign <name> --account dev --csr <full path to dev.csr.pem> --out dev.signed.pem
ds machine sign <name> --account prod --csr <full path to prod.csr.pem> --out prod.signed.pem
./ds runs from the top of the checkout, so give the requests' full paths; the certificates are written to the top of that checkout. It signs with the devstride-mgmt profile unless --mgmt-profile (or DEVSTRIDE_MGMT_PROFILE) names another. The certificate names the machine and account given on the command line, whatever the request says.~/.devstride/machine-identity/<name>/dev.signed.pem and prod.signed.pem, then run ./setup --agent <name> again. It installs them with no sign-in.The repository's ds-machine-identity skill is the full runbook, including how to revoke a machine.
ds verify run is goingWhat you see
ds verify land or ds verify full waits, and about once a minute prints a line ending — one deep check per machine at a time, naming the folder, process and start time of the run it is waiting on.ds work land seems to hang at its local check, which is that same command. Its waiting line goes to .ds/work-land.log, not to the screen.ds machine ready marks its landing check held: another ds verify run is going, then run ds machine ready again once it has finished.Why
A machine runs one deep check at a time. Two full test runs on one machine starve each other of CPU, and the timeouts that follow look exactly like broken code. So every ds verify run takes a machine-wide lock, ~/.cache/devstride/verify.lock, and the next one waits its turn.
What to do
Wait: your run starts as soon as the other finishes. The waiting line tells you whose run it is. There is nothing to clear by hand: a run that died leaves its lock behind, and the next run reclaims it by itself and says so. For ds machine ready, run it again once the other run has finished.
What you see
ds pool status shows a checkout held by machine-setup, with a purpose such as machine setup demo data pid=<n> or machine ready pid=<n>. ds pool checkout on it answers BUSY.Stopped while setup held <checkout>: its lease stays, because a copy may still be finishing inside Docker.ds machine ready marks that checkout held, saying it was left by a stopped ./setup or ds machine ready run and that the next full ds machine ready or ./setup takes it back.Why
Setup and ds machine ready reserve a checkout while they copy demo data into it or start its app. If they are stopped part-way (Ctrl-C, a closed terminal), they deliberately keep the reservation: a copy may still be finishing inside Docker, and handing the checkout over then would give someone a database that is still being rewritten.
What to do
Nothing by hand. Once the stopped process has gone, the next full run takes the reservation back and stops any app it left running: ./setup, a full ds machine ready (not --quick), or on an agent machine the hourly check. It says took back the reservation a stopped setup run left on <checkout>. Never delete the reservation yourself.
The hourly check (ds machine doctor) keeps its own reservation the same way when it is stopped part-way. It prints the exact command that frees the checkout sooner, once nothing is running there, and its next scheduled run takes the checkout back anyway.
What you see
./setup lists a shell startup files line under Needs you:
setup will not edit <file> (it has a "# >>> devstride machine setup >>>" or "# <<< devstride machine setup <<<" line without its pair, or more than one block) — leave one pair of its marker lines, or none, then run setup again
Why
Setup keeps its lines (ds on your PATH, Node, this machine's name and AWS profile) in one block between two marker lines in your shell's startup file, and rewrites only what is between them. When a marker is missing, or there are two blocks, it cannot tell where its block ends: pairing a stray start line with a later end line would delete your own lines in between. So it changes nothing and asks you.
The file depends on your shell: for zsh, ~/.zshrc, plus a second small block in ~/.zshenv, which every zsh reads, scripts included; for bash on a Mac, ~/.bash_profile (or ~/.bash_login or ~/.profile, whichever bash reads); for bash on Linux, ~/.bashrc. The message names the file.
What to do
# >>> devstride machine setup >>> and # <<< devstride machine setup <<<../setup again. It rewrites everything between the pair, or adds a fresh block when there is none.Before its first change to a startup file, setup keeps a copy of it beside the original with .before-devstride-setup added, for example ~/.zshrc.before-devstride-setup. Compare against it to tell your own lines from setup's. Keep your own lines outside the block: setup rewrites the block on every run.
What you see
./setup lists a shell startup files line like this:
setup's shell blocks would make a new terminal run another node (node: v22.11.0 → v24.3.0); they were taken back out and your startup files are as they were — this is a setup defect: tell the team, with this line
Why
After writing its block into your shell's startup files, setup opens a fresh login shell and checks which Node and pnpm it finds. Setup's block must never switch your terminal to a different Node or pnpm than the one it ran before (from nvm, fnm, Homebrew or another version manager), because that could break other projects on your machine. When it would, setup takes its block back out at once, so nothing about your terminal has changed.
What to do
ds on your PATH and the machine's name in new terminals are missing. Use ./ds from inside a checkout until the fix lands..before-devstride-setup (for example ~/.zshrc.before-devstride-setup).What you see
./setup lists the stage under Needs you, or ds machine ready fails a stage: line. The line says which of these it is:
this machine has no Neon key: the team's NEON_API_KEY is not on this machine;holds a database but names no stage: a stage made by hand: this machine had a stage before setup made them;could not look at this machine's stage: the AWS sign-in lapsed or the look failed;Why
Your stage's database is its own Neon branch, made with the team's Neon key, and its first deploy runs from a pool checkout. Setup builds the stage only when everything it needs is in place, and never rebuilds a stage that already exists.
What to do
ds secrets pull and ./setup again.ds stage create --stage <its name> --member wt1../setup again..ds/stage-deploy.log in the checkout it deployed from (setup names it). Fix what it shows, or send it to the team, then run ./setup again: it finishes the stage rather than starting over. ds stage create --member wt1 --dry-run shows what is left to do.What you see
ds machine ready shows a note on stage: database, for example 3 migration(s) behind main's code.
Why
Setup never redeploys your stage because develop moved on. Your checkouts' code moves forward; the stage's database stays where it was until you migrate it.
What to do
From the checkout whose code you will run against the stage, run the command the note gives, for example:
DEVSTRIDE_STAGE=<stage> DEVSTRIDE_REGION=us-east-1 ./ds -b migrations run
What you see
Setup's When you can list, or a ds machine ready note, says Claude in Chrome is offered but not enabled yet.
Why
Setup asks Chrome to offer the Claude extension, but only you can turn an extension on and sign in to it. Agents do their browser work through it, so until it is enabled they hand browser steps back to you.
What to do
Open Chrome. It shows a notice that a new extension, Claude, was added: click Enable. Then open the extension and sign in to Claude. If you dismissed the notice, open chrome://extensions and switch Claude on there. On an agent machine, do this in its DevStride Agent Chrome profile.
What you see
Dozens of test files fail with Test timed out or Hook timed out, or a run takes several times longer than usual. Two signs that the code is not to blame: a different set of files fails on each run, and the failing tests were skipped with no assertions after a setup step timed out, rather than failing an assertion.
Why
Three causes look exactly like broken code:
What to do
docker ps and look for another checkout's devstride-*-drizzle-tests container with a run in progress, and ask the session using it. ds verify already runs one deep check per machine; when two runs genuinely must overlap, cap each with --maxWorkers=7. A high load average alone proves nothing.ds worktree ls names the pair:docker restart devstride-<slug>-drizzle-tests devstride-<slug>-dynamodb-tests
until docker exec devstride-<slug>-drizzle-tests pg_isready -U postgres; do sleep 1; done
Two traps hide the real result:
tail or grep reports the pipe's exit code, so a failed run reads as a success. Write the output to a file instead, from the backend folder: pnpm test:suite:ci:non-golden > run.log 2>&1; echo "GATE_EXIT_CODE=$?".docker rm -f devstride-<slug>-drizzle-tests; the next run recreates it.ds verify land already re-runs, once, up to five files that timed out during setup when no test failed. The full procedure, with measured examples, is in the repository's testing guide: Diagnosing a mass test failure. Running Tests Per Instance explains the per-checkout test databases.
ds pool or ds work says you do not hold the checkoutWhat you see
ds pool reset, ds pool checkin, ds work start or ds work land refuses, saying the checkout is leased under your session name but by a different session, that you do not hold it, and that a lease taken from a plain terminal needs the DEVSTRIDE_LEASE_TOKEN that ds pool checkout printed there.ds work says this checkout is not leased to you, and names the ds pool checkout command to run.Why
A lease records which session took it, not just the name: two sessions can share one name. A Claude Code or Codex session proves it is the holder with its own session id. A plain terminal has no such id, so ds pool checkout handed it a one-time token instead.
What to do
export DEVSTRIDE_LEASE_TOKEN=… line ds pool checkout printed, then try again. Run later commands in that same terminal.ds pool checkout <checkout> --purpose <item>. If that says BUSY, another session holds it: ask that session, or take another checkout.What you see
A git push, git merge, gh pr merge or gh api command is refused before it runs, with a one-sentence reason from the merge guard.
Why
Every Claude Code and Codex session runs the merge guard before each shell command. It refuses the commands that would skip review: a push to develop or master, a forced or deleting push to a release branch, a merge into develop from a branch the merge-path rule does not admit, a merge into master from anything but a release or hotfix branch, and writes to GitHub's branch rules. Develop has the full list.
What to do
ds work land (or /devstride:build-item) for a Story or one-off, a pull request for anything bound for develop, and /devstride:release for production.ds merge-path check --head <branch> --base develop --labels <label1,label2>
mergePath in .claude/ds-config.json), never the guard's switch.DEVSTRIDE_MERGE_GUARD=0 turns the guard off. It is never the way past a refusal: the block is the point. Ask the owner.The health check (ds machine doctor, run hourly on an agent machine and inside every ds pool checkout) fixes what it may and reports the rest. What its messages mean:
| It says | Do this |
|---|---|
on <branch>, … — only reported | A session is working on a branch there. Let it finish and hand the checkout back |
… with uncommitted changes — needs a person or … local commits on no branch — needs a person | The checkout has changes or commits that would be lost. Commit them to a branch, stash them or discard them |
… this code lacks — needs a person | The checkout's database is ahead of its code: it was moved back to older code. Move it forward to develop, or, if it holds only demo data, ./ds pool reset <checkout> |
newer … available — run: ds secrets pull | Run ./ds secrets pull when no session is relying on the current .env. |
every .env still holds the agent's credentials … — run: ds secrets pull | This machine last pulled before agent credentials moved out of .env. One ./ds secrets pull removes them and writes ~/.config/devstride/agent.env |
customer data for org … past the 14-day review | A customer copy is older than 14 days. Drop it with ./ds pool release-client <checkout> once the repro is done; the unattended run does it for you on a checkout nobody holds |
no pool record says this is golden data | Reset it by hand only if that checkout really holds only demo data: ds pool checkout <checkout> --purpose "demo data", then ds pool reset <checkout> and ds pool checkin <checkout>, each with your session's name |
… — run: ds machine setup | Run ./setup from the main checkout (with --agent <name> on an agent machine) |
held by … — only reported | Someone holds it. Nothing to do |
Paths starting with .ds/ are inside a checkout. Setup's and ds machine ready's logs are always in the main checkout, ~/dev/devstride.
| Log | What writes it |
|---|---|
.ds/machine-setup.log (main checkout) | ./setup: every line it printed. A --dry-run writes none |
.ds/machine-setup-data.log (main checkout) | The demo-data build and copies during setup |
.ds/stage-deploy.log | Your AWS stage's first deploy, in the checkout it deployed from |
.ds/machine-ready.log (main checkout) | A full ds machine ready: the apps it started, the landing check, the Codex checks |
.ds/verify-land.log | ds verify land, in the checkout it ran in, every step's whole output. The receipt is .ds/verify-land.json |
.ds/verify-full.log | ds verify full, in the same way. The receipt is .ds/verify-full.json |
.ds/work-land.log | The local check that ds work land ran |
.ds/sst-types.log | Why generating the SST types failed (cli/ensure_sst_types.sh) |
~/.devstride/logs/machine-doctor.log | The hourly check, ds machine doctor |
~/.cache/devstride/merge-guard.log | The merge guard, when it could not run and let a command through |
When you ask for help with a failed setup step, send the line setup printed for that step, .ds/machine-setup.log, and for the demo data .ds/machine-setup-data.log.
Next: back to the Overview
Release
How finished work reaches develop and then production: an Epic's release pull request, the support train, /devstride:release, Seed's deploys, the post-deploy check, and how to tell whether a commit is live.
For admins: access, keys and stages
What an admin does so a developer can set up their machine: grant and remove GitHub and AWS access with one command, add the team's Seed and Neon keys, and look after each machine's AWS stage.