box/drill
claude-bot-andresmgsl ce3a0c5076 Make setup-host privilege-aware; make the drill prove the new contract
Review found two real problems, both confirmed by reproducing them.

setup-host hardcoded 'sudo' for every privileged call, so install.sh's
deliberate root branch — the one that proceeds when id -u is 0 even with no
sudo installed — handed off to a script that died on 'sudo: command not found'
before doing anything (exit 127, reproduced with env -i and a minimal PATH).
The root path was nominal, not real. Privilege is now resolved once: nothing at
UID 0, sudo otherwise, a clear error if neither is possible.

Two things fell out of that. Root does not need incus-admin at all (UID 0 opens
the socket regardless), so adding root to the group was a no-op that also missed
the human — under 'sudo install.sh' that is SUDO_USER, who is now the one
granted the group. And apt must not hang: install.sh runs setup-host with nobody
watching, while a fresh cloud image holds the dpkg lock in apt-daily for its
first minutes, so the calls are now bounded and non-interactive.

The drill did not exercise any of this. It ran setup-host immediately after
install.sh, so the stack existed by the drill's own hand and a run passed
identically whether or not install.sh had done a thing — a fresh run converged
three times while its messages still described the pre-#63 "first pass may only
add you to the group" behaviour. It now asserts the post-install stack in-group,
before the clean or anything else mutates the host, which is the assertion that
actually proves #64. setup-host then runs exactly once more, after the clean —
that one is load-bearing, since the clean deliberately unsets dns.mode and
something has to converge it back. DRILL_OWNS_SETUP=1 hands sequencing back to
the drill. Pre-setup tripwires now read before install.sh, because install.sh is
what triggers setup now; read afterwards they said nothing.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 13:16:06 +00:00
..
doctor.sh fix(doctor): a fresh host is not a dirty one 2026-07-15 00:55:10 +00:00
drill.sh Make setup-host privilege-aware; make the drill prove the new contract 2026-07-17 13:16:06 +00:00
README.md docs: extend the agent-agnostic reframe to design, recipe, and drill docs 2026-07-15 00:21:22 +00:00
RUNS.md docs(drill): record runs 11–13 — the contract measured at zero, from a bare host 2026-07-14 12:52:02 +00:00
wipe.sh fix: pin stdin on every non-interactive exec in the CLI — a mint wedged at 'status: done' 2026-07-14 15:07:01 +00:00

The drill

An end-to-end rehearsal of box against a real Incus: install the CLI, set up the host, mint boxes, drive the whole surface, check that the isolation actually holds — and run the full #15 audit, including a live rehearsal of the hardening #16 proposes. It ends with a block of audit answers to paste into #15.

It rearranges the host it runs on. Incus, a systemd unit, a network, an ACL, a profile, rewritten firewall rules — and, in the last phase, deliberate mutations to the network and profile. Run it on a machine you can format — a spare server, a cloud VM you'll destroy, a VM on your laptop. Not your workstation.

git clone https://github.com/heavy-duty/claudebox && cd claudebox
bash drill/drill.sh --yes      # run and forget; omit --yes to be asked first

Useful flags: --ref <branch> (drill a branch rather than main), --keep-boxes (leave the boxes up to poke at — note the last phase's network and profile mutations stay applied with them).

Exit 0 means every check passed. Roughly 20 minutes, most of it the cold box.

Something wrong with the host? bash drill/doctor.sh — it reports whether the host is fit to drill (network, profile, ACL, leftover boxes, whether a box can still resolve DNS), and --fix reverts what an aborted run left behind. The drill mutates the host in phase D; an aborted run can leave a network that mints boxes with no DNS.

Iterating on the drill? Read RUNS.md first — it is the run log: what the audit has answered so far, the bugs the drill has found in box, the traps this script has already fallen into (every one cost a run), how to diagnose a stall, and how to run a single probe by hand instead of paying for a whole run.

Why it exists

The repo has no tests and no CI, and the CLI is a shell script that shells out to incus. That means the interesting failures are not in the bash — they are in what Incus actually does, which is exactly what unit tests would stub out and get wrong. The drill runs the real thing.

What it checks

A. Incus semantics. The assumptions box is built on, probed directly: that incus config get <inst> user.claudebox returns 1 (this is on the path of every box command — if it lies, everything fails closed); that the user.claudebox=1 list filter selects our instances and excludes an untagged one; that --columns nstS gives four clean CSV fields; that the state column reads RUNNING; that incus rename really does refuse a running instance; that snapshot-list's first CSV field is the label; that an unset config key reads as empty with exit 0 (#15 B4); and that incus copy preserves user.* keys (#15 B2 — the whole template-metadata design in #17 rests on it).

B. The surface. Mint, list, info, snapshot, clone-from-a-snapshot-of-a- renamed-box, rename (running must refuse, stopped must work), the escape hatch and its isolation warning, the rm confirmation guard, and the CLI contract (typo'd command, typo'd flag, list <box>).

The boundary gets its own treatment: the drill launches an instance box did not mint, aims down, rm and the escape hatch at it, and requires all three to refuse — and the instance to still be standing afterwards.

C. Isolation baseline (#15 section A). From inside a real box: public egress works; the box cannot reach a listener on the host's claudenet gateway; RFC1918 is dropped; a sibling box is unreachable (a listener runs on the peer so "refused" — the packet arrived — cannot masquerade as "dropped"); whether DNS enumerates the sibling is recorded (#12 predicts it leaks today; that is audit data, not a failure); IPv6 is off; and the host cannot connect into a box.

D. Hardening rehearsal (#15 section B). The host is disposable, so the drill applies the exact changes #16 proposes and watches what breaks: dns.mode=none (must kill sibling resolution, must not kill egress), security.mac_filtering + security.ipv4_filtering on the NIC (the box and its in-box Docker must keep working), and @internal as an ACL drop destination on a bridge network (if accepted and egress survives, #16's sibling drop is renumber-proof by construction; if not, #16 derives the subnet instead). A FAIL in this phase is a design veto for #16, caught before the code is written.

What it does not check

agent login (e.g. claude /login) — it's interactive by design, and the box is creds-free by design. The drill confirms each coding-agent template's CLI is installed and runnable; authenticating is yours.

If the host has no /dev/kvm, box falls back to container mode. The drill still runs, but it says loudly that the VM trust boundary was not validated rather than passing quietly on a weaker one.