Found by running the drill on a real host, which I could not do before.
The unit is Type=oneshot with no RemainAfterExit, so systemd marks it
'inactive (dead)' the moment ExecStart succeeds. The rules are applied and the
box-to-box drop is live, and the unit still reads as though it died. That is
precisely the question people ask this unit: drill.sh's own failure hint sends
you to 'systemctl status box-firewall.service' to find out whether the firewall
came up, and today the honest answer and the alarming one look identical.
setup-host.sh already believed this was set — 'The unit is RemainAfterExit, so
once it has run it stays "active" forever' — and reasoned from it to explain why
it uses restart instead of 'enable --now'. The reasoning is right and the
restart is right; only the unit was missing the line the comment assumed.
Verified live: before, 'nft list table bridge box' showed the drop present while
is-active said inactive. After, is-active says active (exited) with the drop
still present, and restart still re-applies.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Per @danmt on #66: hatch first, the version-aware migration as its own issue.
Building the host stack from the installer means an upgrade is no longer a tree
swap — it reaches under every box attached to that stack. So the installer now
declines to guess. Same version and ref: it says so and changes nothing.
Version or ref change with boxes on the host: it refuses, lists them, and does
so BEFORE $DEST is touched, so a refusal leaves the working install intact. No
boxes: nothing to lose, proceed. BOX_FORCE_UPGRADE=1 overrides, and the drill
sets it, because arriving on a dirty host and wiping it is the drill's job.
Ref, not just VERSION: a branch and main carry the same VERSION string, so
VERSION alone would call an install of this very branch "unchanged" and skip the
hatch. Both tag generations count as boxes — a pre-rename user.claudebox=1 box
is just as much someone's work as a current one.
The box query runs unprivileged first and escalates only if the socket refuses:
anyone who owns boxes is already in incus-admin, and an installer should not
demand a sudo password merely to look.
The error deliberately does not suggest snapshot -> rm -> restore --from: 'box
rm' deletes a box AND every snapshot it has, so that path loses the data at the
rm. It says to copy anything needed out of the box first. Raised on #67.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Review found two real problems, both confirmed by reproducing them.
setup-host hardcoded 'sudo' for every privileged call, so install.sh's
deliberate root branch — the one that proceeds when id -u is 0 even with no
sudo installed — handed off to a script that died on 'sudo: command not found'
before doing anything (exit 127, reproduced with env -i and a minimal PATH).
The root path was nominal, not real. Privilege is now resolved once: nothing at
UID 0, sudo otherwise, a clear error if neither is possible.
Two things fell out of that. Root does not need incus-admin at all (UID 0 opens
the socket regardless), so adding root to the group was a no-op that also missed
the human — under 'sudo install.sh' that is SUDO_USER, who is now the one
granted the group. And apt must not hang: install.sh runs setup-host with nobody
watching, while a fresh cloud image holds the dpkg lock in apt-daily for its
first minutes, so the calls are now bounded and non-interactive.
The drill did not exercise any of this. It ran setup-host immediately after
install.sh, so the stack existed by the drill's own hand and a run passed
identically whether or not install.sh had done a thing — a fresh run converged
three times while its messages still described the pre-#63 "first pass may only
add you to the group" behaviour. It now asserts the post-install stack in-group,
before the clean or anything else mutates the host, which is the assertion that
actually proves #64. setup-host then runs exactly once more, after the clean —
that one is load-bearing, since the clean deliberately unsets dns.mode and
something has to converge it back. DRILL_OWNS_SETUP=1 hands sequencing back to
the drill. Pre-setup tripwires now read before install.sh, because install.sh is
what triggers setup now; read afterwards they said nothing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
box setup-host stopped halfway when it had to add you to incus-admin: it
usermod'd, printed a NOTE telling you to re-login and re-run, and exited 0 —
a success-shaped no-op with no boxnet, no ACL, no box-net profile and no
firewall behind it. It now re-execs itself under 'sg incus-admin' and
finishes in that same invocation.
The membership check was also asking the wrong question. 'id -nG "$USER"'
names a user, so it reads the group database — which lists incus-admin the
instant usermod returns, while the shell's own credentials still lack it
(supplementary groups are fixed at login). A same-session re-run therefore
passed the check and died further down on a bare permission error from incus
that mentioned neither the group nor the re-login. Argless 'id -nG' asks the
process what it actually holds, which is what incus checks when it opens
/var/lib/incus/unix.socket.
With one run now sufficient, install.sh runs the setup itself instead of
printing a warning and leaving the user a command: the install reported
success and 'box new' then failed on a host with no Incus. setup-host is
idempotent, so doing this on every install is also how an upgraded host picks
up stack changes. BOX_SKIP_SETUP_HOST=1 opts out, and a failed setup leaves
the install standing and says what to re-run.
Fixes#63Fixes#64
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A bare host got two DIRTYs (the box-to-box drop 'MISSING', the host's
Tailscale resolver) and a 'NOT fit to mint (or to drill)' verdict — and
the very next drill run went 84/84 green from that exact state. Missing
from a stack and never set up are different findings: the network,
profile and ACL sections already knew this; the firewall and resolver
sections now do too. FRESH (no boxnet) downgrades both to information —
the VPN resolver is still named, as a fact about the host that setup-host
pins around, not a fault in a stack that does not exist. The clean
verdict on a fresh host now says what to actually do: run setup-host, or
the drill, which sets the host up itself.
Verified both paths live: standing stack → 'clean', post-teardown →
'fresh' with no DIRTYs.
Also: the measured drill count is 84 (README said 83 — the box-info
exposure check was a NOTE when last counted and is a PASS now).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tarball's top dir is <repo>-<ref>, and the extractor globbed
claudebox-* — dead the moment the repo became heavy-duty/box: every
'curl install.sh | bash' and every drill died with 'could not find the
extracted source directory'. The archive has exactly one top-level
directory; take it whatever it is called, and let the existing bin/box
check judge whether it is the right tree. Survives the next rename too.
Verified against the live tarball: BOX_HOME/BOX_BIN scratch install from
heavy-duty/box@main lands and 'box --version' answers 0.5.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The drill section showed the commands but not the step that actually
bites: making sure the checkout you run is the code you mean to judge.
Two versions are in play — the drill script itself, and the (repo, ref)
the drill installs from GitHub and asserts before any verdict. Spell both
out, plus --repo/--ref for drilling a release or a PR branch.
Also catch drill.sh's REPO default up with the rename — it still said
heavy-duty/claudebox (GitHub redirects it, but the default should name
the repo that exists).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Snapshots, ownership-rule, and Isolation sections still named the
pre-0.4.0 stack (claudenet/claude-dev/claude-isolate) while the rest of the
repo — README, host/setup-host.sh, host/box-firewall.sh — uses boxnet/box-net/
box-isolate. Rename the doc references to match ground truth: claudenet→boxnet,
claude-dev→box-net (profile), claude-isolate→box-isolate (ACL). Stale wording
only; the mechanism described is unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
83 checks once the resource-flag assertions land; the '81 passing' claim
belongs to the last green run and the next one re-earns it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Same treatment as the README, applied to the prose that stood in "Claude" for
"the coding agent": box-recipe.md and box-design.md now describe the `.box/`
runbook and creds-free flow around whichever agent the box was minted with, and
name the per-template context file (`~/.claude/CLAUDE.md`, `~/.codex/AGENTS.md`,
`~/.grok/AGENTS.md`) instead of hardcoding Claude's. drill/README.md's
"does not check" note says the drill confirms each template's CLI, not just
Claude Code. `claude` stays as the concrete login example throughout.
Left untouched (out of scope, literal identifiers): the legacy isolation-stack
names (claudenet/claude-dev/claude-isolate, user.claudebox), the claude template
files themselves, and drill/RUNS.md history.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
One entry for now — 0.5.0 as merged plus the inline resource flags that
fold into it. Pre-0.5.0 history stays in git and drill/RUNS.md, which
this points at.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolution is most-specific-first: flag > BOX_* env (kept — it is the
scripting form and how the drill shrinks boxes on small hosts) > the
template's box.env > defaults. Values pass to Incus verbatim (limits.cpu,
limits.memory, root size=) — its units, its validation; box adds no
parser. Resources are all a flag can touch: there is still no flag for a
network or a security.* key, on purpose.
Flags shape a fresh mint only — --from refuses them, a clone carries its
source's resources. An explicit --disk on a container mint gets a note
instead of a silent drop (a container's root rides the pool).
The drill's blank mint now carries --cpu 1 --memory 1GiB and asserts the
limits landed — which is also the precedence proof, since the drill
exports BOX_CPU/BOX_MEMORY on small hosts — plus a negative check that
--from refuses resource flags.
Verified live (container mint, image cached): BOX_CPU=3 + --cpu 1
--memory 1GiB → limits.cpu=1, limits.memory=1GiB; --from + --cpu exits 2
before touching anything; container --disk prints the note.
Closes#57
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The README told the story around Claude Code even though the tool ships
`codex` and `grok` templates and treats every coding agent the same. Reframe
the generic prose — the intro, creds-free line, `.box/` runbook, quick start,
snapshot copy, isolation boundary, recipes, and non-goals — to speak of "the
coding agent" while keeping `claude` as a named, concrete example (it's still
where the project started). No behavior or command changes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Same straggler phrase the README carried — the instance description is
operator-visible in 'incus list', so it follows the rename as well.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
README title, install URL, and the claude-template blurb now use the
post-rename name; install.sh's REPO default follows (GitHub redirects the
pre-rename URLs, and BOX_REPO still overrides). The two survivors are
literal legacy artifact names that must keep their spelling: the pre-0.4.0
user.claudebox=1 tag and the old claudebox symlink the installer retires.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The version callout now leads with what 0.5.0 adds (codex+grok templates,
box expose, setup-host/teardown-host/migrate-host as verbs) and keeps
0.4.0's clean-cut terms beneath it. New 'See a dev server' section
documents expose's contract: loopback-only listen, 0.0.0.0 in-box,
per-port, visible in box info. Commands block synced to the actual table
(expose synopsis, host verbs, --remote gone, default template is blank —
the text said claude). Drill paragraph now names the full sweep: every
template cold, the expose door, the legacy re-home.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reproduced the drill's E failure on a live stack (Incus 6.0.4, container
box, same setup-host/box-firewall): tcpdump on boxnet shows the SYN
leaving masqueraded as the gateway and the box answering SYN/ACK
instantly — which then dies at the host's input hook. The inet-box input
chain dropped ALL boxnet input except DNS/DHCP, stateless: the reply to
the very connection the door opened. UFW hosts never had this hole
(before.rules accepts RELATED,ESTABLISHED); the nft fallback now matches
that semantics with a ct state established,related accept ahead of the
drop. Boxes still cannot INITIATE toward the host — a box-originated SYN
is a NEW flow, which is what the drop is for.
Also rebuild the chains on every run (add chain + flush + re-add) instead
of skip-if-present: the existence guard pinned every host to the rule set
of the release that first ran there, so an upgraded rule never landed.
Verified end-to-end on the repro stack: curl 127.0.0.1:18091 → HTTP 200;
box→host initiation still times out; DNS carve-out intact; non-exposed
port still dropped; --remove kills the door; re-expose returns 200.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The drill's E phase moved one layer down: the device now adds, but
127.0.0.1:<hport> never reaches the box. Incus's NAT-mode proxy installs
only the DNAT (prerouting + output); a loopback-sourced packet then dies
twice — the kernel refuses to route it out a non-loopback interface
without route_localnet on the bridge, and the box would reply to its OWN
127.0.0.1 without a masquerade. This is the exact plumbing Docker
installs on docker0 for '-p 127.0.0.1❌y'.
box-firewall.sh now sets route_localnet=1 on boxnet and masquerades
loopback-sourced traffic leaving it (chain expose-snat, table inet box).
route_localnet's known risk — 127/8 becomes a routable destination on
the bridge — is covered by the existing iifname-boxnet input drop, which
fires regardless of destination address. The no-UFW guard now checks the
input CHAIN, not the table, since expose-snat shares the table.
expose warns (root-free, via /proc) when the host firewall predates this
plumbing instead of handing over a door that silently does not answer,
and 'box info' now lists open exposures — the drill's nice-to-have: a
box with a hole says so.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The drill's E phase failed with `Instance has no static IPv4 address
assigned to be used as the connect IP`: Incus NAT-mode proxy devices
read the NIC's static ipv4.address device config, never the neighbour
table — the previous comment claimed otherwise. First cut pinned the
wrong address (docker0's), second cut removed the pin instead of
correcting it; this pins the box's current boxnet lease (same address
it already holds) before adding the device, and unpins when the last
exposure is removed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Incus finally said it, once the drill stopped swallowing the error:
Connect IP "172.17.0.1" must be one of the instance's static IPv4 addresses
172.17.0.1 is DOCKER0. box_ipv4() returns whatever Incus lists first, and
a box running docker lists docker0 first — so expose has been aiming the
proxy at the wrong interface all along. This is drill trap 4, verbatim
('docker0 (172.17.x) is the decoy'), which the drill has known since run
4 and the CLI never learned. Now bin/box has its own box_net_ip(): the
address ON boxnet, with the prefix derived from the network rather than
hardcoded.
And the connect address is now the wildcard 0.0.0.0: in NAT mode Incus
resolves the instance's own current address off the bridge's neighbour
table. Naming an address makes it demand a *static* one — the very
demand that produced the error, for an address that was wrong anyway.
Ask Incus for less and it finds the box itself.
Run 18: 74/2. grok passes (reading the installer worked), migrate's
retire passes. Two left, both mine:
1. box expose — read the Incus proxy docs instead of guessing again.
Proxy IS supported on VMs, but NAT mode only (correct already), and
crucially it does NOT need a static address: with no static IP, Incus
reads the box's address off the bridge's neighbour table and keeps
the NAT rules in step. My first cut pinned the lease with a device
override anyway — unnecessary, and almost certainly the step that was
failing before the proxy was ever added. Dropped it. Ask for less.
The drill also stopped throwing incus's reason away: it swallowed
stderr, then RE-RAN the command and printed only the last line ('box:
expose failed') — box's own words, never incus's. It now captures the
first attempt and prints all of it.
2. 'legacy box never came up' was unfixable by any timeout. wait_box
uses 'box exec', which for a legacy-tagged box resolves the user to
'claude' — and the drill's synthetic legacy box was a BARE image with
no claude user, so sudo -u claude could never answer, on a box that
was perfectly healthy (every migration check against it passed). The
fake legacy box now creates a claude user, like a real pre-0.4.0 box
had.
Fetched https://x.ai/cli/install.sh and read it, rather than inferring
the layout from docs. Three facts, every one of which the template had
wrong:
· the CLI installs as 'grok' (with an 'agent' alias) — NOT 'grok-build'.
So the drill was checking a command that never existed.
· BIN_DIR defaults to $HOME/.grok/bin, and what lands there is a
SYMLINK into the versioned download dir — which is exactly why the
template's 'find -type f' found nothing.
· GROK_BIN_DIR can override the directory.
The install was almost certainly succeeding the whole time; the template
was hunting for the wrong name, as the wrong file type, in the wrong
place. Now it links the known path onto the system PATH, asserts
'grok --version' answers, and dumps what the installer actually left if
the upstream layout ever moves.
Also corrected: the drill's version check (grok, not grok-build), the
agent briefing ('grok login'), the template description, and the README
row. The lesson is the repo's oldest one — read the thing, don't reason
about it.
The first drill run where every failure was the RELEASE CODE, not the
environment. 71 passed, 5 failed; all five traced to four bugs:
1. migrate-host --retire-legacy could NEVER succeed. Re-homing ADDS
user.box=1 but never removed user.claudebox=1, and legacy_boxes()
counted the old tag — so retire saw its own freshly-migrated box as
un-migrated and refused forever ('legacy boxes still exist:
legacybox'), leaving claudenet + claude-dev behind. Now: a verified
re-home drops the legacy tag LAST (after the move is proven, so a
failure anywhere above still leaves the box valid under one tag or
the other), and legacy_boxes() ignores boxes already carrying
user.box=1.
2. box expose died with a bare 'could not add the proxy device' — it
swallowed incus's reason, exactly the sin this repo keeps punishing.
Now it prints incus's error. And the mechanism is corrected: a VM's
proxy needs NAT mode, which requires a static NIC address, so expose
pins the box's current lease first (which also fixes the restart
caveat — the exposure no longer points at a lease the box may lose).
3. wait_box's 2-minute window was too short: the legacy box was declared
dead and then every migration check against it passed. 4 minutes.
4. The grok template hunted for a regular file named exactly
'grok-build' under /home/grok and found nothing — an installer's drop
may be a SYMLINK, and its binary name is upstream's to choose. Now it
tries the plausible names and paths, falls back to any executable
grok*, links both names, and SAYS what it found — or dumps what the
installer actually left when it finds nothing. The drill likewise
dumps the on-disk evidence and the cloud-init log on a --version
failure instead of discarding the box.
Wall 2 is solved: 'EFI stub: Failed to decompress kernel' was a CORRUPT
IMAGE — the --purge-storage re-download produced a bad blob. Deleting
the cached image and re-pulling booted the box immediately. The storage
pool was innocent (1.29GiB used of 30GiB).
Both walls cost hours to diagnose by hand. The next box to hit them
should be told the answer, not the symptom — so wait_agent now reads
the console log and names the failure:
· 'Failed to decompress kernel' -> the cached image is corrupt; here
is the incus image delete command to re-pull it
· 'bad shim signature' -> Secure Boot rejected the kernel (shouldn't
happen now; box mints with security.secureboot=false)
· GRUB/firmware menu -> never booted; re-pull or pin BOX_IMAGE
The 0.5.0 env-var rename (CLAUDEBOX_* -> BOX_*) created a silent trap:
a STALE local drill.sh passes CLAUDEBOX_REPO/REF, today's install.sh
reads BOX_REPO/REF, the vars are ignored, main gets installed — and the
drill runs to a green summary having drilled the wrong tree entirely.
The same class already cost an hour once via a lagged CDN tarball.
install.sh now records what it installed (/INSTALLED_FROM), and the
drill ASSERTS it matches the requested repo@ref before touching the
host — failing loudly, and naming the stale-checkout cause, instead of
drilling a lie.
Two things:
1. The host scripts are first-class verbs now — 'box setup-host',
'box teardown-host [--purge-incus]', 'box migrate-host --box <n>'.
Nobody should have to run ~/.local/share/box/host/<script>.sh; that
read like an external script and exposed an install path. Each verb
execs the installed script with its flags passed through (same
pattern as 'box doctor'). README, doctor hints, and the uninstall
section point at the verbs now.
2. The repo-runbook convention is '.box/', not '.claudebox/'. Renamed
across docs and the templates' agent briefing; the briefing tells
the agent to read either, and box-recipe.md notes the rename, so
repos still shipping '.claudebox/' keep working through the
transition. (Consuming repos rename their own folder — tracked
separately.)
Also widened the help command column for the longer verb names, and
fixed one sed-casualty where a broad '.claudebox'→'.box' pass had
turned the README's legacy user.claudebox tag into user.box.
The 0.4.0 rename left surface leftovers the user hit through the
CLAUDEBOX_* env vars. Sweep them, drawing a clean line:
box = everything the user touches — env vars (BOX_REPO/REF/HOME/
BIN), installer messages (box-install:), the install tree
(~/.local/share/box, with the installer sweeping the old
~/.local/share/claudebox on upgrade), tool prose, and the
docs (docs/box-{design,recipe}.md).
claudebox = the GitHub repo name (URLs, the claudebox-<ref> tarball
dir, issue refs), the legacy user.claudebox=1 tag, the
old-stack cleanup code (claudenet/claude-dev/claude-isolate/
claudebox-firewall), and the .claudebox/ runbook convention
— a deliberate v1 hold, since renaming it breaks consuming
repos.
Renamed the two doc files and their links; updated drill.sh/doctor.sh
paths and BOX_REPO/BOX_REF; fixed the claude template's in-box briefing
to say 'box'. RUNS.md left as-is (append-only history). No behavior
change beyond the install-dir move, which the installer migrates.
The console log finally showed the real error behind the GRUB-menu hang:
error: prohibited by secure boot policy.
error: bad shim signature.
Failed to boot both default and fallback entries.
Incus defaults VMs to security.secureboot=true. A Debian cloud image
whose shim is signed with a key this host's OVMF does not trust then
fails signature verification, the kernel never loads, and the VM sits
at the GRUB menu forever — which is exactly the 5-min agent timeout on
every box. It worked in runs 11–15 on the old cached image and broke
the moment --purge-storage re-downloaded a build with a different shim.
security.secureboot=false on VM launch (cmd_new, and the drill's legacy
box). Secure Boot inside a throwaway box is not part of its threat
model — the VM boundary is — and off, it boots reliably across image
rebuilds. Container mode has no firmware and is unaffected.
Bare repro that isolated it: 'incus launch images:debian/13/cloud x --vm'
alone reproduced the hang, proving it was never the 0.5.0 code.
The first sanitize stripped only the ESC byte, leaving visible '[1m[37m'
halves as noise. Strip full CSI/escape sequences first (while ESC is
present), then residual control bytes — clean text.
And read the log: a box sitting at 'GNU GRUB / Press enter to boot /
UEFI Firmware Settings' never booted — that is the IMAGE, not box. Say
so, and point at re-pulling the image or pinning BOX_IMAGE. Surfaced
running the 0.5.0 drill after --purge-storage re-downloaded a
debian/13/cloud build that hangs at the GRUB menu on the serial console.
Two bugs surfaced running the 0.5.0 drill with --purge-storage on a
cold btrfs pool:
1. wait_agent dumped the VM's RAW console log on timeout — full of
terminal escape sequences and a firmware menu — which scrambled the
operator's terminal, and doubly so when it landed in a log they were
tail -f'ing ('*Debian GNU/Linux', 'ESC to return previous menu',
^[^[^[). Now: capture to /tmp/box-console-<n>.log, strip everything
but printable ASCII + tab/newline, print only a short sanitized tail.
Nothing raw reaches a terminal.
2. A failed mint left its stuck VM running, starving the NEXT box's boot
and cascading more 5-min timeouts (tpl failed → codex failed). Every
mint-failure branch now tears the box down before continuing.
The 'no inbound path' contract is one notch too absolute for the tool's
own flagship workflow: coding in a box, a dev server on :3000, and no
way to open it in your browser. expose is the deliberate un-screwing.
- Loopback only, always: the host side listens on 127.0.0.1, never
0.0.0.0 — no other machine can reach the box; only this host gets a
door. No flag widens it (that is the escape hatch's job).
- A verb, per-port, reversible, visible: each exposure is a named proxy
device (expose-<port>); --list and box info show it, --remove undoes
it. A box with a hole says so.
- Mechanism (VMs): an Incus proxy device forwards host loopback to the
box's ip:port, plus a SCOPED ingress ACL allow (this box's ip + this
port only) so the forkproxy's connection survives the default drop —
the drill decides whether that allow is needed or redundant. The
in-box server must listen on 0.0.0.0 (a VM's forwarder reaches it over
the network); inside an isolated box that is safe.
Drill phase E: start a detached listener in a box, expose it, prove the
HOST loopback reaches it, prove a NON-exposed port is still dropped (A7
survives), prove --remove shuts the door.
Closes#55
The 0.4.0 transition is zero-ceremony (install + setup-host = a
dual-stack host). This script is the two things that path does not do:
- --box <name> / --all-boxes: re-home a pre-rename box onto the new
stack, PRESERVING its authed state (no re-login). Order is
load-bearing — tag first (additive, reversible), profile-assign last
(the network move), then verify the box actually resolves + reaches
the internet on its 10.88 leg before declaring it migrated. A box
never ends up tagless or profileless.
- --retire-legacy: remove claudenet/claude-dev/claude-isolate and the
old firewall unit + nft tables, but REFUSE while any legacy box still
references them; assert their absence rather than trust exit codes.
Drill phase M builds a faithful legacy stack (claudenet on 10.87, a
claude-dev profile pinned to it, a box on the old tag), then proves:
retire refuses with a legacy box present, re-home flips the tag +
reassigns box-net + lands a 10.88 address + resolves, and retire then
succeeds and leaves nothing. The transition is measured, not asserted.
Closes#53
Two coding-CLI templates mirroring claude's shape: a box.env + verbatim
cloud-init, inheriting the box-net placement contract structurally, no
new design.
- codex: OpenAI Codex CLI via 'npm i -g @openai/codex' (the SCOPED
package; needs Node 22), symlinked onto the non-interactive exec PATH
via 'npm prefix -g' — the same PATH fix the claude template needed.
- grok: xAI Grok Build via the official 'curl x.ai/cli/install.sh',
run AS the grok user (the installer drops into $HOME); the binary is
found and symlinked to /usr/local/bin.
Install commands verified upstream at implementation time, per the
issue's rule (npmjs.com/package/@openai/codex, x.ai/cli). Each gets an
AGENTS.md-style context file telling the agent it lives in a
disposable, isolated, creds-free box.
Drill: templates listing now expects four; a compact per-template smoke
(mint, '<cli> --version' via box exec, remove) validates each payload
installs and lands on the exec PATH — the generic mechanic is already
proven by blank+claude and not repeated.
Closes#54
Operator watched /tmp/new.log through a claude mint and saw not one
message: cloud-init's progress dots are block-buffered the moment
stdout is not a tty, so a redirected mint shows nothing for the whole
install and then one burst — which reads exactly like a hang, on the
very night three real hangs happened.
PYTHONUNBUFFERED=1 on the cloud-init wait makes the dots arrive as
dots; box new also prints how to watch the box's own full narration
(incus exec <box> -- tail -f /var/log/cloud-init-output.log); and the
drill's logs are named for the box being minted (/tmp/mint-drill.log),
not for the verb that mints it.
Run 14: the blank mint — the FIRST VM launch on the fresh btrfs pool,
which unpacks the image into a pool volume and takes the coldest boot —
died at wait_agent's 3-minute window ('Processes: -1' well past it),
while the identical claude mint sixty seconds later booted in the warm
path and passed in 96s. The window was tuned on a dir pool with a
cached, unpacked image.
150×2s now, and on failure box new prints the VM's console log tail
before dying — this run's evidence was torn down with the box before
anyone could read it.
All four mints (blank, claude, clone, peer) now run through mint_box:
box new's narration lands in the log as before, the drill prints where
to tail it, and a dot every 5s on the drill's own terminal proves the
run is alive. A silent multi-minute mint is indistinguishable from a
wedge, and that ambiguity has cost whole evenings — the operator said
so, verbatim.
Run 14's second catch: the blank mint's cloud-init finished, printed
'status: done' — and 'box new' hung for 15+ minutes on an exec session
that never closed ('incus operation list' showed it still RUNNING).
With a TTY on stdin (the drill redirects only stdout/stderr), incus
exec goes interactive, and the session can wedge open after the remote
command has exited. Same disease as drill trap 2 and doctor trap 13;
the CLI's own execs never got the cure.
Every non-interactive exec now pins stdin: wait_agent's probe, the
cloud-init wait, both failure-path reads, and the clone identity
reset. shell/exec/tmux keep the terminal — owning it is their job.
Also: wipe.sh keeps cached images on a plain wipe. An image is
upstream's artifact, content-addressed by fingerprint — deleting it
buys zero cleanliness and costs the next mint a full re-download. It
goes only with --purge-storage, where the pool it lives in goes too
(and it must go first: images block pool deletion).
The profile rename changed the file's name, header and limits but not
the device's 'network:' field; the claudenet→boxnet sed covered
host/*.sh only. On a wiped host (no claudenet to silently latch onto)
'incus profile edit box-net' refused the YAML and setup died — run 14's
first catch, before a single box was minted. The sweep this fix rode in
on found exactly one other stale reference, in the same file's comment.
Follow-up to the rename, per operator direction — the divergence is
reversed and the cut is complete:
- Host stack: boxnet (10.88.0.0/24 — a pre-rename host may still carry
claudenet on 10.87, two bridges must not claim one subnet),
box-isolate, nft tables 'inet box'/'bridge box', box-firewall.{sh,
service}. teardown-host now strips BOTH name generations, so one
script uninstalls a host of any age.
- Default template is blank: 'box new --name x' mints bare Debian;
the claude box is '--template claude'. The login hint follows the
EFFECTIVE template read off the instance, so clones of claude boxes
still get it and blank boxes are not told to run a binary they lack.
- The drill validates templates: listing, unknown-template refusal,
the allowlist rejecting BOX_NETWORK by name, and a full blank mint —
default resolves to blank, metadata stamped, box-net placement, exec
lands in 'dev', no claude binary, and isolation parity (egress +
pinned DNS) on the same contract as every template.
- drill/wipe.sh: scorched earth for drill hosts. Both tag generations,
every drill-named instance, networks/ACLs/profiles/firewall of both
generations, cached images, and (--purge-storage) the default pool.
Ends by asserting the ABSENCE of every artifact rather than trusting
the removals' exit codes.
The tool underneath was already generic: a thin, honest wrapper over
Incus. What was Claude-specific was welded on — one image, one profile,
one cloud-init file, one hardcoded 'sudo -u claude'. The weld is now a
template.
The mechanic: 'box new' stamps the template's identity onto the
instance (user.box=1, user.box.template, user.box.user); shell/exec/
tmux read the user back off the instance, and 'incus copy' carries
user.* keys (audit B2), so a clone knows what it is without consulting
the template. Templates are box.env (parsed against a strict allowlist,
never sourced — no key for a network exists, on purpose) plus a
verbatim cloud-init. Every template launches with the shared box-net
profile: the isolated NIC and root disk, nothing template-controlled —
resources land per-instance from box.env, overridable via BOX_CPU/
BOX_MEMORY/BOX_DISK (which is also how the drill shrinks boxes on a
small host now that profile edits can't).
The three open calls, taken as recommended: clean cut at 0.4.0 (no
claudebox shim; the installer retires the old symlink); default
template = claude (muscle memory survives); repo stays heavy-duty/
claudebox, binary is box.
Compat is the tag, not the name: resolve_box and list honor the legacy
user.claudebox=1 forever, and the legacy tag maps to the claude user —
a pre-rename box lists, shells, clones, unchanged.
Deliberate divergence from #17's table: the host-stack resource names
(claudenet, claude-isolate, nft tables, claudebox-firewall.*) are NOT
renamed — they are host-internal, invisible to users, and renaming
them breaks every provisioned host for zero user-visible gain.
claude-dev is no longer created; setup-host creates box-net, teardown
removes both.
Closes#17