release: box 0.5.0 — codex+grok templates, migrate-host, box expose #56

Merged
dan-claude-bot merged 23 commits from integration/0.5.0 into main 2026-07-15 00:04:54 +00:00

23 commits

Author SHA1 Message Date
b1c430c448 chore(templates): debrand the claude template's description too
Same straggler phrase the README carried — the instance description is
operator-visible in 'incus list', so it follows the rename as well.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:00:15 +00:00
794c81c520 docs: debrand the last project-name references — the repo becomes heavy-duty/box
README title, install URL, and the claude-template blurb now use the
post-rename name; install.sh's REPO default follows (GitHub redirects the
pre-rename URLs, and BOX_REPO still overrides). The two survivors are
literal legacy artifact names that must keep their spelling: the pre-0.4.0
user.claudebox=1 tag and the old claudebox symlink the installer retires.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 23:59:34 +00:00
75acd5c84d docs(readme): 0.5.0 — expose, codex/grok, host verbs, drill at 81/81
The version callout now leads with what 0.5.0 adds (codex+grok templates,
box expose, setup-host/teardown-host/migrate-host as verbs) and keeps
0.4.0's clean-cut terms beneath it. New 'See a dev server' section
documents expose's contract: loopback-only listen, 0.0.0.0 in-box,
per-port, visible in box info. Commands block synced to the actual table
(expose synopsis, host verbs, --remote gone, default template is blank —
the text said claude). Drill paragraph now names the full sweep: every
template cold, the expose door, the legacy re-home.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 23:55:07 +00:00
edf8309f99 fix(expose): accept established flows back from boxnet — the input drop was eating the door's replies
Reproduced the drill's E failure on a live stack (Incus 6.0.4, container
box, same setup-host/box-firewall): tcpdump on boxnet shows the SYN
leaving masqueraded as the gateway and the box answering SYN/ACK
instantly — which then dies at the host's input hook. The inet-box input
chain dropped ALL boxnet input except DNS/DHCP, stateless: the reply to
the very connection the door opened. UFW hosts never had this hole
(before.rules accepts RELATED,ESTABLISHED); the nft fallback now matches
that semantics with a ct state established,related accept ahead of the
drop. Boxes still cannot INITIATE toward the host — a box-originated SYN
is a NEW flow, which is what the drop is for.

Also rebuild the chains on every run (add chain + flush + re-add) instead
of skip-if-present: the existence guard pinned every host to the rule set
of the release that first ran there, so an upgraded rule never landed.

Verified end-to-end on the repro stack: curl 127.0.0.1:18091 → HTTP 200;
box→host initiation still times out; DNS carve-out intact; non-exposed
port still dropped; --remove kills the door; re-expose returns 200.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 23:13:25 +00:00
32bb203ddb fix(expose): install the loopback door's missing half — route_localnet + masquerade on boxnet
The drill's E phase moved one layer down: the device now adds, but
127.0.0.1:<hport> never reaches the box. Incus's NAT-mode proxy installs
only the DNAT (prerouting + output); a loopback-sourced packet then dies
twice — the kernel refuses to route it out a non-loopback interface
without route_localnet on the bridge, and the box would reply to its OWN
127.0.0.1 without a masquerade. This is the exact plumbing Docker
installs on docker0 for '-p 127.0.0.1y'.

box-firewall.sh now sets route_localnet=1 on boxnet and masquerades
loopback-sourced traffic leaving it (chain expose-snat, table inet box).
route_localnet's known risk — 127/8 becomes a routable destination on
the bridge — is covered by the existing iifname-boxnet input drop, which
fires regardless of destination address. The no-UFW guard now checks the
input CHAIN, not the table, since expose-snat shares the table.

expose warns (root-free, via /proc) when the host firewall predates this
plumbing instead of handing over a door that silently does not answer,
and 'box info' now lists open exposures — the drill's nice-to-have: a
box with a hole says so.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 22:55:10 +00:00
44b9d512db fix(expose): pin the boxnet lease as static — NAT proxy resolves connect=0.0.0.0 against ipv4.address, not the lease
The drill's E phase failed with `Instance has no static IPv4 address
assigned to be used as the connect IP`: Incus NAT-mode proxy devices
read the NIC's static ipv4.address device config, never the neighbour
table — the previous comment claimed otherwise. First cut pinned the
wrong address (docker0's), second cut removed the pin instead of
correcting it; this pins the box's current boxnet lease (same address
it already holds) before adding the device, and unpins when the last
exposure is removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 22:31:40 +00:00
claude-hdb
a5d54e4b70 fix(expose): it was pointing the proxy at docker0 — the drill's own oldest trap
Incus finally said it, once the drill stopped swallowing the error:

  Connect IP "172.17.0.1" must be one of the instance's static IPv4 addresses

172.17.0.1 is DOCKER0. box_ipv4() returns whatever Incus lists first, and
a box running docker lists docker0 first — so expose has been aiming the
proxy at the wrong interface all along. This is drill trap 4, verbatim
('docker0 (172.17.x) is the decoy'), which the drill has known since run
4 and the CLI never learned. Now bin/box has its own box_net_ip(): the
address ON boxnet, with the prefix derived from the network rather than
hardcoded.

And the connect address is now the wildcard 0.0.0.0: in NAT mode Incus
resolves the instance's own current address off the bridge's neighbour
table. Naming an address makes it demand a *static* one — the very
demand that produced the error, for an address that was wrong anyway.
Ask Incus for less and it finds the box itself.
2026-07-14 20:45:09 +00:00
claude-hdb
e4b546cd29 fix: expose asks for too much, and the fake legacy box had no claude user
Run 18: 74/2. grok passes (reading the installer worked), migrate's
retire passes. Two left, both mine:

1. box expose — read the Incus proxy docs instead of guessing again.
   Proxy IS supported on VMs, but NAT mode only (correct already), and
   crucially it does NOT need a static address: with no static IP, Incus
   reads the box's address off the bridge's neighbour table and keeps
   the NAT rules in step. My first cut pinned the lease with a device
   override anyway — unnecessary, and almost certainly the step that was
   failing before the proxy was ever added. Dropped it. Ask for less.

   The drill also stopped throwing incus's reason away: it swallowed
   stderr, then RE-RAN the command and printed only the last line ('box:
   expose failed') — box's own words, never incus's. It now captures the
   first attempt and prints all of it.

2. 'legacy box never came up' was unfixable by any timeout. wait_box
   uses 'box exec', which for a legacy-tagged box resolves the user to
   'claude' — and the drill's synthetic legacy box was a BARE image with
   no claude user, so sudo -u claude could never answer, on a box that
   was perfectly healthy (every migration check against it passed). The
   fake legacy box now creates a claude user, like a real pre-0.4.0 box
   had.
2026-07-14 20:30:01 +00:00
claude-hdb
8729c3e522 fix(grok): read the actual installer instead of guessing — the binary is 'grok'
Fetched https://x.ai/cli/install.sh and read it, rather than inferring
the layout from docs. Three facts, every one of which the template had
wrong:

  · the CLI installs as 'grok' (with an 'agent' alias) — NOT 'grok-build'.
    So the drill was checking a command that never existed.
  · BIN_DIR defaults to $HOME/.grok/bin, and what lands there is a
    SYMLINK into the versioned download dir — which is exactly why the
    template's 'find -type f' found nothing.
  · GROK_BIN_DIR can override the directory.

The install was almost certainly succeeding the whole time; the template
was hunting for the wrong name, as the wrong file type, in the wrong
place. Now it links the known path onto the system PATH, asserts
'grok --version' answers, and dumps what the installer actually left if
the upstream layout ever moves.

Also corrected: the drill's version check (grok, not grok-build), the
agent briefing ('grok login'), the template description, and the README
row. The lesson is the repo's oldest one — read the thing, don't reason
about it.
2026-07-14 20:04:10 +00:00
claude-hdb
6f7c3bfd60 fix: run 17's four real findings — migrate retire, expose proxy, wait_box, grok PATH
The first drill run where every failure was the RELEASE CODE, not the
environment. 71 passed, 5 failed; all five traced to four bugs:

1. migrate-host --retire-legacy could NEVER succeed. Re-homing ADDS
   user.box=1 but never removed user.claudebox=1, and legacy_boxes()
   counted the old tag — so retire saw its own freshly-migrated box as
   un-migrated and refused forever ('legacy boxes still exist:
   legacybox'), leaving claudenet + claude-dev behind. Now: a verified
   re-home drops the legacy tag LAST (after the move is proven, so a
   failure anywhere above still leaves the box valid under one tag or
   the other), and legacy_boxes() ignores boxes already carrying
   user.box=1.

2. box expose died with a bare 'could not add the proxy device' — it
   swallowed incus's reason, exactly the sin this repo keeps punishing.
   Now it prints incus's error. And the mechanism is corrected: a VM's
   proxy needs NAT mode, which requires a static NIC address, so expose
   pins the box's current lease first (which also fixes the restart
   caveat — the exposure no longer points at a lease the box may lose).

3. wait_box's 2-minute window was too short: the legacy box was declared
   dead and then every migration check against it passed. 4 minutes.

4. The grok template hunted for a regular file named exactly
   'grok-build' under /home/grok and found nothing — an installer's drop
   may be a SYMLINK, and its binary name is upstream's to choose. Now it
   tries the plausible names and paths, falls back to any executable
   grok*, links both names, and SAYS what it found — or dumps what the
   installer actually left when it finds nothing. The drill likewise
   dumps the on-disk evidence and the cloud-init log on a --version
   failure instead of discarding the box.
2026-07-14 19:56:17 +00:00
claude-hdb
c31fbc63f7 fix: name the boot failure — a corrupt image, Secure Boot, or a GRUB hang
Wall 2 is solved: 'EFI stub: Failed to decompress kernel' was a CORRUPT
IMAGE — the --purge-storage re-download produced a bad blob. Deleting
the cached image and re-pulling booted the box immediately. The storage
pool was innocent (1.29GiB used of 30GiB).

Both walls cost hours to diagnose by hand. The next box to hit them
should be told the answer, not the symptom — so wait_agent now reads
the console log and names the failure:

  · 'Failed to decompress kernel' -> the cached image is corrupt; here
    is the incus image delete command to re-pull it
  · 'bad shim signature' -> Secure Boot rejected the kernel (shouldn't
    happen now; box mints with security.secureboot=false)
  · GRUB/firmware menu -> never booted; re-pull or pin BOX_IMAGE
2026-07-14 19:31:41 +00:00
claude-hdb
e09e4b62ee fix: assert the install landed the ref we asked for — a silent wrong-install is worse than a failure
The 0.5.0 env-var rename (CLAUDEBOX_* -> BOX_*) created a silent trap:
a STALE local drill.sh passes CLAUDEBOX_REPO/REF, today's install.sh
reads BOX_REPO/REF, the vars are ignored, main gets installed — and the
drill runs to a green summary having drilled the wrong tree entirely.
The same class already cost an hour once via a lagged CDN tarball.

install.sh now records what it installed (/INSTALLED_FROM), and the
drill ASSERTS it matches the requested repo@ref before touching the
host — failing loudly, and naming the stale-checkout cause, instead of
drilling a lie.
2026-07-14 19:10:48 +00:00
claude-hdb
455fbc656e feat: host lifecycle as verbs (setup-host/teardown-host/migrate-host) + .box/ convention
Two things:

1. The host scripts are first-class verbs now — 'box setup-host',
   'box teardown-host [--purge-incus]', 'box migrate-host --box <n>'.
   Nobody should have to run ~/.local/share/box/host/<script>.sh; that
   read like an external script and exposed an install path. Each verb
   execs the installed script with its flags passed through (same
   pattern as 'box doctor'). README, doctor hints, and the uninstall
   section point at the verbs now.

2. The repo-runbook convention is '.box/', not '.claudebox/'. Renamed
   across docs and the templates' agent briefing; the briefing tells
   the agent to read either, and box-recipe.md notes the rename, so
   repos still shipping '.claudebox/' keep working through the
   transition. (Consuming repos rename their own folder — tracked
   separately.)

Also widened the help command column for the longer verb names, and
fixed one sed-casualty where a broad '.claudebox'→'.box' pass had
turned the README's legacy user.claudebox tag into user.box.
2026-07-14 18:01:34 +00:00
claude-hdb
4eb6b35a7b chore: finish the debrand — env vars, install dir, docs are 'box', not 'claudebox'
The 0.4.0 rename left surface leftovers the user hit through the
CLAUDEBOX_* env vars. Sweep them, drawing a clean line:

  box       = everything the user touches — env vars (BOX_REPO/REF/HOME/
              BIN), installer messages (box-install:), the install tree
              (~/.local/share/box, with the installer sweeping the old
              ~/.local/share/claudebox on upgrade), tool prose, and the
              docs (docs/box-{design,recipe}.md).
  claudebox = the GitHub repo name (URLs, the claudebox-<ref> tarball
              dir, issue refs), the legacy user.claudebox=1 tag, the
              old-stack cleanup code (claudenet/claude-dev/claude-isolate/
              claudebox-firewall), and the .claudebox/ runbook convention
              — a deliberate v1 hold, since renaming it breaks consuming
              repos.

Renamed the two doc files and their links; updated drill.sh/doctor.sh
paths and BOX_REPO/BOX_REF; fixed the claude template's in-box briefing
to say 'box'. RUNS.md left as-is (append-only history). No behavior
change beyond the install-dir move, which the installer migrates.
2026-07-14 17:44:24 +00:00
claude-hdb
912e0621ca fix: disable Secure Boot on box VMs — 'bad shim signature' hung every mint at GRUB
The console log finally showed the real error behind the GRUB-menu hang:

  error: prohibited by secure boot policy.
  error: bad shim signature.
  Failed to boot both default and fallback entries.

Incus defaults VMs to security.secureboot=true. A Debian cloud image
whose shim is signed with a key this host's OVMF does not trust then
fails signature verification, the kernel never loads, and the VM sits
at the GRUB menu forever — which is exactly the 5-min agent timeout on
every box. It worked in runs 11–15 on the old cached image and broke
the moment --purge-storage re-downloaded a build with a different shim.

security.secureboot=false on VM launch (cmd_new, and the drill's legacy
box). Secure Boot inside a throwaway box is not part of its threat
model — the VM boundary is — and off, it boots reliably across image
rebuilds. Container mode has no firmware and is unaffected.

Bare repro that isolated it: 'incus launch images:debian/13/cloud x --vm'
alone reproduced the hang, proving it was never the 0.5.0 code.
2026-07-14 17:05:39 +00:00
claude-hdb
b9480fd9ce fix: strip whole escape sequences from the console dump, and name the GRUB hang
The first sanitize stripped only the ESC byte, leaving visible '[1m[37m'
halves as noise. Strip full CSI/escape sequences first (while ESC is
present), then residual control bytes — clean text.

And read the log: a box sitting at 'GNU GRUB / Press enter to boot /
UEFI Firmware Settings' never booted — that is the IMAGE, not box. Say
so, and point at re-pulling the image or pinning BOX_IMAGE. Surfaced
running the 0.5.0 drill after --purge-storage re-downloaded a
debian/13/cloud build that hangs at the GRUB menu on the serial console.
2026-07-14 16:56:52 +00:00
claude-hdb
c4cc9f43d1 fix: sanitize the console dump (no more scrambled terminal) + tear down failed mints
Two bugs surfaced running the 0.5.0 drill with --purge-storage on a
cold btrfs pool:

1. wait_agent dumped the VM's RAW console log on timeout — full of
   terminal escape sequences and a firmware menu — which scrambled the
   operator's terminal, and doubly so when it landed in a log they were
   tail -f'ing ('*Debian GNU/Linux', 'ESC to return previous menu',
   ^[^[^[). Now: capture to /tmp/box-console-<n>.log, strip everything
   but printable ASCII + tab/newline, print only a short sanitized tail.
   Nothing raw reaches a terminal.

2. A failed mint left its stuck VM running, starving the NEXT box's boot
   and cascading more 5-min timeouts (tpl failed → codex failed). Every
   mint-failure branch now tears the box down before continuing.
2026-07-14 16:38:54 +00:00
claude-hdb
5deef69621 chore: VERSION → 0.5.0 (templates, migrate-host, expose) 2026-07-14 16:13:08 +00:00
claude-hdb
449d750eb5 Merge remote-tracking branch 'fork/feat/expose' into integration/0.5.0 2026-07-14 16:12:53 +00:00
claude-hdb
05287c96d7 Merge remote-tracking branch 'fork/feat/migrate-host' into integration/0.5.0
# Conflicts:
#	drill/drill.sh
2026-07-14 16:12:53 +00:00
claude-hdb
de6467a728 feat: 'box expose <box> <port>' — a deliberate, loopback-only door to a dev server
The 'no inbound path' contract is one notch too absolute for the tool's
own flagship workflow: coding in a box, a dev server on :3000, and no
way to open it in your browser. expose is the deliberate un-screwing.

- Loopback only, always: the host side listens on 127.0.0.1, never
  0.0.0.0 — no other machine can reach the box; only this host gets a
  door. No flag widens it (that is the escape hatch's job).
- A verb, per-port, reversible, visible: each exposure is a named proxy
  device (expose-<port>); --list and box info show it, --remove undoes
  it. A box with a hole says so.
- Mechanism (VMs): an Incus proxy device forwards host loopback to the
  box's ip:port, plus a SCOPED ingress ACL allow (this box's ip + this
  port only) so the forkproxy's connection survives the default drop —
  the drill decides whether that allow is needed or redundant. The
  in-box server must listen on 0.0.0.0 (a VM's forwarder reaches it over
  the network); inside an isolated box that is safe.

Drill phase E: start a detached listener in a box, expose it, prove the
HOST loopback reaches it, prove a NON-exposed port is still dropped (A7
survives), prove --remove shuts the door.

Closes #55
2026-07-14 16:07:17 +00:00
claude-hdb
c9712834f2 feat: host/migrate-host.sh — re-home legacy boxes onto the new stack, then retire it
The 0.4.0 transition is zero-ceremony (install + setup-host = a
dual-stack host). This script is the two things that path does not do:

- --box <name> / --all-boxes: re-home a pre-rename box onto the new
  stack, PRESERVING its authed state (no re-login). Order is
  load-bearing — tag first (additive, reversible), profile-assign last
  (the network move), then verify the box actually resolves + reaches
  the internet on its 10.88 leg before declaring it migrated. A box
  never ends up tagless or profileless.
- --retire-legacy: remove claudenet/claude-dev/claude-isolate and the
  old firewall unit + nft tables, but REFUSE while any legacy box still
  references them; assert their absence rather than trust exit codes.

Drill phase M builds a faithful legacy stack (claudenet on 10.87, a
claude-dev profile pinned to it, a box on the old tag), then proves:
retire refuses with a legacy box present, re-home flips the tag +
reassigns box-net + lands a 10.88 address + resolves, and retire then
succeeds and leaves nothing. The transition is measured, not asserted.

Closes #53
2026-07-14 16:02:27 +00:00
claude-hdb
c6bb6cb0e0 feat: codex and grok templates — the mechanic's second and third tenants
Two coding-CLI templates mirroring claude's shape: a box.env + verbatim
cloud-init, inheriting the box-net placement contract structurally, no
new design.

- codex: OpenAI Codex CLI via 'npm i -g @openai/codex' (the SCOPED
  package; needs Node 22), symlinked onto the non-interactive exec PATH
  via 'npm prefix -g' — the same PATH fix the claude template needed.
- grok: xAI Grok Build via the official 'curl x.ai/cli/install.sh',
  run AS the grok user (the installer drops into $HOME); the binary is
  found and symlinked to /usr/local/bin.

Install commands verified upstream at implementation time, per the
issue's rule (npmjs.com/package/@openai/codex, x.ai/cli). Each gets an
AGENTS.md-style context file telling the agent it lives in a
disposable, isolated, creds-free box.

Drill: templates listing now expects four; a compact per-template smoke
(mint, '<cli> --version' via box exec, remove) validates each payload
installs and lands on the exec PATH — the generic mechanic is already
proven by blank+claude and not repeated.

Closes #54
2026-07-14 15:59:39 +00:00