Commit graph

20 commits

Author SHA1 Message Date
dan-claude-bot
8e6f3a4bb8 grant/rehearsal: the codex round — verified rollback, loud partial states, and the raw-attach guarantee measured (#75)
Review 4727756972 (A2): the backout no longer trusts gpasswd — it re-reads
the live group database after removal; verified-absent gets the safe
message, anything else screams ROLLBACK INCOMPLETE, exits nonzero, and
names the exact remediation. The concurrent-login window (a session begun
between usermod and backout keeps the group) is CLOSED to the extent the
database can't reach: the backout detects live processes and names
loginctl terminate-user, and the success wording claims only what was
verified.

Review 4727641752 (A1): a failed grant for a user whose membership predates
the run (the hand-added-user scenario) now fails LOUDLY — they retain
socket access on part-converged policy, and the message says so with both
remediations (box revoke now, or fix and re-run). Their membership is not
stripped: breaking a working user over a failed re-grant is its own hazard.
The default-profile eth0 removal is deliberately not restored on failure —
that mutation only reduces capability, and restoring it would move the
failure state AWAY from fail-closed. Injected-failure coverage is criterion
(n), both flavors: fresh-user backout (fault at the LAST mutation, so the
rollback runs after every earlier one) with the group's absence verified
and a converging re-run; blocked narrowing staged for real with an
instance-local NIC parked on the private bridge.

Review A3, resolution 3 with the measurement demanded: criterion (m)
launches exactly 'incus launch --network boxnet' as the restricted user and
probes the raw NIC from inside — egress works, RFC1918 dropped (the ACL is
the network's), sibling probes dropped BOTH directions (the nft drop is the
host's), name enumeration blocked. The scoped guarantee is now stated in
box-design.md and measured on every run: box-minted instances carry per-NIC
port_isolation; raw attachments keep every network- and host-owned control,
losing only that redundant L2 layer. Instrument lesson kept as MU-5: the
probe's first cut minted the non-cloud image — no DHCP client, no lease,
and a dead NIC passes every negative probe vacuously; it now requires the
lease before believing its own answers.

Rehearsal: 54/54 (containers). test/cli.sh: 82 checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 06:41:05 +00:00
dan-claude-bot
3f9b38ac96 doctor: name an interrupted rehearsal's leftover users instead of absorbing them
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 04:12:33 +00:00
dan-claude-bot
0429a11020 feat: the restricted tier — box grant/revoke converge users onto the hardened boxnet (#74)
incus-user confines an incus-group user to their own project, but its
defaults miss box's contract three measured ways (Debian 13 / Incus 6.0.4):
a private UNHARDENED NAT bridge per user (ipv6.nat=true, no ACL, no DNS
isolation), snapshots blocked, and the box-net profile invisible to their
project. So the tier is an admin-run idempotent convergence:

  box grant <user>   # incus group; touch incus-user (the project is lazy);
                     # drop the private-bridge eth0 from their default
                     # profile; restricted.networks.access=boxnet — and ONLY
                     # boxnet, or the unhardened bridge stays one --network
                     # flag away; restricted.snapshots=allow; install the
                     # shipped box-net profile into their project
  box revoke <user>  # group removal closes the socket, boxes keep running
         --purge     # ...or delete their world, and assert the absence

box_tier() (live credentials, argless id -nG; byte-identical copy in
setup-host.sh) drives the tier-aware surface: new pre-flights the profile
and names the right fix per tier, expose refuses before any daemon call
(without the guard the failure is a lie — restricted certs cannot read
boxnet's redacted config, so box_net_ip claims a running box has no
address), setup-host exits 0 with the honest note, doctor judges only what
the caller can see.

Also fixed while the rehearsal exercised the lifecycle: box restore
dispatched 'incus restore', which does not exist in Incus 6 (it is
'incus snapshot restore') — the verb had never worked. Fixed for every tier.

Convergence survives incus-user restarts by that tool's own design (it
configures a project only at creation) — read in its source, then measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 04:09:48 +00:00
dan-claude-bot
437a3a8e35 test+ci: add CI workflow and a dependency-free test suite
box had no CI and no unit tests — only the live-host drill. Mirror rig's CI:
one `check` job = globstar `shellcheck -x` over bin/* and **/*.sh, then
`bash test/cli.sh`. The suite is dependency-free and runs non-root with no
Incus: the full CLI contract; install.sh's DEST/BINDIR branch driven
functionally against a shim `id` (both tiers + the BOX_HOME/BOX_BIN overrides);
the root-only a+rX and #66's confirm/no-op flow grep-guarded; tmux asserted in
every template. Pre-existing repo shellcheck findings (bin/box SC2034/SC2015/
SC2020, and file-level SC2015 idioms in doctor.sh/wipe.sh/migrate-host.sh) were
resolved — real fixes where behaviour allows, reasoned disables otherwise — so
the new CI is green over the whole repo.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 00:01:15 +00:00
42c1f63c95 fix(doctor): a fresh host is not a dirty one
A bare host got two DIRTYs (the box-to-box drop 'MISSING', the host's
Tailscale resolver) and a 'NOT fit to mint (or to drill)' verdict — and
the very next drill run went 84/84 green from that exact state. Missing
from a stack and never set up are different findings: the network,
profile and ACL sections already knew this; the firewall and resolver
sections now do too. FRESH (no boxnet) downgrades both to information —
the VPN resolver is still named, as a fact about the host that setup-host
pins around, not a fault in a stack that does not exist. The clean
verdict on a fresh host now says what to actually do: run setup-host, or
the drill, which sets the host up itself.

Verified both paths live: standing stack → 'clean', post-teardown →
'fresh' with no DIRTYs.

Also: the measured drill count is 84 (README said 83 — the box-info
exposure check was a NOTE when last counted and is a PASS now).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 00:55:10 +00:00
claude-hdb
455fbc656e feat: host lifecycle as verbs (setup-host/teardown-host/migrate-host) + .box/ convention
Two things:

1. The host scripts are first-class verbs now — 'box setup-host',
   'box teardown-host [--purge-incus]', 'box migrate-host --box <n>'.
   Nobody should have to run ~/.local/share/box/host/<script>.sh; that
   read like an external script and exposed an install path. Each verb
   execs the installed script with its flags passed through (same
   pattern as 'box doctor'). README, doctor hints, and the uninstall
   section point at the verbs now.

2. The repo-runbook convention is '.box/', not '.claudebox/'. Renamed
   across docs and the templates' agent briefing; the briefing tells
   the agent to read either, and box-recipe.md notes the rename, so
   repos still shipping '.claudebox/' keep working through the
   transition. (Consuming repos rename their own folder — tracked
   separately.)

Also widened the help command column for the longer verb names, and
fixed one sed-casualty where a broad '.claudebox'→'.box' pass had
turned the README's legacy user.claudebox tag into user.box.
2026-07-14 18:01:34 +00:00
claude-hdb
4eb6b35a7b chore: finish the debrand — env vars, install dir, docs are 'box', not 'claudebox'
The 0.4.0 rename left surface leftovers the user hit through the
CLAUDEBOX_* env vars. Sweep them, drawing a clean line:

  box       = everything the user touches — env vars (BOX_REPO/REF/HOME/
              BIN), installer messages (box-install:), the install tree
              (~/.local/share/box, with the installer sweeping the old
              ~/.local/share/claudebox on upgrade), tool prose, and the
              docs (docs/box-{design,recipe}.md).
  claudebox = the GitHub repo name (URLs, the claudebox-<ref> tarball
              dir, issue refs), the legacy user.claudebox=1 tag, the
              old-stack cleanup code (claudenet/claude-dev/claude-isolate/
              claudebox-firewall), and the .claudebox/ runbook convention
              — a deliberate v1 hold, since renaming it breaks consuming
              repos.

Renamed the two doc files and their links; updated drill.sh/doctor.sh
paths and BOX_REPO/BOX_REF; fixed the claude template's in-box briefing
to say 'box'. RUNS.md left as-is (append-only history). No behavior
change beyond the install-dir move, which the installer migrates.
2026-07-14 17:44:24 +00:00
claude-hdb
5defd40bca feat!: rename the host stack too, default to blank, drill the templates, add wipe
Follow-up to the rename, per operator direction — the divergence is
reversed and the cut is complete:

- Host stack: boxnet (10.88.0.0/24 — a pre-rename host may still carry
  claudenet on 10.87, two bridges must not claim one subnet),
  box-isolate, nft tables 'inet box'/'bridge box', box-firewall.{sh,
  service}. teardown-host now strips BOTH name generations, so one
  script uninstalls a host of any age.

- Default template is blank: 'box new --name x' mints bare Debian;
  the claude box is '--template claude'. The login hint follows the
  EFFECTIVE template read off the instance, so clones of claude boxes
  still get it and blank boxes are not told to run a binary they lack.

- The drill validates templates: listing, unknown-template refusal,
  the allowlist rejecting BOX_NETWORK by name, and a full blank mint —
  default resolves to blank, metadata stamped, box-net placement, exec
  lands in 'dev', no claude binary, and isolation parity (egress +
  pinned DNS) on the same contract as every template.

- drill/wipe.sh: scorched earth for drill hosts. Both tag generations,
  every drill-named instance, networks/ACLs/profiles/firewall of both
  generations, cached images, and (--purge-storage) the default pool.
  Ends by asserting the ABSENCE of every artifact rather than trusting
  the removals' exit codes.
2026-07-14 14:36:39 +00:00
claude-hdb
c11f3d7552 feat!: claudebox becomes box — the Claude box is one template among several
The tool underneath was already generic: a thin, honest wrapper over
Incus. What was Claude-specific was welded on — one image, one profile,
one cloud-init file, one hardcoded 'sudo -u claude'. The weld is now a
template.

The mechanic: 'box new' stamps the template's identity onto the
instance (user.box=1, user.box.template, user.box.user); shell/exec/
tmux read the user back off the instance, and 'incus copy' carries
user.* keys (audit B2), so a clone knows what it is without consulting
the template. Templates are box.env (parsed against a strict allowlist,
never sourced — no key for a network exists, on purpose) plus a
verbatim cloud-init. Every template launches with the shared box-net
profile: the isolated NIC and root disk, nothing template-controlled —
resources land per-instance from box.env, overridable via BOX_CPU/
BOX_MEMORY/BOX_DISK (which is also how the drill shrinks boxes on a
small host now that profile edits can't).

The three open calls, taken as recommended: clean cut at 0.4.0 (no
claudebox shim; the installer retires the old symlink); default
template = claude (muscle memory survives); repo stays heavy-duty/
claudebox, binary is box.

Compat is the tag, not the name: resolve_box and list honor the legacy
user.claudebox=1 forever, and the legacy tag maps to the claude user —
a pre-rename box lists, shells, clones, unchanged.

Deliberate divergence from #17's table: the host-stack resource names
(claudenet, claude-isolate, nft tables, claudebox-firewall.*) are NOT
renamed — they are host-internal, invisible to users, and renaming
them breaks every provisioned host for zero user-visible gain.
claude-dev is no longer created; setup-host creates box-net, teardown
removes both.

Closes #17
2026-07-14 14:22:50 +00:00
claude-hdb
16a4c129fc feat: 'claudebox doctor' — the host-health checks as a first-class verb
Every fault the drill's doctor diagnoses is a user's fault first: a
wedged Incus daemon, a dnsmasq that silently isn't serving, a VPN
resolver boxes inherit, isolation claimed by config but off in the
kernel — each has killed a cold mint or weakened a boundary, with a
cloud-init error that names none of them. The CLI half-admitted it:
cmd_new's failure path hand-pointed at issue #33, doing one special
case of a doctor's job inline.

The verb delegates to the installed drill/doctor.sh (the tree ships
whole) — one hardened script, two audiences. Its flags pass through;
the verdict now reads 'fit to mint boxes (and to drill)'; the mint-
failure hint ends with 'claudebox doctor' instead of the hand-rolled
diagnosis.

Closes #46
2026-07-14 13:13:18 +00:00
claude-hdb
a2c0758ba6 fix: pin claudenet's resolver — a box's DNS is not a function of the host's VPN
The bridge's dnsmasq forwarded to whatever sat in the host's
/etc/resolv.conf at that moment. On a Tailscale host that is MagicDNS:
box DNS flapped with the tailnet (killing cold mints), and tailnet peer
names and split-DNS zones resolved from inside a box — name-level
reconnaissance of a private network, the same class as the sibling
enumeration dns.mode=none already closes.

setup-host.sh now sets raw.dnsmasq to no-resolv + pinned public
upstreams (BOX_DNS overrides the default 1.1.1.1 8.8.8.8), answering
the issue's three open questions from live measurement: raw.dnsmasq is
the lever (no first-class upstream key on the bridge; verified by
doctor --pin-dns followed by a box resolving), upstreams are a setting
with a sane default, and the pin is unconditional.

The doctor's unpinned-state messages now point at setup-host.sh as the
durable fix, keeping --pin-dns as the quick test.

Closes #33
2026-07-14 12:15:45 +00:00
claude-hdb
2c03624013 fix(doctor): the gateway does not answer ping — by design, so stop asking
claudebox-firewall.sh drops everything from a box to the host except
DNS (53) and DHCP (67); ICMP to 10.87.0.1 dies in that trailing drop on
every healthy host. The doctor used exactly that ping as its routing
probe, so it reported 'cannot even reach the gateway' — and a NOT-fit-
to-drill verdict — on a host whose very next line proved DNS working
through that same gateway.

Probe routing the way the contract states it: a box reaches the public
internet. curl to 1.1.1.1 by address, reusing the one probe for the
DNS-failure diagnosis instead of running it twice.
2026-07-14 11:58:36 +00:00
claude-hdb
2ab3ac8244 fix(doctor): pin the probes' stdin — an interactive exec cannot be timed out
With a TTY on stdin, 'incus exec' goes interactive and puts the terminal
in raw mode. The DNS probes then hang forever: timeout's bare TERM never
lands (no -k escalation), and ^C is forwarded into the box as a keystroke
instead of killing the script. Observed live: doctor hung 15+ minutes at
'Can a box actually resolve DNS?' and survived Ctrl-C; the operator had
to kill the shell.

The drill already learned this exact lesson in #22 (exec_in pins stdin
and uses 'timeout -k'); the doctor's probe section was added later in
#34/#35 and never got the cure. All four probes now pin stdin to
/dev/null and escalate to SIGKILL.
2026-07-14 11:50:08 +00:00
f1748b681d fix(doctor): four checks that lied, and none of them about a real fault
Run the doctor on a healthy host and it reported three problems, all of
them its own:

  · dns.mode=none was flagged as leftover rehearsal dirt. It is now the
    SHIPPED setting — it is what stops a box enumerating its siblings.
    Its absence is the fault; its presence was being "fixed" away.
  · the nft check used 'sudo -n', which fails without cached credentials,
    so it reported the box-to-box rule MISSING on a host where 'sudo nft
    list' plainly shows it.
  · the kernel's bridge view — the one fact that settles the isolation
    question — was skipped with "'bridge' not installed". It is installed;
    it lives in /usr/sbin, which is not on a normal user's PATH.
  · and the DNS probe looked only for the drill's own box names, so it
    said "no box to probe with" while two boxes sat there RUNNING.

A diagnostic that cries wolf is worse than no diagnostic: it costs the
same trust as a real failure and teaches you to ignore it. All four now
report what is actually true.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 02:24:20 +00:00
a0400606c2 fix(doctor): read the isolation off the bridge, not off the config
Two bugs, one in each direction.

setup-host's new assertion ran 'nft list table bridge claudebox' without
sudo. nft needs root, so it failed with permission denied and printed
"the box-to-box drop is NOT active" about a rule that was demonstrably
there. A check that cries wolf is worse than no check.

And the deeper one: every check so far has asked the CONFIG whether
boxes are isolated. The config is a claim. Incus can accept
security.port_isolation and the kernel can still leave 'isolated off' on
the tap — and then boxes reach each other while every config in sight
says they cannot. That is precisely the shape of the original bug: the
ACL looked airtight and never saw the traffic.

So the doctor now reads the kernel's own view — 'bridge -d link show'
on claudenet's ports — and reports the isolated flag as the fact it is.
If the profile says true and the kernel says off, we learn that in a
second instead of after another ten-minute drill.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 02:16:57 +00:00
3bed7c2f2a fix: isolate boxes with the bridge's port-isolation flag, not an nft rule
The nft bridge-family rule from the previous commit is LIVE on the host
and boxes still reach each other:

    table bridge claudebox {
      chain forward { ... meta ibrname "claudenet" meta obrname "claudenet" drop }
    }

    FAIL  BOX A REACHES BOX B — sibling isolation does NOT hold [tcp: refused]

So the rule is not wrong about intent, it is wrong about mechanism —
whatever path these frames take, that hook does not stop them. Rather
than reason harder about netfilter (reasoning is what put the hole there
in the first place), use the mechanism Incus provides for exactly this:
security.port_isolation on the bridged NIC, which sets the kernel bridge
port's isolated flag so two isolated ports cannot exchange frames at all.

The nft rule stays as a second layer — it costs nothing — but the
profile flag is what carries the guarantee. doctor.sh checks it, because
the absence of this one is invisible: everything works and boxes can
simply reach each other.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 01:41:18 +00:00
bd849181bd fix: the firewall unit never re-ran, so new rules were never applied
The box-to-box drop shipped, the drill still found boxes reaching each
other, and the rule was simply not on the host. setup-host.sh ended with
'systemctl enable --now claudebox-firewall.service' — but the unit is
RemainAfterExit, so once it has run it stays "active" forever, and
'--now' does nothing to an active unit. Re-running setup-host after
upgrading claudebox therefore installed the new script to
/usr/local/sbin and never executed it. The host silently kept its old
firewall, and the box-to-box hole stayed open through the release that
claimed to close it.

This is worse than the original bug: every future firewall change would
have landed only on hosts that had never run setup-host before.

Restart the unit instead — the script is idempotent by design. Then
ASSERT the rule is live rather than assume it, because the absence of
this particular rule is invisible: everything keeps working and boxes
can simply reach each other. doctor.sh checks it too.

Also: dns.mode=none is now part of the shipped stack, so the drill must
stop treating it as leftover rehearsal dirt and reverting it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 01:40:02 +00:00
9736977196 fix(doctor): check that something is actually serving the network
Two cold mints died in cloud-init with "Temporary failure resolving
deb.debian.org", on a host the doctor had just certified clean. The
cause was not DNS forwarding, not leftover mutations, and not the drill:
claudenet had NO dnsmasq. It never respawned after the SIGKILL that
recovered the wedged daemon in run 6, so boxes got no DHCP lease at all
— no address, no gateway, no DNS.

Incus does not surface this. The bridge is up, the config is pristine,
'incus network show' says status: Created. Only the process table knows.
So the doctor now asks the process table, and --fix restarts incus to
respawn it.

An hour of hunting and two dead mints went into learning this. It is a
five-second check.

RUNS.md gains trap 11: a network incus calls "Created" may have nothing
serving it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 01:13:38 +00:00
28be7a0ec7 fix: a failed cold mint must say why, and the doctor must find the cause
Two cold mints in a row died with cloud-init 'status: error' on a host
the doctor had just certified clean — so the earlier "leftover mutations
poisoned the network" theory is dead, and the DNS failure is
reproducible rather than transient.

'claudebox new' printed four hundred dots and the word "error", leaving
the user with nothing to act on: the reason was in the box's own log and
nobody was told the log existed. It now prints cloud-init's status, the
fetch/resolve errors from the box's log, and how to inspect the box —
which is left running, because a box that failed to build is evidence,
not garbage. It also names the usual culprit: the host's resolver.

doctor.sh gains the diagnosis that keeps being done by hand:
  · what the HOST resolves through, and whether that is a CGNAT/Tailscale
    resolver the boxes inherit (issue #33);
  · whether claudenet's resolver is pinned;
  · and inside a box, the question that settles it — DNS is broken, but
    can it still reach 1.1.1.1 BY ADDRESS? If yes, egress is fine and the
    fault is purely the inherited forwarder.
  · --pin-dns applies the #33 fix (raw.dnsmasq: no-resolv + public
    servers) so the hypothesis can be TESTED rather than argued.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 00:55:39 +00:00
b0eefd8369 fix(drill): stop poisoning the host, and add a doctor to prove it
Run 8's cold mint failed with 'cloud-init status: error' — the box could
not resolve deb.debian.org, or claude.ai, or anything. The cause was not
in that run at all: run 7's phase D set dns.mode=none on claudenet, the
run ended before reverting it, and every box minted afterwards came up
with no DNS.

This is the worst failure mode the drill has: a poisoned host does not
fail the next run honestly, it produces confident wrong answers. It is
how a false design veto against #16 got posted, and it wasted a cold
mint plus an hour of diagnosis that had nothing to do with the code
under test.

Three defences:
  · the phase-D revert is armed with a trap BEFORE the first mutation,
    so it fires on any exit, Ctrl-C included;
  · the revert is VERIFIED rather than fired into /dev/null, so a failed
    unset can no longer masquerade as a successful one;
  · the drill refuses to start on a host still carrying the mutations.

And drill/doctor.sh answers the question that kept being answered by
hand: what state is this host actually in? Network, profile, ACL,
leftover boxes, and whether a box can still resolve DNS — with --fix to
revert the leftovers.

RUNS.md gains trap 10.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 00:32:03 +00:00