Commit graph

5 commits

Author SHA1 Message Date
b0eefd8369 fix(drill): stop poisoning the host, and add a doctor to prove it
Run 8's cold mint failed with 'cloud-init status: error' — the box could
not resolve deb.debian.org, or claude.ai, or anything. The cause was not
in that run at all: run 7's phase D set dns.mode=none on claudenet, the
run ended before reverting it, and every box minted afterwards came up
with no DNS.

This is the worst failure mode the drill has: a poisoned host does not
fail the next run honestly, it produces confident wrong answers. It is
how a false design veto against #16 got posted, and it wasted a cold
mint plus an hour of diagnosis that had nothing to do with the code
under test.

Three defences:
  · the phase-D revert is armed with a trap BEFORE the first mutation,
    so it fires on any exit, Ctrl-C included;
  · the revert is VERIFIED rather than fired into /dev/null, so a failed
    unset can no longer masquerade as a successful one;
  · the drill refuses to start on a host still carrying the mutations.

And drill/doctor.sh answers the question that kept being answered by
hand: what state is this host actually in? Network, profile, ACL,
leftover boxes, and whether a box can still resolve DNS — with --fix to
revert the leftovers.

RUNS.md gains trap 10.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 00:32:03 +00:00
769287eb7e fix(drill): refuse to judge #16 on a broken baseline
Run 7's box had no network — a clone/source IP collision had taken it
down before a single hardening change was applied. Phase D measured
anyway and reported "L2 filtering BREAKS the box — design veto". It does
not. The box was already broken, and #16 would have been redesigned
around a fiction.

Phase D is now gated on baseline egress passing, and says loudly that it
is skipping rather than quietly producing a verdict. The dns.mode
rejection also stops swallowing incus's error message — the message IS
the finding.

RUNS.md gains trap 9: check that the thing you are measuring WITH still
works before you trust what it tells you. Same failure as the B3 flip,
different costume.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 00:04:01 +00:00
6899fc3626 fix(drill): the NIC inside a VM is enp5s0, not eth0 — read by subnet instead
A3, the one probe the whole audit exists for, has never fired in six
runs. It was never the network: the profile names the DEVICE eth0, but
inside a VM guest predictable naming renames it enp5s0, so every address
lookup — first the '(eth0)' CSV match, then 'ip addr show dev eth0' —
was hunting an interface that does not exist. I fixed that symptom twice
without ever questioning the assumption underneath it.

Read the address from inside the box and select by SUBNET (10.87.x, what
claudenet hands out) rather than by interface name. docker0's 172.17.x
is the decoy; the NIC's name is the guest's business, not ours.

A3 also gains a guard it should have had from the start: if the peer and
the source hold the SAME address, refuse to probe. That is not
hypothetical — clones were inheriting their source's machine-id, hence
its DHCP lease, hence its address, so 'archive → peer' was archive
probing itself and would have reported a cheerful 'reachable' as an
isolation failure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 23:32:22 +00:00
dcc9f001b3 fix(drill): clean the host BEFORE setup-host, and bound it
Run 6 stalled in setup-host.sh, which should be seconds on a host that
already has incus. The ordering was wrong: cleanup ran AFTER setup, so
setup-host reconfigured claudenet's ACLs while an aborted run's boxes
were still attached to that network — 'incus network set' then has to
push the change onto every live NIC. The same aborted run also left the
D-phase mutations (dns.mode=none, NIC filtering) in place, so setup was
converging against a moving target.

Boxes are now deleted and the mutations reverted first, so setup-host is
the no-op it should be. It is also bounded (5 min) and, on timeout,
prints the three things worth checking instead of hanging: instances
still on the network, the firewall unit, the incus daemon. Every cleanup
call gets its own timeout, so a wedged instance cannot stall the run
before the drill has printed a single line.

RUNS.md gains this as trap 8.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 23:09:51 +00:00
a3e5d1226e docs(drill): keep the run log in the repo, not in PR descriptions
Five runs of hard-won knowledge — which probes are answered, which bugs
the drill found in claudebox, and the seven traps the drill itself fell
into — has been living in PR bodies and chat, where the next person
debugging this cannot find it.

RUNS.md carries: the audit scoreboard (A3 still unanswered, and why that
matters), the claudebox findings, a traps section to read BEFORE adding
a probe, stall diagnostics, how to run a single probe by hand instead of
paying for a full run, and the B3 flip-flop as a standing lesson about
verdicts drawn from one observation of a system with restart semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 23:06:38 +00:00