fix: boxes could reach each other — isolate them at the bridge

A live probe (drill run 10) found box A's SYN arriving at box B, and B
answering with a RST. Boxes were not isolated from each other at all,
and the README's contract — "a box reaches the public internet and
nothing else" — was false.

The ACL was not wrong; it simply never saw the traffic. Two boxes on one
bridge share an L2 segment, so their frames are SWITCHED between bridge
ports and never traverse the netfilter path where an L3 rule lives. The
drop on 10.0.0.0/8 (which contains claudenet) and the default ingress
drop both looked airtight and neither ever fired. This is why the
original reasoning — "belt and braces" — was plausible and wrong.

The bridge family does see it. Its forward hook fires exactly when a
frame passes from one bridge port to another, which on claudenet means
box→box and nothing else: frames for the gateway are delivered locally,
and so is anything routed out to the internet. Dropping every forwarded
frame on the bridge isolates the boxes and costs them nothing — DHCP and
ARP are unaffected, being broadcast and delivered on INPUT.

Also: dns.mode=none, so a box can no longer ENUMERATE its siblings
through the gateway's dnsmasq. Blocked connections with open
reconnaissance is not isolation.

security.ipv4_filtering is deliberately NOT used: it breaks the box's
networking (dockerd comes up but cannot pull or run a container).

The drill now ASSERTS all of this in phase C against the real stack;
phase D's rehearsal is retired, its findings recorded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
claude-hdb 2026-07-14 01:29:33 +00:00
parent 08b6b2ce90
commit dbefd66dd5
4 changed files with 106 additions and 122 deletions

View file

@ -90,12 +90,39 @@ you cannot aim it at an instance claudebox didn't mint. If the command can move
the box off the isolation stack (profile, network, device, `security.*`), it the box off the isolation stack (profile, network, device, `security.*`), it
warns and proceeds — from there the trust boundary is yours to keep. warns and proceeds — from there the trust boundary is yours to keep.
## Isolation (unchanged) ## Isolation
Dedicated NAT bridge `claudenet` + Incus `claude-isolate` ACL dropping all Dedicated NAT bridge `claudenet` + Incus `claude-isolate` ACL dropping all
RFC1918/CGNAT/link-local egress, plus host-firewall rules blocking instance → RFC1918/CGNAT/link-local egress, plus host-firewall rules blocking instance →
host. The instance reaches the internet and nothing else. Entry is `incus exec` host. Entry is `incus exec` over the local socket — no inbound path. The VM is
over the local socket — no inbound path. The VM is the trust boundary. the trust boundary.
**A box reaches the public internet and nothing else — including no other box.**
That last clause is the one that was assumed and turned out to be false, so it
is spelled out here with the mechanism, and `drill/` tests it on every run.
- **Box → host, LAN, RFC1918, CGNAT, link-local:** the `claude-isolate` ACL.
- **Box → box: an nftables *bridge-family* rule** (`host/claudebox-firewall.sh`).
It cannot be an ACL rule. Two boxes on one bridge share an L2 segment, so
their frames are *switched* between bridge ports and never traverse the
netfilter path an L3 ACL lives on — the ACL looked airtight (it drops
`10.0.0.0/8`, which contains `claudenet`) while box→box was in fact wide open.
A live probe found box A's SYN arriving at box B. The bridge family's forward
hook fires exactly on port-to-port frames, which on this bridge means box→box
and nothing else: gateway traffic and routed egress are delivered locally, not
forwarded. Dropping every forwarded frame therefore isolates the boxes and
costs them nothing.
- **Box → box by NAME:** `dns.mode=none`. dnsmasq on the gateway held a record
for every instance, so a box could enumerate its siblings even where it could
not reach them. Blocked connections with open reconnaissance is not isolation.
- **IPv6:** off (`ipv6.address=none`), and that is a *contract*, not a default —
every rule above is IPv4-only, so IPv6 would be an uncovered path.
- **`security.ipv4_filtering`: deliberately NOT used.** It breaks the box's
networking (in-box Docker cannot pull or run a container). Tested, vetoed.
The rule that keeps this honest: **isolation claims are tested, never reasoned
about.** The box→box hole existed because a plausible code reading said it could
not. See `drill/RUNS.md`.
## Non-goals ## Non-goals

View file

@ -519,16 +519,16 @@ else
aud "A3 sibling: NOT PROBED (no 10.87.x address on peer)" aud "A3 sibling: NOT PROBED (no 10.87.x address on peer)"
fi fi
# C5 — DNS enumeration (#15 A4): #12 predicts this LEAKS today. Either way it # C5 — DNS enumeration (#15 A4). Now a CONTRACT, not an observation: setup-host
# is audit data, not a code failure — #16's dns.mode=none is the fix. # sets dns.mode=none precisely so a box cannot enumerate its siblings.
e1="$(in_box archive getent hosts peer)" e1="$(in_box archive getent hosts peer)"
e2="$(in_box archive getent hosts peer.incus)" e2="$(in_box archive getent hosts peer.incus)"
if [ -n "$e1$e2" ]; then if [ -n "$e1$e2" ]; then
note "DNS enumeration leaks, as #12 predicted: a box resolves its sibling ($(printf '%s' "$e1$e2" | head -1 | cut -c1-50))" no "a box can still RESOLVE its sibling ($(printf '%s' "$e1$e2" | head -1 | cut -c1-40)) — dns.mode=none is not holding"
aud "A4 dns enumeration: LEAKS (bare='${e1:+yes}' fqdn='${e2:+yes}') — #16's dns.mode=none earns its place" aud "A4 dns enumeration: LEAKS — the fix is not in effect"
else else
note "DNS enumeration did NOT leak — #16's dns.mode item shrinks to an assertion" ok "a box cannot resolve its sibling's name (no DNS enumeration)"
aud "A4 dns enumeration: no leak observed" aud "A4 dns enumeration: blocked (dns.mode=none)"
fi fi
# C6 — IPv6 off (#15 A6): every ACL rule is IPv4-only; off is the only cover. # C6 — IPv6 off (#15 A6): every ACL rule is IPv4-only; off is the only cover.
@ -558,123 +558,37 @@ else
aud "A7 inbound host→box: NOT PROBED" aud "A7 inbound host→box: NOT PROBED"
fi fi
# The mutations below OUTLIVE the run if it dies: dns.mode=none leaves every
# box minted afterwards with NO DNS at all (cloud-init then fails with
# "Temporary failure resolving deb.debian.org"), and NIC filtering can stop a
# box getting on the network. A poisoned host does not fail the NEXT run
# honestly — it produces confident, wrong answers, which is how a false design
# veto against #16 got posted. So arm the revert BEFORE making the first
# mutation, and let it fire on any exit, including Ctrl-C.
revert_phase_d() {
incus network unset claudenet dns.mode >/dev/null 2>&1
incus profile device unset claude-dev eth0 security.mac_filtering >/dev/null 2>&1
incus profile device unset claude-dev eth0 security.ipv4_filtering >/dev/null 2>&1
incus network acl rule remove claude-isolate egress action=drop destination=@internal >/dev/null 2>&1
}
trap 'revert_phase_d' EXIT INT TERM
# =========================================================================== # ===========================================================================
phase "D. Hardening rehearsal — #16's changes, applied live (#15 section B)" phase "D. The isolation contract, stated"
# =========================================================================== # ===========================================================================
# The host is disposable, so rehearse the exact changes #16 proposes and watch # Phase D used to REHEARSE the hardening on a throwaway host, because nobody
# what breaks. A FAIL here vetoes a piece of #16's design before it is written. # knew whether it would work. That question is settled: the hardening now ships
# in setup-host.sh and claudebox-firewall.sh, so phase C tests the real thing
# and there is nothing left to rehearse. What the rehearsal established, kept
# here so it is not re-litigated:
# #
# But ONLY if the box was healthy to begin with. On run 7 the baseline was # · @internal is REJECTED as an ACL destination on a bridge network
# broken (a clone/source IP collision took the box's networking down), and phase # ("Unsupported nftables subject") — so the sibling drop is an nftables
# D dutifully reported "L2 filtering BREAKS the box — design veto". It did not. # bridge-family rule, not an ACL rule. It has to be: an L3 ACL never sees
# A verdict measured on a broken box is not a verdict, it is a slander; #16 # frames switched between two ports of one bridge, which is why box→box was
# would have been redesigned around a fiction. Refuse to judge instead. # wide open while the ACL looked airtight.
# · dns.mode=none closes the enumeration leak and public egress survives it.
# · security.ipv4_filtering BREAKS the box's networking (dockerd comes up but
# cannot pull or run a container). VETOED — it is not in the shipped stack.
#
inf "@internal: unsupported on bridge ACLs ⇒ the sibling drop is an nft bridge rule"
inf "dns.mode=none: shipped (closes DNS enumeration, egress unaffected)"
inf "security.ipv4_filtering: VETOED — it breaks the box. Not shipped."
inf "the contract is now tested in phase C against the real stack, not rehearsed"
# The lesson that cost the most: a verdict measured on a broken box is not a
# verdict. Run 7 reported "L2 filtering BREAKS the box" from a box whose network
# was already dead, and #16 was nearly redesigned around it. If the baseline
# failed, say plainly that phase C's isolation results cannot be trusted.
if [ "$BASELINE_OK" -ne 1 ]; then if [ "$BASELINE_OK" -ne 1 ]; then
note "SKIPPING phase D — the box could not reach the internet BEFORE any hardening was applied" no "the box could not reach the internet AT ALL — every isolation result above is suspect, not a pass"
inf "a design verdict measured on a broken baseline is worthless; fix the baseline and re-run" inf "a boundary that 'holds' on a box with no network holds nothing. fix the baseline, re-run."
aud "B1/B3/B5: NOT MEASURED — baseline egress was already broken (see A1)" inf "start with: bash drill/doctor.sh"
fi
if [ "$BASELINE_OK" -eq 1 ]; then
# D1 — dns.mode=none (#15 B3): must kill sibling resolution, must NOT kill egress.
# Runs 2 and 3 DISAGREED on the egress half (broken, then intact) — a verdict
# that flips is a verdict you cannot design on. Setting dns.mode restarts the
# network's dnsmasq, so a probe fired immediately can catch it mid-restart.
# Distinguish TRANSIENT (recovers) from BROKEN (still dead after 30s), and say so.
if incus network set claudenet dns.mode=none 2>/tmp/dnsmode.err; then
sleep 2
[ -n "$(in_box archive getent hosts peer.incus)" ] \
&& { no "dns.mode=none did not stop sibling resolution"; aud "B3 dns.mode=none: does NOT close the leak"; } \
|| { ok "dns.mode=none: sibling names no longer resolve"; aud "B3 dns.mode=none: closes the enumeration leak"; }
if [ "$(box_curl archive https://api.github.com 20)" = 0 ]; then
ok "dns.mode=none: public egress still works, immediately"
aud "B3 egress under dns.mode=none: intact, no outage window"
else
settled=0
for _i in $(seq 1 6); do
sleep 5
[ "$(box_curl archive https://api.github.com 20)" = 0 ] && { settled=1; break; }
done
if [ "$settled" = 1 ]; then
note "dns.mode=none caused a TRANSIENT DNS outage (dnsmasq restart), recovered within 30s"
aud "B3 egress under dns.mode=none: recovers after a brief outage — shippable, but #16 must not probe DNS mid-restart (this is what made runs 2/3 disagree)"
else
no "dns.mode=none BROKE public DNS and it did not recover in 30s — #16 cannot ship it as-is"
aud "B3 egress under dns.mode=none: BROKEN, no recovery — design veto"
fi
fi
else
no "incus rejected dns.mode=none on claudenet: $(head -1 /tmp/dnsmode.err 2>/dev/null)"
aud "B3 dns.mode=none: REJECTED by incus — $(head -1 /tmp/dnsmode.err 2>/dev/null) — #16 needs another mechanism"
fi
# D2 — L2 filtering (#15 B5): the box must keep working with it on.
# Filtering binds at NIC attach, so the boxes restart here.
if incus profile device set claude-dev eth0 security.mac_filtering=true security.ipv4_filtering=true 2>/dev/null; then
incus restart -f archive peer >/dev/null 2>&1
wait_box archive || note "archive slow to return after the filtering restart"
[ "$(box_curl archive https://api.github.com 20)" = 0 ] \
&& { ok "mac+ipv4 filtering on: the box still reaches the internet"; aud "B5 L2 filtering: box networking intact — safe for the claude workload"; } \
|| { no "mac+ipv4 filtering BROKE the box's networking"; aud "B5 L2 filtering: BREAKS the box — design veto"; }
# 'docker info' as the claude user needs the docker group, which a fresh exec
# session may not have picked up; sudo is the honest probe of the DAEMON.
if in_box archive docker run --rm alpine:latest true >/dev/null 2>&1; then
ok "…and in-box Docker still pulls and runs a container under ipv4_filtering"
aud "B5 in-box docker under filtering: WORKS — pulls + runs (NAT hides behind eth0, as #12 argued)"
elif in_box archive docker info >/dev/null 2>&1; then
note "dockerd is up under filtering but could not pull/run a container — check egress from docker0"
aud "B5 in-box docker under filtering: daemon up, container run FAILED — #16 must check docker0 egress"
else
note "in-box Docker daemon is not running at all (so filtering is not what broke it)"
aud "B5 in-box docker under filtering: UNVERIFIED — dockerd not running ($(timeout 30 claudebox incus archive -- exec {} -- systemctl is-active docker 2>&1 | grep -v '^claudebox:' | tail -1))"
fi
else
no "incus rejected security filtering on the profile NIC"
aud "B5 L2 filtering: REJECTED on a profile NIC"
fi
# D3 — @internal as an ACL destination on a bridge network (#15 B1). If it
# holds AND egress survives (the gateway carve-out must win the ordering),
# #16's sibling drop is renumber-proof by construction.
if incus network acl rule add claude-isolate egress action=drop destination=@internal 2>/tmp/ai.err; then
ok "@internal accepted as an ACL destination on a bridge network"
[ "$(box_curl archive https://api.github.com 20)" = 0 ] \
&& { ok "…and public egress survives the @internal drop (the carve-out wins)"; aud "B1 @internal: accepted, egress survives ⇒ #16 uses @internal"; } \
|| { no "the @internal drop killed egress/DNS — ordering does not protect the carve-out"; aud "B1 @internal: accepted but kills DNS ⇒ #16 derives the subnet instead"; }
incus network acl rule remove claude-isolate egress action=drop destination=@internal >/dev/null 2>&1
else
note "@internal rejected on a bridge ACL ($(head -1 /tmp/ai.err 2>/dev/null | cut -c1-60))"
aud "B1 @internal: rejected ⇒ #16 derives the subnet in setup-host.sh (mask the gateway CIDR)"
fi
fi # end of the BASELINE_OK guard around phase D
# Revert now, and CHECK it — the old code fired these into /dev/null and moved
# on, so a failed unset was indistinguishable from a successful one.
revert_phase_d
d="$(incus network get claudenet dns.mode 2>/dev/null)"
f="$(incus profile device get claude-dev eth0 security.ipv4_filtering 2>/dev/null)"
if [ -z "$d" ] && [ -z "$f" ]; then
ok "phase D reverted: dns.mode and NIC filtering are back to shipped defaults"
else
no "PHASE D DID NOT REVERT (dns.mode='$d' ipv4_filtering='$f') — this host will poison the next run"
inf "fix it with: bash drill/doctor.sh --fix"
fi fi
# =========================================================================== # ===========================================================================

View file

@ -27,6 +27,32 @@ else
fi fi
fi fi
# --- Sibling isolation: a box must not reach another box --------------------
#
# This is the ONE rule that makes "isolated even from each other" true, and it
# is not the one anyone expected. The Incus ACL drops egress to 10.0.0.0/8, and
# claudenet's 10.87.0.0/24 sits inside it — so on paper box→box was already
# blocked twice over (the ingress default is drop as well). It was not: a live
# probe found box A's SYN arriving at box B and B answering with a RST.
#
# Why: two boxes on one bridge are on the same L2 segment. Their frames are
# SWITCHED between bridge ports, never routed — so they never traverse the
# netfilter path where an L3 ACL lives. The ACL is not wrong, it simply never
# sees this traffic.
#
# The bridge family DOES see it. Its forward hook fires exactly when a frame is
# passed from one bridge port to another — which, on claudenet, means box→box
# and nothing else: frames addressed to the gateway are delivered locally (the
# INPUT hook), and so is anything being routed out to the internet. So dropping
# every forwarded frame on this bridge isolates the boxes from one another and
# costs them nothing else. DHCP and ARP still work: they are broadcast, and the
# local delivery to dnsmasq happens on INPUT, not FORWARD.
if ! nft list table bridge claudebox >/dev/null 2>&1; then
nft add table bridge claudebox
nft "add chain bridge claudebox forward { type filter hook forward priority -200 ; policy accept ; }"
nft add rule bridge claudebox forward meta ibrname "$NET" meta obrname "$NET" drop
fi
# Docker rewrites FORWARD policy to DROP; DOCKER-USER is its escape hatch. # Docker rewrites FORWARD policy to DROP; DOCKER-USER is its escape hatch.
if command -v docker >/dev/null && iptables -L DOCKER-USER -n >/dev/null 2>&1; then if command -v docker >/dev/null && iptables -L DOCKER-USER -n >/dev/null 2>&1; then
iptables -C DOCKER-USER -i "$NET" -j ACCEPT 2>/dev/null || iptables -I DOCKER-USER -i "$NET" -j ACCEPT iptables -C DOCKER-USER -i "$NET" -j ACCEPT 2>/dev/null || iptables -I DOCKER-USER -i "$NET" -j ACCEPT

View file

@ -38,6 +38,23 @@ incus network set claudenet security.acls=claude-isolate \
security.acls.default.egress.action=allow \ security.acls.default.egress.action=allow \
security.acls.default.ingress.action=drop security.acls.default.ingress.action=drop
# A box must not be able to ENUMERATE its siblings, either. dnsmasq on the
# gateway serves DNS (that carve-out is what makes egress resolution work) and
# it holds a record for every instance on the network — so 'getent hosts <box>'
# from inside one box resolved another's name and address. Connection blocked,
# reconnaissance wide open. dns.mode=none stops it registering instance records;
# forwarding for public names is unaffected (verified live).
incus network set claudenet dns.mode=none
# Sibling isolation itself is NOT an ACL rule — an L3 ACL never sees frames
# switched between two ports of one bridge. It lives in claudebox-firewall.sh
# as an nftables bridge-family rule. See the comment there; it is the reason
# boxes cannot reach each other.
# IPv6 stays off (ipv6.address=none, above). Every rule in the ACL and every
# rule in the firewall is IPv4-only, so IPv6 would be an uncovered path, not a
# feature. That is a contract, not a default.
# --- Firewall coexistence --------------------------------------------------- # --- Firewall coexistence ---------------------------------------------------
# Hosts running UFW (INPUT drop) and/or Docker (FORWARD drop) silently eat # Hosts running UFW (INPUT drop) and/or Docker (FORWARD drop) silently eat
# claudenet traffic. Punch minimal, ordered holes; the Incus ACL still layers # claudenet traffic. Punch minimal, ordered holes; the Incus ACL still layers