box/docs/plans/2026-07-18-restricted-tier.md
dan-claude-bot 2097982ac2 docs: the restricted tier — README, design doc, plan doc with measured results (#74)
The plan doc records what was measured and why each decision fell where it
did: the private bridge is worse than unhardened (a live NAT bridge with
IPv6 on), incus-user blocks snapshots, no daemon-level template exists (read
in incus-user's source), widening survives re-sync (same source, then
measured live). box-design gains the access-tiers section — including why
narrowing to boxnet-only is the load-bearing decision and why the nft bridge
drop is the layer a restricted user cannot strip. RUNS.md logs MU-1..3.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 04:09:48 +00:00

6.6 KiB
Raw Blame History

Restricted incus tier — design and measured results (#74)

Status: implemented and rehearsed. 41/41 rehearsal criteria green on the design host (Debian 13 trixie, Incus 6.0.4, nested KVM), in container mode and VM mode. This doc records the design, what was measured, and why each decision fell where it did. It supersedes the vetoed #72 design (docs/plans/2026-07-17-multiuser-hosts.md on feat/restricted-tier-wip).

What #74 asked

A restricted (incus-group) user can box new/list/shell/snapshot/rm their own boxes, on a network carrying box's full isolation contract, seeing no one else's; the convergence is a documented, idempotent path; the rehearsal passes criteria (a)(f); the admin tier is unchanged.

The three facts that shaped the design

Task-0 (#72) measured one: incus-user confines users to user-<uid> projects (sound), but pins them to a private auto-created bridge incusbr-<uid> and restricted.networks.access: incusbr-<uid> — they cannot even see boxnet.

This round measured two more:

  1. The private bridge is worse than unhardened. incusbr-<uid> is a fully functional NAT bridge — ipv4.nat=true, ipv6.nat=true — with no ACL, no dns.mode=none, no resolver pin, no port isolation, and IPv6 egress box's contract explicitly forbids. Any instance placed on it holds a door to the host's LAN.
  2. incus-user projects block snapshots (Project "user-<uid>" doesn't allow for snapshot creation) — box's entire reuse workflow.

And two open questions from #74, answered from incus-user's own source (cmd/incus-user/server.go, stable-6.0) and then confirmed live:

  • Daemon-level project template? None exists — the project config is hardcoded in serverSetupUser(). A per-user admin hook is the only path.
  • Does widening survive a re-sync? Yes. Setup runs only when the project does not exist (and early-outs when the user's certificate is already trusted); incus-user never rewrites an existing project. Confirmed live: systemctl restart incus-user.socket leaves the convergence intact.

The decision: option 1, tightened

#74 offered (1) converge users onto the shared boxnet or (2) harden each private bridge. Option 2 multiplies every mechanism per user (ACL, resolver pin, dnsmasq, nft rules, firewall coexistence) and turns the shipped static profile into N generated ones. Option 1 keeps one hardened network and one shipped profile — and the existing box↔box mechanisms already make cross-user isolation free: the nft bridge-family drop and dns.mode=none are host/network-owned, so they bind every instance on boxnet no matter whose project it lives in.

One tightening beyond the issue's sketch: the issue proposed restricted.networks.access boxnet,incusbr-<uid> ("must list both" — true as long as the default profile still references the private bridge). Listing both leaves fact 1's unhardened bridge one --network flag away, forever. Instead, box grant:

  • removes eth0 from the project's default profile (nothing references the private bridge anymore, so the narrowing validates), and
  • sets restricted.networks.access boxnetonly.

After which the hardened network is not the user's default placement but the only placement their certificate can express. Measured: incus launch --network incusbr-<uid> as the user → Network not found; the user cannot widen their own project (Error: Certificate is restricted); they cannot touch boxnet's config or the ACL (no permission for project "default").

A restricted user CAN edit the box-net profile copy in their own project (they own project profiles — features.profiles=true), including stripping security.port_isolation. That is why the host-owned nft bridge drop is the second layer: meta ibrname boxnet obrname boxnet drop fires on every port-to-port frame regardless of per-NIC flags. Cross-user sibling probes are dropped either way — measured from inside the boxes.

What box grant <user> converges (idempotent, re-run to refresh)

  1. usermod -aG incus (not incus-admin — that is the tier)
  2. first-touch incus-user as the user (runuser/sudo -u, stdin pinned) — the project is created lazily and cannot be pre-created by an admin
  3. remove the default profile's private-bridge eth0
  4. restricted.networks.access boxnet
  5. restricted.snapshots allow
  6. install/refresh the shipped box-net profile into the project
  7. verify from the user's side of the socket

box revoke <user> is the inverse, two strengths: bare = group removal (the socket closes; boxes keep running; re-grant restores), --purge = boxes, images, project, private bridge, trust-store certificate, incus-user state — then asserts the absence (the wipe.sh discipline).

The tier in the CLI

box_tier() — UID 0 / incus-admin → admin, incus alone → restricted, neither → none — decided from live process credentials (argless id -nG), byte-identical in bin/box and host/setup-host.sh (diffed by a test). Tier-aware surface: new pre-flights the profile and names the right fix per tier; expose refuses before any daemon call (its plumbing is daemon-global; without the guard the failure is a lie — a restricted user cannot read boxnet's redacted config, so box_net_ip would claim their running box has no address); setup-host exits 0 with the honest note; doctor runs a restricted check-set (is the tier granted, does their box resolve/route) instead of judging host state they cannot see.

Found along the way, fixed for every tier: box restore dispatched incus restore, which does not exist in Incus 6 (incus snapshot restore). It had never worked.

Rehearsal and CI

drill/multiuser.sh (root, opt-in via BOX_MULTIUSER_REHEARSAL=1) proves criteria (a)(f) from #74 plus the measured extensions (g)(l): the in-box isolation contract (egress, DNS, box→host, RFC1918, cross-user sibling drop, name enumeration, IPv6-off), the closed escape hatches, re-sync survival, and scoped revoke. Two real users, real grants, real mints, probes from inside; --container for CI, VM mode on real hardware; cleanup deletes everything it made.

CI gains a rehearsal job (ubuntu-latest): install incus, stage the checkout at /opt/box (the #71 layout), setup-host, doctor, then the rehearsal in container mode. The tier's semantics are proven on every PR against a live daemon; the VM trust boundary stays a real-hardware ritual, like the drill.

Environment

Debian 13 (trixie), Incus 6.0.4, incus-user.socket shipped in the incus package (Debian 13 and Ubuntu 24.04 both), /dev/kvm present (VM-mode run), btrfs pool. Companion work: rig#24 (box role), rig#12/#25 (host-class).