2026-07-18 20:57:13 +00:00
|
|
|
# Changelog
|
|
|
|
|
|
|
|
|
|
History before 0.1.0 lives in git — rig grew its version surface (`VERSION`,
|
|
|
|
|
`rig --version`, the side-by-side `versions/<v>` install layout; #35/#36)
|
|
|
|
|
on the way to cutting its first release, and this file starts there.
|
|
|
|
|
|
fix: dropping the box role revokes through box, not behind its back
`users apply` converged group `incus` with a bare `gpasswd -d`, the same
move it makes for `rig-admin` and `rig`. Those two are rig's. `incus` is
box's, and `box revoke` does strictly more with it: it says out loud that
supplementary groups are read AT LOGIN, so a session the dropped operator
already holds keeps the Incus socket until that session dies, and it hands
over `loginctl terminate-user <user>` as the remedy.
rig logged "removed <user> from incus" and moved on. An operator who
dropped someone from the users file and watched apply succeed believed the
VM access was gone — and was wrong for as long as that user held a session.
Both removal paths — the per-user convergence loop and the dropped-user
sweep — now route the incus group through one `drop_incus` helper that
calls `box revoke`, keeping a single owner for the group. Never `--purge`:
that deletes the user's boxes, images and project, and destroying someone's
running machines is not a convergence step; it stays an explicit admin act.
The exit code is not trusted (the #12 lesson bootstrap already applies to
box's installer): a revoke that returns 0 with the membership still
standing has not closed the socket, so the effective state is checked and
rig falls back to removing the group itself — as it also does on a host
where box is not installed. Every fallback path carries the session warning
in rig's own voice, because the silence was the bug. The absent-group case
needs no new guard: `id -nG` cannot report a group that does not exist, so
the existing `in_group` test at both call sites is already false on a
host=no box or one where `box setup-host` never ran.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:18:21 +00:00
|
|
|
## Unreleased
|
2026-07-18 20:57:13 +00:00
|
|
|
|
fix(labels): state:needs-human means a human could merge it right now
Ported from heavy-duty/box#137 (heavy-duty/box#136) so the three repos'
reconcilers stay byte-identical. The state machine here was byte-identical to
box's before this change, and remains so after -- only the scope:* taxonomy
differs, correctly.
decide_state() derived state from three inputs -- draft flag, requested
reviewers, submitted reviews -- and read NOTHING about mergeability or checks.
With the `if requested "$HUMAN"` short-circuit at the top of its precedence,
the label was sticky: once the maintainer was requested, a PR read
state:needs-human through conflicts, through red CI, through a force-push that
staled every approval.
This repo paid for it directly. During the ten-PR batch merged today, every
merge re-conflicted the PRs below it through CHANGELOG.md, and each kept its
state:needs-human label throughout -- inviting merges that could not happen.
It was caught only by opening them one at a time, which is the work the label
exists to save.
The rule the label now keeps: state:needs-human means a human could merge this
RIGHT NOW, so anything making that false outranks the request that put it
there.
CONFLICTING or failing checks -> state:needs-rebase (new; the agent's to fix)
approvals staled by a push -> state:addressing (nobody reviewed this tree)
An UNFINISHED round still yields to an explicit human request -- MISSING
(nobody has reviewed yet) is a different fact from STALE (everyone reviewed
something else). UNKNOWN mergeability is NOT treated as unmergeable: GitHub
reports it for about a minute after every merge, and flapping every open PR
through needs-rebase on each merge would be worse than the bug. A failed read
degrades to the same "do not know" value.
Also adds merge-next: queue order is intent, so the reconciler never sets it,
only CLEARS it once the PR stops being mergeable-by-a-human.
Fixtures 19 -> 29, including that UNKNOWN does not trigger needs-rebase and a
draft outranks a conflict. No live dry-run evidence here -- this repo has no
open PRs right now -- so the fixtures and box's live dry-run are the proof.
Closes #87
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 15:26:54 +00:00
|
|
|
### Fixed
|
|
|
|
|
|
|
|
|
|
- **`state:needs-human` no longer appears on PRs a human cannot merge**
|
|
|
|
|
(#87, heavy-duty/box#136) — `decide_state()` derived state from three inputs
|
|
|
|
|
(draft flag, requested reviewers, submitted reviews) and read *nothing* about
|
|
|
|
|
mergeability or checks. Combined with the `if requested "$HUMAN"`
|
|
|
|
|
short-circuit at the top of its precedence, the label was **sticky**: once
|
|
|
|
|
the maintainer was requested, the PR read `state:needs-human` through
|
|
|
|
|
conflicts, through red CI, through a force-push that staled every approval.
|
|
|
|
|
Nothing demoted it.
|
|
|
|
|
|
|
|
|
|
This repo paid for it directly. During the ten-PR batch merged on 2026-07-20,
|
|
|
|
|
every merge re-conflicted the PRs below it through `CHANGELOG.md` — and each
|
|
|
|
|
one kept its `state:needs-human` label the whole time, inviting a merge that
|
|
|
|
|
could not happen. It was noticed only by opening them one at a time, which is
|
|
|
|
|
the exact work the label exists to save.
|
|
|
|
|
|
|
|
|
|
The rule the label now keeps is that **`state:needs-human` means a human
|
|
|
|
|
could merge this right now**, so anything making that false outranks the
|
|
|
|
|
request that put it there. A `CONFLICTING` branch or a failing check is the
|
|
|
|
|
agent's to fix: new `state:needs-rebase`. Approvals staled by a push mean
|
|
|
|
|
nobody reviewed this tree: `state:addressing`, because the agent owes a
|
|
|
|
|
re-request. An *unfinished* round still yields to an explicit human request —
|
|
|
|
|
a maintainer pulling a PR to themselves early is deliberate, and `MISSING`
|
|
|
|
|
(nobody has reviewed yet) is a different fact from `STALE` (everyone reviewed
|
fix(labels): unrecognised check outcomes block, and STALE outranks MISSING
Round 2 review found two ways the "a human could merge this right now"
invariant still leaked, both of which let state:needs-human land on a PR
the button would refuse.
The check-rollup classifier enumerated the outcomes that block and
defaulted the rest to SUCCESS, so ERROR, CANCELLED and STALE fell through
into green. Inverted to an allow-list of the outcomes that DON'T block
(SUCCESS, NEUTRAL, SKIPPED, plus the pending set); everything else,
including an outcome neither enum has today, blocks. The rollup mixes
CheckRun.conclusion with StatusContext.state and an outcome the list
forgets is one we cannot certify as mergeable — a false FAILURE parks the
PR on the agent, a false SUCCESS invites a bad merge. The classifier also
moved out of main() into checks_state(), which is why no fixture caught
this: it was inline in the fetch loop and the jq itself was untestable.
Once CANCELLED blocks, superseded runs must be dropped first — a re-run
does not evict the run it replaced, and judging every entry would strand
every re-run PR in needs-rebase. Each context now collapses to its newest
entry, keyed on workflow + job name because a bare job name is only
unique within its workflow.
decide_state() returned from inside the bot loop on the first MISSING, so
a STALE belonging to a later bot in BOTS was never read: a round that was
both unfinished and staled came out needs-human with nothing bound to the
head. The whole round is now collected before any precedence is applied,
STALE ahead of MISSING. The MISSING-yields-to-an-explicit-human-request
rule is untouched.
Fixtures 29 -> 44, pinning the whole check-outcome enum, the supersede
rule in both orders, and the mixed round at both ends of BOTS.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:05:39 +00:00
|
|
|
something else). Precedence is applied to the round as a whole, after every
|
|
|
|
|
verdict is collected: deciding inside the loop let the order of `BOTS` pick
|
|
|
|
|
the answer, so a round that was *both* unfinished and staled returned on the
|
|
|
|
|
`MISSING` before any later bot's `STALE` was read — and came out
|
|
|
|
|
`needs-human` over a head nobody had reviewed, the original bug wearing a
|
|
|
|
|
different hat.
|
|
|
|
|
|
|
|
|
|
Whether a check blocks is judged by listing the outcomes that *don't* —
|
|
|
|
|
`SUCCESS`, `NEUTRAL`, `SKIPPED`, and the pending set — rather than the
|
|
|
|
|
outcomes that do. The rollup mixes two closed enums (`CheckRun.conclusion`
|
|
|
|
|
and `StatusContext.state`), and an outcome the list forgets is one the label
|
|
|
|
|
cannot certify as mergeable: `ERROR`, `CANCELLED` and `STALE` all read as
|
|
|
|
|
green under an allow-list of failures. The costs are not symmetric — a false
|
|
|
|
|
failure parks the PR on the agent, who looks; a false success invites a human
|
|
|
|
|
to merge a tree that will not merge. Superseded runs are dropped first, each
|
|
|
|
|
context collapsing to its newest entry: a re-run does not evict the run it
|
|
|
|
|
replaced, so box#137's own tip carried a `CANCELLED` `scope` beside the
|
|
|
|
|
`SUCCESS` `scope` that superseded it, and judging every entry would have
|
|
|
|
|
stranded every re-run PR in `needs-rebase`.
|
fix(labels): state:needs-human means a human could merge it right now
Ported from heavy-duty/box#137 (heavy-duty/box#136) so the three repos'
reconcilers stay byte-identical. The state machine here was byte-identical to
box's before this change, and remains so after -- only the scope:* taxonomy
differs, correctly.
decide_state() derived state from three inputs -- draft flag, requested
reviewers, submitted reviews -- and read NOTHING about mergeability or checks.
With the `if requested "$HUMAN"` short-circuit at the top of its precedence,
the label was sticky: once the maintainer was requested, a PR read
state:needs-human through conflicts, through red CI, through a force-push that
staled every approval.
This repo paid for it directly. During the ten-PR batch merged today, every
merge re-conflicted the PRs below it through CHANGELOG.md, and each kept its
state:needs-human label throughout -- inviting merges that could not happen.
It was caught only by opening them one at a time, which is the work the label
exists to save.
The rule the label now keeps: state:needs-human means a human could merge this
RIGHT NOW, so anything making that false outranks the request that put it
there.
CONFLICTING or failing checks -> state:needs-rebase (new; the agent's to fix)
approvals staled by a push -> state:addressing (nobody reviewed this tree)
An UNFINISHED round still yields to an explicit human request -- MISSING
(nobody has reviewed yet) is a different fact from STALE (everyone reviewed
something else). UNKNOWN mergeability is NOT treated as unmergeable: GitHub
reports it for about a minute after every merge, and flapping every open PR
through needs-rebase on each merge would be worse than the bug. A failed read
degrades to the same "do not know" value.
Also adds merge-next: queue order is intent, so the reconciler never sets it,
only CLEARS it once the PR stops being mergeable-by-a-human.
Fixtures 19 -> 29, including that UNKNOWN does not trigger needs-rebase and a
draft outranks a conflict. No live dry-run evidence here -- this repo has no
open PRs right now -- so the fixtures and box's live dry-run are the proof.
Closes #87
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 15:26:54 +00:00
|
|
|
|
fix(labels): order check runs by one consistent quantity — when they began
The dating expression took the newest stamp each run carries. That reads
`completedAt` for a finished run and `startedAt` for a live one, so the
comparison comes down to "when this one ended" against "when that one
began" — which is not an ordering on runs at all.
A run cancelled by the concurrency group does not stop the instant its
replacement starts: the runner has to receive the signal and wind down. So
`predecessor.completedAt > successor.startedAt` is the ordinary case, not a
corner. On the box#137 tip that motivated the supersede rule the window was
13s wide — the superseding run started 15:19:38, the run it cancelled did
not finish until 15:19:51 — and for that whole window the dying predecessor
out-dated its own live replacement, so `last` discarded the replacement and
judged the corpse.
Both round-3 failure modes came back inside that window, narrowed rather
than closed: a draining CANCELLED predecessor reported FAILURE and sent the
agent to fix nothing, and a draining SUCCESS predecessor reported SUCCESS —
mergeable, all bots approve, state:needs-human — over a tree whose merge
button branch protection had already disabled. #136 again, one field over.
Dated by `first` of the preference-ordered stamps rather than `max` of
them: start time if the run recorded one, falling back only if it did not.
The sentinel filtering is unchanged, and finished runs still date by
completion when that is all they carry, so the supersede rule keeps the
case it exists for.
Found independently by claude-bot-andresmgsl and codex-bot-andresmgsl.
No existing fixture could express it — `run_()` carries no startedAt, so
every supersede fixture spaced the predecessor's completion safely before
the successor's start, the same blind spot as round 3 one field over. New
`drained_()` helper pins both directions; fixtures 48 -> 51 (with the
reverse-direction in-flight fixture ported from cast#128).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:32:53 +00:00
|
|
|
Which entry is newest is dated by **when the run began** — `startedAt`,
|
|
|
|
|
falling back only if it never recorded one — discarding *both* spellings of
|
|
|
|
|
absent. A run still in flight has no completion, but `gh` does not omit the
|
|
|
|
|
field: its Go struct marshals the zero time as the string
|
|
|
|
|
`0001-01-01T00:00:00Z`, and `//` falls through `null` and `false` only.
|
|
|
|
|
Dating by completion therefore sorted the *live* re-run to the bottom and
|
|
|
|
|
let `last` pick the very run it superseded — a green context with a
|
|
|
|
|
replacement mid-flight read `SUCCESS`, inviting a merge the button had
|
|
|
|
|
already disabled, which is #136 restored by the fix for it.
|
|
|
|
|
|
|
|
|
|
Taking the *newest* stamp each run carries is not a fix either, and this is
|
|
|
|
|
the subtle part: it compares a finished predecessor by when it **ended**
|
|
|
|
|
against a live successor by when it **began**, which is not an ordering on
|
|
|
|
|
runs at all. A run cancelled by the concurrency group does not stop the
|
|
|
|
|
instant its replacement starts — the runner has to receive the signal and
|
|
|
|
|
wind down — so the predecessor completing *after* the successor started is
|
|
|
|
|
the ordinary case, 13s wide on the box#137 tip that motivated the supersede
|
|
|
|
|
rule. For that whole drain window the dying predecessor out-dated its own
|
|
|
|
|
replacement. One consistent quantity, start time, is the only ordering that
|
|
|
|
|
holds. An entry carrying no usable timestamp at all sorts last rather than
|
|
|
|
|
first, so something that cannot be dated is never discarded in favour of a
|
|
|
|
|
stale success. Every ambiguity here resolves toward "not settled".
|
fix(labels): date a check run by when it started, not by a zero completion
The supersede collapse added in the previous commit dated each run by
`.completedAt // .startedAt // .createdAt`. A run still in flight has no
completion, but `gh` does not omit the field: its Go struct marshals the
zero time as the string "0001-01-01T00:00:00Z", and jq's `//` falls
through null and false only. The sentinel was therefore taken as the sort
key, and it sorts before every real timestamp — so the LIVE re-run became
the oldest entry in its context, `last` discarded it, and the run it
superseded was judged instead.
That restored #136 through the fix for it: a green context with a
replacement mid-flight reported SUCCESS, so a PR read mergeable, green,
all bots approve — state:needs-human — while branch protection had the
merge button disabled. It also narrowed rather than removed the flap the
supersede rule exists to prevent: between "run A cancelled by the
concurrency group" and "run B finishes", the PR reported FAILURE and the
agent was sent to fix something that was not broken.
Runs are now dated by the newest timestamp they actually carry, with both
spellings of absent discarded (null, and the zero sentinel). Entries that
carry no usable timestamp sort LAST rather than first: something we cannot
date is most likely the thing just created, and treating it as newest
keeps an undateable in-flight run from being discarded in favour of a
stale success. Every ambiguity resolves toward "not settled".
Found independently by claude-bot-andresmgsl and codex-bot-andresmgsl.
The fixtures could not have caught it: `run_()` always emits a real
completedAt, so every supersede fixture was a race between two finished
runs, and the bug lived in the one shape the helper could not express.
New `inflight_()` helper covers it; fixtures 44 -> 48.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:18:50 +00:00
|
|
|
|
fix(labels): state:needs-human means a human could merge it right now
Ported from heavy-duty/box#137 (heavy-duty/box#136) so the three repos'
reconcilers stay byte-identical. The state machine here was byte-identical to
box's before this change, and remains so after -- only the scope:* taxonomy
differs, correctly.
decide_state() derived state from three inputs -- draft flag, requested
reviewers, submitted reviews -- and read NOTHING about mergeability or checks.
With the `if requested "$HUMAN"` short-circuit at the top of its precedence,
the label was sticky: once the maintainer was requested, a PR read
state:needs-human through conflicts, through red CI, through a force-push that
staled every approval.
This repo paid for it directly. During the ten-PR batch merged today, every
merge re-conflicted the PRs below it through CHANGELOG.md, and each kept its
state:needs-human label throughout -- inviting merges that could not happen.
It was caught only by opening them one at a time, which is the work the label
exists to save.
The rule the label now keeps: state:needs-human means a human could merge this
RIGHT NOW, so anything making that false outranks the request that put it
there.
CONFLICTING or failing checks -> state:needs-rebase (new; the agent's to fix)
approvals staled by a push -> state:addressing (nobody reviewed this tree)
An UNFINISHED round still yields to an explicit human request -- MISSING
(nobody has reviewed yet) is a different fact from STALE (everyone reviewed
something else). UNKNOWN mergeability is NOT treated as unmergeable: GitHub
reports it for about a minute after every merge, and flapping every open PR
through needs-rebase on each merge would be worse than the bug. A failed read
degrades to the same "do not know" value.
Also adds merge-next: queue order is intent, so the reconciler never sets it,
only CLEARS it once the PR stops being mergeable-by-a-human.
Fixtures 19 -> 29, including that UNKNOWN does not trigger needs-rebase and a
draft outranks a conflict. No live dry-run evidence here -- this repo has no
open PRs right now -- so the fixtures and box's live dry-run are the proof.
Closes #87
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 15:26:54 +00:00
|
|
|
`UNKNOWN` mergeability is deliberately not treated as unmergeable: GitHub
|
|
|
|
|
reports it for about a minute after every merge while it recomputes, and
|
|
|
|
|
flapping every open PR through `needs-rebase` on each merge would be worse
|
|
|
|
|
than the bug. A failed read of either fact degrades to the same "do not know"
|
|
|
|
|
value, for the same reason.
|
|
|
|
|
|
|
|
|
|
Also adds `merge-next`, because a correct `needs-human` still does not say
|
|
|
|
|
*which* PR to merge first, and order matters when they conflict. Queue order
|
|
|
|
|
is intent, so the reconciler never sets it — it only **clears** it once the
|
|
|
|
|
PR stops being mergeable-by-a-human, which is precisely the staleness that
|
fix(labels): date a check run by when it started, not by a zero completion
The supersede collapse added in the previous commit dated each run by
`.completedAt // .startedAt // .createdAt`. A run still in flight has no
completion, but `gh` does not omit the field: its Go struct marshals the
zero time as the string "0001-01-01T00:00:00Z", and jq's `//` falls
through null and false only. The sentinel was therefore taken as the sort
key, and it sorts before every real timestamp — so the LIVE re-run became
the oldest entry in its context, `last` discarded it, and the run it
superseded was judged instead.
That restored #136 through the fix for it: a green context with a
replacement mid-flight reported SUCCESS, so a PR read mergeable, green,
all bots approve — state:needs-human — while branch protection had the
merge button disabled. It also narrowed rather than removed the flap the
supersede rule exists to prevent: between "run A cancelled by the
concurrency group" and "run B finishes", the PR reported FAILURE and the
agent was sent to fix something that was not broken.
Runs are now dated by the newest timestamp they actually carry, with both
spellings of absent discarded (null, and the zero sentinel). Entries that
carry no usable timestamp sort LAST rather than first: something we cannot
date is most likely the thing just created, and treating it as newest
keeps an undateable in-flight run from being discarded in favour of a
stale success. Every ambiguity resolves toward "not settled".
Found independently by claude-bot-andresmgsl and codex-bot-andresmgsl.
The fixtures could not have caught it: `run_()` always emits a real
completedAt, so every supersede fixture was a race between two finished
runs, and the bug lived in the one shape the helper could not express.
New `inflight_()` helper covers it; fixtures 44 -> 48.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:18:50 +00:00
|
|
|
made `needs-human` untrustworthy. Both live shapes, the mixed round, the
|
fix(labels): order check runs by one consistent quantity — when they began
The dating expression took the newest stamp each run carries. That reads
`completedAt` for a finished run and `startedAt` for a live one, so the
comparison comes down to "when this one ended" against "when that one
began" — which is not an ordering on runs at all.
A run cancelled by the concurrency group does not stop the instant its
replacement starts: the runner has to receive the signal and wind down. So
`predecessor.completedAt > successor.startedAt` is the ordinary case, not a
corner. On the box#137 tip that motivated the supersede rule the window was
13s wide — the superseding run started 15:19:38, the run it cancelled did
not finish until 15:19:51 — and for that whole window the dying predecessor
out-dated its own live replacement, so `last` discarded the replacement and
judged the corpse.
Both round-3 failure modes came back inside that window, narrowed rather
than closed: a draining CANCELLED predecessor reported FAILURE and sent the
agent to fix nothing, and a draining SUCCESS predecessor reported SUCCESS —
mergeable, all bots approve, state:needs-human — over a tree whose merge
button branch protection had already disabled. #136 again, one field over.
Dated by `first` of the preference-ordered stamps rather than `max` of
them: start time if the run recorded one, falling back only if it did not.
The sentinel filtering is unchanged, and finished runs still date by
completion when that is all they carry, so the supersede rule keeps the
case it exists for.
Found independently by claude-bot-andresmgsl and codex-bot-andresmgsl.
No existing fixture could express it — `run_()` carries no startedAt, so
every supersede fixture spaced the predecessor's completion safely before
the successor's start, the same blind spot as round 3 one field over. New
`drained_()` helper pins both directions; fixtures 48 -> 51 (with the
reverse-direction in-flight fixture ported from cast#128).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:32:53 +00:00
|
|
|
whole check-outcome enum, and the re-run in every temporal arrangement it
|
|
|
|
|
occurs in — superseded, still in flight in both spellings of an absent
|
|
|
|
|
completion, undateable, and still draining past its replacement's start —
|
|
|
|
|
are pinned in `test/labels-reconcile.sh`. Ported from heavy-duty/box#137 so
|
|
|
|
|
the three repos' reconcilers stay byte-identical; fixtures 19 → 51.
|
fix(labels): state:needs-human means a human could merge it right now
Ported from heavy-duty/box#137 (heavy-duty/box#136) so the three repos'
reconcilers stay byte-identical. The state machine here was byte-identical to
box's before this change, and remains so after -- only the scope:* taxonomy
differs, correctly.
decide_state() derived state from three inputs -- draft flag, requested
reviewers, submitted reviews -- and read NOTHING about mergeability or checks.
With the `if requested "$HUMAN"` short-circuit at the top of its precedence,
the label was sticky: once the maintainer was requested, a PR read
state:needs-human through conflicts, through red CI, through a force-push that
staled every approval.
This repo paid for it directly. During the ten-PR batch merged today, every
merge re-conflicted the PRs below it through CHANGELOG.md, and each kept its
state:needs-human label throughout -- inviting merges that could not happen.
It was caught only by opening them one at a time, which is the work the label
exists to save.
The rule the label now keeps: state:needs-human means a human could merge this
RIGHT NOW, so anything making that false outranks the request that put it
there.
CONFLICTING or failing checks -> state:needs-rebase (new; the agent's to fix)
approvals staled by a push -> state:addressing (nobody reviewed this tree)
An UNFINISHED round still yields to an explicit human request -- MISSING
(nobody has reviewed yet) is a different fact from STALE (everyone reviewed
something else). UNKNOWN mergeability is NOT treated as unmergeable: GitHub
reports it for about a minute after every merge, and flapping every open PR
through needs-rebase on each merge would be worse than the bug. A failed read
degrades to the same "do not know" value.
Also adds merge-next: queue order is intent, so the reconciler never sets it,
only CLEARS it once the PR stops being mergeable-by-a-human.
Fixtures 19 -> 29, including that UNKNOWN does not trigger needs-rebase and a
draft outranks a conflict. No live dry-run evidence here -- this repo has no
open PRs right now -- so the fixtures and box's live dry-run are the proof.
Closes #87
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 15:26:54 +00:00
|
|
|
|
feat: rig platform — what is this machine, computed not stored
rig read no hardware at all. The single exception was `uname -m` in
runner-install.sh, used to pick a runner tarball and then discarded — so
"is this the 32GB one, or the M900?" was a question you answered by
logging in and running free -h, nproc, df -h and uname -r by hand, four
commands deep, on a machine you were already unsure about.
`rig platform` prints hostname, OS, kernel, CPU, memory, disk and
virtualization, then a provenance block: which rig, when, and the role
marker's traits.
It COMPUTES rather than stores, and that is the design rather than an
implementation detail. Specs change without rig doing anything — RAM
added, root disk resized, the unattended-upgrades bootstrap itself
enables patching the kernel — so a stored spec is stale the moment the
machine changes, and refreshing one on every run would collide with
bootstrap's "safe to re-run; a second run changes nothing" contract.
Nothing is written, so nothing can go stale.
The corollary is deliberate: reading only /proc, uname, /etc/os-release,
df and systemd-detect-virt means no root, no network, and it runs on a
pristine Debian box rig has never bootstrapped — useful for deciding
what to converge a machine into, not only for auditing it afterwards.
That also makes it the rare rig command the harness can RUN for real
rather than grep: the tests assert the answer describes the actual test
machine (kernel and hostname compared against independently computed
values), and assert it writes nothing.
Both known traps are handled explicitly. /etc/os-release is sourced in a
SUBSHELL — it defines VERSION, NAME and ID and would otherwise clobber
same-named script variables, the form every other site in this tree uses
and test/cli.sh already greps for. systemd-detect-virt exits non-zero on
bare metal while printing 'none', a normal answer that set -e would
otherwise turn into a failed run, so it is wrapped in `|| true`.
Provenance is read, never written, and degrades per file.
/etc/rig/manifest is #61 and does not exist yet, so that line reads
'not bootstrapped' on every machine today; the command ships complete
without it and neither blocks the other.
Named `platform` and not `status`: `users status` and `runner status`
cross-check recorded against live state and print DRIFT, and a command
that records nothing cannot drift, so calling it status would borrow a
promise it structurally cannot make. It also leaves `rig status` free
for the machine-wide roll-up it will eventually want to be.
Refs #64
2026-07-19 23:35:58 +00:00
|
|
|
### Added
|
|
|
|
|
|
|
|
|
|
- **`rig platform` — what is this machine, calculated at run time, stored
|
|
|
|
|
nowhere** (#64) — rig read no hardware at all; the single exception was
|
|
|
|
|
`uname -m` in `runner-install.sh`, used to pick a tarball and then
|
|
|
|
|
discarded. So "is this the 32GB one, or the M900?" was answered by logging
|
|
|
|
|
in and running `free -h`, `nproc`, `df -h` and `uname -r` by hand, four
|
|
|
|
|
commands deep, on a machine you were already unsure about. `rig platform`
|
|
|
|
|
prints hostname, OS, kernel, CPU, memory, disk and virtualization, then a
|
|
|
|
|
provenance block (which rig, when, and the role marker's traits). It
|
|
|
|
|
**computes rather than stores**: specs change without rig doing anything —
|
|
|
|
|
RAM added, root disk resized, the unattended-upgrades bootstrap itself
|
|
|
|
|
enables patching the kernel — so a stored spec is stale the moment the
|
|
|
|
|
machine changes, and refreshing one per run would collide with bootstrap's
|
|
|
|
|
"a second run changes nothing" contract. Nothing is written, so nothing can
|
|
|
|
|
go stale. The corollary is deliberate: reading only `/proc`, `uname`,
|
|
|
|
|
`/etc/os-release`, `df` and `systemd-detect-virt` means it needs no root,
|
|
|
|
|
makes no network call, and **runs on a pristine Debian box rig has never
|
|
|
|
|
bootstrapped** — useful for deciding what to converge a machine into, not
|
|
|
|
|
only for auditing it afterwards. That also makes it the rare rig command
|
|
|
|
|
the harness can RUN for real instead of grepping: the tests assert the
|
|
|
|
|
actual answer describes the actual test machine. Provenance is read, never
|
|
|
|
|
written, and degrades per-file — `/etc/rig/manifest` is #61 and does not
|
2026-07-20 00:22:01 +00:00
|
|
|
exist yet, so those lines read `not bootstrapped` on every machine today and
|
|
|
|
|
nothing else depends on it. The reader is keyed to #61's documented schema
|
|
|
|
|
(`schema`, `bootstrapped_by`/`_at`, `converged_by`/`_at`) and fixtures pin
|
|
|
|
|
that exact spelling, so the integration cannot land silently broken; birth
|
2026-07-20 00:34:26 +00:00
|
|
|
and latest stay separate rather than one being inferred from the other. A
|
|
|
|
|
fresh machine writes both pairs equal, so two identical lines read as
|
|
|
|
|
"never re-converged"; a manifest missing the pair is partial rather than
|
|
|
|
|
fresh, and says so instead of backfilling from birth. Named `platform` and not `status` on purpose:
|
feat: rig platform — what is this machine, computed not stored
rig read no hardware at all. The single exception was `uname -m` in
runner-install.sh, used to pick a runner tarball and then discarded — so
"is this the 32GB one, or the M900?" was a question you answered by
logging in and running free -h, nproc, df -h and uname -r by hand, four
commands deep, on a machine you were already unsure about.
`rig platform` prints hostname, OS, kernel, CPU, memory, disk and
virtualization, then a provenance block: which rig, when, and the role
marker's traits.
It COMPUTES rather than stores, and that is the design rather than an
implementation detail. Specs change without rig doing anything — RAM
added, root disk resized, the unattended-upgrades bootstrap itself
enables patching the kernel — so a stored spec is stale the moment the
machine changes, and refreshing one on every run would collide with
bootstrap's "safe to re-run; a second run changes nothing" contract.
Nothing is written, so nothing can go stale.
The corollary is deliberate: reading only /proc, uname, /etc/os-release,
df and systemd-detect-virt means no root, no network, and it runs on a
pristine Debian box rig has never bootstrapped — useful for deciding
what to converge a machine into, not only for auditing it afterwards.
That also makes it the rare rig command the harness can RUN for real
rather than grep: the tests assert the answer describes the actual test
machine (kernel and hostname compared against independently computed
values), and assert it writes nothing.
Both known traps are handled explicitly. /etc/os-release is sourced in a
SUBSHELL — it defines VERSION, NAME and ID and would otherwise clobber
same-named script variables, the form every other site in this tree uses
and test/cli.sh already greps for. systemd-detect-virt exits non-zero on
bare metal while printing 'none', a normal answer that set -e would
otherwise turn into a failed run, so it is wrapped in `|| true`.
Provenance is read, never written, and degrades per file.
/etc/rig/manifest is #61 and does not exist yet, so that line reads
'not bootstrapped' on every machine today; the command ships complete
without it and neither blocks the other.
Named `platform` and not `status`: `users status` and `runner status`
cross-check recorded against live state and print DRIFT, and a command
that records nothing cannot drift, so calling it status would borrow a
promise it structurally cannot make. It also leaves `rig status` free
for the machine-wide roll-up it will eventually want to be.
Refs #64
2026-07-19 23:35:58 +00:00
|
|
|
`users status` and `runner status` cross-check recorded against live state
|
|
|
|
|
and print `DRIFT`, and a command that records nothing cannot drift — which
|
|
|
|
|
also leaves `rig status` free for the machine-wide roll-up. Known
|
|
|
|
|
limitation, stated rather than guessed at: `CPU`/`MEMORY` are read from
|
|
|
|
|
`/proc` with no cgroup awareness, and whether an `lxc` guest sees its own
|
|
|
|
|
limits or the host's totals depends on whether `lxcfs` is in play — it is
|
|
|
|
|
unverified, so those two lines are unreliable there.
|
|
|
|
|
|
feat: /etc/rig/manifest — which rig converged this machine, and when
A rig-managed machine recorded nothing about its own provenance. The entire
durable output of a bootstrap run was one line in /etc/rig/role, and that line
says what the box IS, never what built it. VERSION was read in exactly one
place (bin/rig:9, for --version) and reports the currently INSTALLED tree, not
the one that ran; there was no timestamp anywhere in the codebase.
bootstrap now stamps a second file beside the marker: schema=1, a birth pair
(bootstrapped_by/_at, pinned forever) and a latest pair (converged_by/_at).
key=value, one per line, 0644 — the one file that must stay readable on the
most broken machine in the fleet, where there is no YAML parser and no jq.
`rig manifest [<key>]` reads it back.
Only DECIDED facts go in, which is what keeps bootstrap.sh:3's convergence
contract intact: bootstrapped_* is first-write-wins, and converged_* updates
only when the version actually differs — it is the time the converging version
last changed, not the time of the last run. The renderer is pure, so a re-run
by the same rig is byte-identical no matter where the clock is, and the
cmp-guard stays silent. OBSERVED facts (cores, RAM, disk, kernel) stay out:
they go stale on their own and belong to `rig platform` (#64).
/etc/rig/role is untouched.
Closes #61
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 10:05:13 +00:00
|
|
|
- **`/etc/rig/manifest` records which rig converged a machine, and when**
|
|
|
|
|
(#61) — the entire durable output of a bootstrap run was one line in
|
|
|
|
|
`/etc/rig/role`, and that line is about what the box *is*, never about what
|
|
|
|
|
built it. `VERSION` was read in exactly one place (`bin/rig:9`, for
|
|
|
|
|
`--version`) and that reports the *currently installed* tree, not the one
|
|
|
|
|
that ran; there was no timestamp anywhere in the codebase. SSH into a
|
|
|
|
|
control plane six months on and a machine converged by `0.1.0-dev` was
|
|
|
|
|
indistinguishable from one converged by `0.4.0`. Bootstrap — both the
|
|
|
|
|
machine roles and the box tenant roles — now stamps a second file beside
|
|
|
|
|
the marker: `schema=1`, `bootstrapped_by`/`bootstrapped_at` (the rig that
|
|
|
|
|
*first* converged this machine, pinned forever) and
|
|
|
|
|
`converged_by`/`converged_at` (the newest rig to have converged it), read
|
|
|
|
|
back with a new `rig manifest [<key>]`. `key=value`, one per line, `0644` —
|
|
|
|
|
never JSON or YAML, because this is the one file that must stay readable on
|
|
|
|
|
the most broken machine in the fleet and a rig-bootstrapped box has no YAML
|
|
|
|
|
parser and no `jq`.
|
|
|
|
|
|
|
|
|
|
Only **decided** facts go in, which is what keeps `bootstrap.sh:3`'s
|
|
|
|
|
contract ("a second run changes nothing", enforced by cmp-guards at nine
|
|
|
|
|
sites) intact: `bootstrapped_*` is first-write-wins, and `converged_*`
|
|
|
|
|
updates **only when the version actually differs** — it is the time the
|
|
|
|
|
converging version last changed, not the time of the last run. A naive
|
|
|
|
|
timestamp would have made every re-run a diff and had rig report a change it
|
|
|
|
|
did not make. **Observed** facts — cores, RAM, disk, kernel — are
|
|
|
|
|
deliberately absent: they go stale without rig doing anything, so they
|
|
|
|
|
belong to `rig platform` (#64), which computes them fresh and stores
|
|
|
|
|
nothing. `/etc/rig/role` is untouched — the marker holds traits and has six
|
|
|
|
|
readers; the manifest holds provenance. Readers must ignore keys they do not
|
|
|
|
|
know, and the writer preserves lines it does not own, so a manifest written
|
|
|
|
|
by a newer rig stays readable to (and survives a rewrite by) an older one.
|
|
|
|
|
|
feat(bootstrap)!: machine roles carry a -server suffix; staging-server restored
rig builds two kinds of thing on opposite sides of a trust boundary --
tailnet machines it converges, and guests a box mints -- and both families
lived in one flat namespace with nothing in a role name saying which you
meant. `staging` is where that stopped being cosmetic: the word names the
metal that hosts guests and the guests on it, only one could have it, and
#31 gave it to the guests. The VM-host shape was left nameless, spelled
`custom --class server --host yes --join authkey`, which is what every
refusal recited at an operator who had confused the two.
The suffix now names the family: control-plane-server, workload-server,
runner-server, dev-server, plus the restored staging-server (class=server
host=yes join=authkey). host=yes already installs the box CLI and runs box's
setup-host, so staging-server is a table row, not new machinery. It stays
OUT of the tag:server allow-list deliberately -- a host is never managed by
the control plane, its guests are -- so its key is minted tag:local.
custom and workstation keep bare names as the rule, not an exception to it:
custom presets nothing and can be any shape including a guest, so a family
claim is one it cannot make; a workstation is somebody's own device, joined
by interactive login, user-owned and untagged, never tailnet-managed.
Hard cut, no aliases -- old names are refused as unknown. Two consequences
this reaches beyond the CLI surface. TS_HOSTNAME defaults to the role name,
so a box taking the default now comes up control-plane-server. And the two
coolify commands match the ROLE NAME in /etc/rig/role, not the traits, so
they now look for role=control-plane-server; a pre-rename control plane
takes their warning branch, which is advisory and never a gate, so the run
proceeds and the message names the repair.
dev-server is class=human, which reads like a contradiction and is not: the
suffix names the family, the class names the root-SSH door policy. The two
axes share the word "server", which is a real wart -- #77 renames the class
trait to what it controls, kept separate because it reaches markers on live
machines that guard root SSH.
Tests cover both directions of the cut: every new name resolves, every old
name is refused as unknown, and the two deliberately-bare roles are proven
NOT to have been swept up -- the inverse error, which would otherwise only
surface at somebody's laptop.
Closes #76 (machine-role half)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:01:16 +00:00
|
|
|
### Changed
|
|
|
|
|
|
feat(users)!: --class human|server becomes --root-door closed|open
The trait was named for who lives on a box; what it decides is whether
root SSH stays open as the control plane's automation door. Those are
different questions, and `dev-server` proved it: an unattended VM-host
appliance nobody lives on, correctly class=human because its root door
must close. After #76 gave `-server` the job of naming the machine
family, that box carried a suffix saying server and a trait saying
human. `dev-server --root-door closed` says what is true, once.
Unlike #76's role rename this field is read back on live machines, so
the compat read is mandatory rather than courteous: one resolver,
root_door_of, reads both vocabularies and every consumer goes through
it — close-root's gate, apply's note, and bootstrap-tenant's
machine-marker guard, which used the presence of `class=` as its "is
this a real fleet machine?" test and would otherwise have let a tenant
converge clobber a live box. New markers are written as `root-door=`
only. Markers carrying both fields in disagreement, or neither, fail
closed with a re-run-bootstrap repair.
Fixture markers are kept deliberately at the retired spelling (the
convention #76's pre-rename-cp fixture established) and pinned at both
consumers; deleting the compat arm turns ten checks red.
Closes #77
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 10:02:50 +00:00
|
|
|
- **BREAKING: `--class human|server` is now `--root-door closed|open`** (#77) —
|
|
|
|
|
the trait was named for who *lives on* a box; what it decides is one thing,
|
|
|
|
|
and it is not occupancy: whether root SSH stays open as the control plane's
|
|
|
|
|
automation door, or `rig users close-root` shuts it once named operators can
|
|
|
|
|
get in. The roles had been saying so for a while. `dev-server` is an
|
|
|
|
|
unattended VM-host appliance — nobody lives there, operators visit it to mint
|
|
|
|
|
boxes and leave — and by occupancy it is plainly a server. It was
|
|
|
|
|
`class=human` anyway, and correctly so, because operators enter it *as
|
|
|
|
|
themselves* and its root door must close. The trait was right; its name
|
|
|
|
|
described the wrong axis.
|
|
|
|
|
|
|
|
|
|
That stayed cheap until a second thing wanted the word "server". After #76
|
|
|
|
|
the `-server` suffix names the machine *family*, so `dev-server` carried a
|
|
|
|
|
suffix saying server and a trait saying human, and nothing in the name told a
|
|
|
|
|
reader that the two words were answering unrelated questions. `dev-server
|
|
|
|
|
--root-door closed` says exactly what is true, and `-server` means one thing
|
|
|
|
|
everywhere. The values moved with the name: `human` → `closed`, `server` →
|
|
|
|
|
`open`, and the marker field follows as `root-door=`.
|
|
|
|
|
|
|
|
|
|
Other names were considered. `--root-door open|closed` describes a
|
|
|
|
|
*destination* rather than the state at bootstrap time — bootstrap leaves root
|
|
|
|
|
SSH open on every box, and the door only shuts later, when `close-root` runs
|
|
|
|
|
— so `--root-door closes|stays` was on the table for naming the fate as a
|
|
|
|
|
verb, as was `--automation-door yes|no` for naming the thing itself. Both were
|
|
|
|
|
rejected in favour of the plainer pair: the marker is already a declaration of
|
|
|
|
|
*intent* rather than a report of observed state everywhere else in this repo
|
|
|
|
|
(`host=yes` claims a box hosts VMs; #58 settled that the marker's claim wins
|
|
|
|
|
over probing the machine), so a trait that states the door's designed end
|
|
|
|
|
state is consistent with how every other field is read. Every string that
|
|
|
|
|
prints the trait says "once operators exist" or names `close-root` explicitly,
|
|
|
|
|
so the tense never has to be inferred.
|
|
|
|
|
|
|
|
|
|
**Old markers still resolve, permanently, and that is the substance of this
|
|
|
|
|
change.** Unlike #76's role rename — role names are informational, nothing
|
|
|
|
|
reads them back — this field is written into `/etc/rig/role` and read *from*
|
|
|
|
|
there on live machines, where it gates `rig users close-root`. Every box
|
|
|
|
|
bootstrapped before this carries `class=human` or `class=server` and carries
|
|
|
|
|
it until someone re-bootstraps it, which for a fleet is never. Dropping the
|
|
|
|
|
old read would have broken in both directions at once and both are incidents:
|
|
|
|
|
a machine whose door is supposed to close loses the ability to close it, and
|
|
|
|
|
— through `bootstrap-tenant`'s machine-marker guard, which used the presence
|
|
|
|
|
of `class=` as its "is this a real fleet machine?" test — a live box stops
|
|
|
|
|
looking like a machine at all, so a tenant converge sails past the refusal
|
|
|
|
|
that exists to protect it and clobbers its marker. That second one is the
|
|
|
|
|
fail-*open* direction and was the least obvious part of the change.
|
|
|
|
|
|
|
|
|
|
So one resolver, `root_door_of`, reads both vocabularies, and every consumer
|
|
|
|
|
goes through it — close-root's gate, apply's root-SSH note, and the tenant
|
|
|
|
|
guard — because a compat read that lives at three call sites is three chances
|
|
|
|
|
to drift. `root-door=` wins where both fields are present and agree;
|
|
|
|
|
`class=` answers alone on every pre-#77 marker. A marker carrying **both and
|
|
|
|
|
disagreeing** resolves to a refusal rather than a winner: bootstrap writes one
|
|
|
|
|
line fresh and never produces that state, so a marker in it was hand-edited,
|
|
|
|
|
and rig declines to arbitrate between two equally-authored claims about a root
|
|
|
|
|
door. A marker naming **neither** refuses too, unchanged from before. Both
|
|
|
|
|
refusals fail closed, which here means the door stays open and the operator is
|
|
|
|
|
told to re-run bootstrap — never a door welded shut on a machine whose only
|
|
|
|
|
entrance it was.
|
|
|
|
|
|
2026-07-20 11:18:37 +00:00
|
|
|
The resolver matches **whole fields, not substrings** — the marker is one
|
|
|
|
|
line of space-separated `key=value` pairs, so it pads both ends and matches
|
|
|
|
|
on field boundaries. Review caught the first cut doing unanchored matching,
|
|
|
|
|
which resolved any value that *extended* a real one: `root-door=closedish`
|
|
|
|
|
read as `closed` and passed close-root's gate — the single arm that
|
|
|
|
|
authorizes an irreversible act — and `class=humanoid` did the same through
|
|
|
|
|
the compat arm, both contradicting the resolver's own promise that a value
|
|
|
|
|
outside the set resolves empty and fails closed. Only reachable by hand-
|
|
|
|
|
editing a marker, so never a live incident, but this is the one function
|
|
|
|
|
every consumer trusts and it owes them exactness rather than nearly. Both
|
|
|
|
|
vocabularies are anchored: fixing only the current spelling would have left
|
|
|
|
|
every pre-#77 box carrying the hole. Whitespace is normalised first, so a
|
|
|
|
|
hand-edit using tabs reads the same rather than trading one silent misread
|
|
|
|
|
for another.
|
|
|
|
|
|
feat(users)!: --class human|server becomes --root-door closed|open
The trait was named for who lives on a box; what it decides is whether
root SSH stays open as the control plane's automation door. Those are
different questions, and `dev-server` proved it: an unattended VM-host
appliance nobody lives on, correctly class=human because its root door
must close. After #76 gave `-server` the job of naming the machine
family, that box carried a suffix saying server and a trait saying
human. `dev-server --root-door closed` says what is true, once.
Unlike #76's role rename this field is read back on live machines, so
the compat read is mandatory rather than courteous: one resolver,
root_door_of, reads both vocabularies and every consumer goes through
it — close-root's gate, apply's note, and bootstrap-tenant's
machine-marker guard, which used the presence of `class=` as its "is
this a real fleet machine?" test and would otherwise have let a tenant
converge clobber a live box. New markers are written as `root-door=`
only. Markers carrying both fields in disagreement, or neither, fail
closed with a re-run-bootstrap repair.
Fixture markers are kept deliberately at the retired spelling (the
convention #76's pre-rename-cp fixture established) and pinned at both
consumers; deleting the compat arm turns ten checks red.
Closes #77
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 10:02:50 +00:00
|
|
|
**New markers are written in the new vocabulary only.** Writing both would
|
|
|
|
|
keep an old rig reading a new marker, but it would entrench the retired
|
|
|
|
|
spelling on every box rig ever converges and make the disagreement case
|
|
|
|
|
reachable from rig's own hand instead of only from a text editor. The compat
|
|
|
|
|
obligation runs the other way and only the other way: new rig reads old
|
|
|
|
|
markers. The bounded consequence to know about is downgrade — flipping a
|
|
|
|
|
box back to a pre-#77 rig with `rig use` leaves that older code unable to
|
|
|
|
|
recognize the new marker; the flip already WARNS on a bootstrapped host (#35),
|
|
|
|
|
and re-running bootstrap under whichever rig you settle on rewrites the line.
|
|
|
|
|
|
|
|
|
|
The suite proves the compat read rather than asserting it. Fixture markers are
|
|
|
|
|
kept **deliberately** at the retired spelling — byte for byte as a real
|
|
|
|
|
pre-#77 box reads, the same convention #76's `pre-rename-cp` fixture
|
|
|
|
|
established — and pinned at both consumers: `close-root` still passes on
|
|
|
|
|
`class=human` and still *refuses* on `class=server`, with today's refusal text
|
|
|
|
|
naming today's flag, and the tenant guard still recognizes a pre-#77 machine
|
|
|
|
|
marker as a machine. Deleting the compat arm turns ten of them red.
|
|
|
|
|
|
feat(bootstrap)!: box tenant roles carry a -box suffix
The other half of #76. claude -> claude-box, codex -> codex-box, grok ->
grok-box, staging -> staging-box, so a role name always says which family it
belongs to: -server builds a fleet machine, -box converges a guest a box
minted. With both halves in, the two families can no longer collide on a
word the way `staging` did.
The role carries the suffix; nothing inside the guest does. A tenant user is
the account the box SEED created (BOX_USER) and each agent CLI reads its own
dotdir, so claude-box still converges the `claude` user and still writes
~/.claude/CLAUDE.md. Every rename here is a $ROLE comparison or a case arm --
no CLI binary name, no dotdir path, and no account moved. README's tenant
table now shows role and user in adjacent columns, because that distinction
stopped being cosmetic the moment they differed.
Hard cut, no aliases. The old names are refused as unknown at BOTH
entrypoints -- `rig bootstrap <name>` and bootstrap-tenant.sh directly -- and
the suite asserts each of the four at each, because bootstrap.sh keeps its
own dispatch list and a name could survive in one and not the other. An alias
left in for a single tenant is the shape that survives review: the taxonomy
reads complete while one old name still quietly converges.
The consequence is cross-repo. A seed carrying BOX_BOOTSTRAP_ROLE="claude"
now fails its own mint-time bootstrap, so heavy-duty/box#123 updates the
seeds and must land after this.
Closes #76 (tenant half)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:05:19 +00:00
|
|
|
- **BREAKING: the box tenant roles carry a `-box` suffix** (#76) — the other
|
|
|
|
|
half of the rename below. `claude` → `claude-box`, `codex` → `codex-box`,
|
|
|
|
|
`grok` → `grok-box`, `staging` → `staging-box`, so a role name always says
|
|
|
|
|
which family it belongs to: `-server` builds a fleet machine, `-box`
|
|
|
|
|
converges a guest a box minted.
|
|
|
|
|
|
|
|
|
|
**The role carries the suffix; nothing inside the guest does.** A tenant user
|
|
|
|
|
is the account the box *seed* created (`BOX_USER`) and each agent CLI reads
|
|
|
|
|
its own dotdir, so `claude-box` still converges the `claude` user and still
|
|
|
|
|
writes `~/.claude/CLAUDE.md`. The suffix is rig's word for "this is a guest",
|
|
|
|
|
not a rename of anything the guest contains — no path, no account, and no CLI
|
|
|
|
|
binary moved.
|
|
|
|
|
|
|
|
|
|
**Migration: hard cut, no aliases**, same as the machine roles. The old names
|
|
|
|
|
are refused as unknown tenant roles at both entrypoints — `rig bootstrap
|
|
|
|
|
<name>` and the tenant script directly — and the suite asserts each one at
|
|
|
|
|
both, because an alias left in for a single tenant is exactly the shape that
|
|
|
|
|
survives review: the taxonomy reads complete while one old name still quietly
|
|
|
|
|
converges. The practical consequence is cross-repo: a box seed carrying
|
|
|
|
|
`BOX_BOOTSTRAP_ROLE="claude"` now fails its own mint-time bootstrap, so
|
|
|
|
|
heavy-duty/box#125 (closing heavy-duty/box#123) updates the seeds and must
|
|
|
|
|
land after this.
|
|
|
|
|
|
feat(bootstrap)!: machine roles carry a -server suffix; staging-server restored
rig builds two kinds of thing on opposite sides of a trust boundary --
tailnet machines it converges, and guests a box mints -- and both families
lived in one flat namespace with nothing in a role name saying which you
meant. `staging` is where that stopped being cosmetic: the word names the
metal that hosts guests and the guests on it, only one could have it, and
#31 gave it to the guests. The VM-host shape was left nameless, spelled
`custom --class server --host yes --join authkey`, which is what every
refusal recited at an operator who had confused the two.
The suffix now names the family: control-plane-server, workload-server,
runner-server, dev-server, plus the restored staging-server (class=server
host=yes join=authkey). host=yes already installs the box CLI and runs box's
setup-host, so staging-server is a table row, not new machinery. It stays
OUT of the tag:server allow-list deliberately -- a host is never managed by
the control plane, its guests are -- so its key is minted tag:local.
custom and workstation keep bare names as the rule, not an exception to it:
custom presets nothing and can be any shape including a guest, so a family
claim is one it cannot make; a workstation is somebody's own device, joined
by interactive login, user-owned and untagged, never tailnet-managed.
Hard cut, no aliases -- old names are refused as unknown. Two consequences
this reaches beyond the CLI surface. TS_HOSTNAME defaults to the role name,
so a box taking the default now comes up control-plane-server. And the two
coolify commands match the ROLE NAME in /etc/rig/role, not the traits, so
they now look for role=control-plane-server; a pre-rename control plane
takes their warning branch, which is advisory and never a gate, so the run
proceeds and the message names the repair.
dev-server is class=human, which reads like a contradiction and is not: the
suffix names the family, the class names the root-SSH door policy. The two
axes share the word "server", which is a real wart -- #77 renames the class
trait to what it controls, kept separate because it reaches markers on live
machines that guard root SSH.
Tests cover both directions of the cut: every new name resolves, every old
name is refused as unknown, and the two deliberately-bare roles are proven
NOT to have been swept up -- the inverse error, which would otherwise only
surface at somebody's laptop.
Closes #76 (machine-role half)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:01:16 +00:00
|
|
|
- **BREAKING: machine roles carry a `-server` suffix, and the VM host gets its
|
|
|
|
|
name back** (#76) — rig builds two kinds of thing that sit on opposite sides
|
|
|
|
|
of a trust boundary: tailnet **machines** it converges, and **guests** a box
|
|
|
|
|
mints. Both families lived in one flat namespace, and no role name said which
|
|
|
|
|
one you were asking for. `staging` is where that stopped being cosmetic — the
|
|
|
|
|
word names the metal that hosts guests *and* the guests on it, only one of
|
|
|
|
|
them could have the name, and #31 gave it to the guests. The VM-host shape
|
|
|
|
|
was left with no name at all, spelled `custom --class server --host yes
|
|
|
|
|
--join authkey`, which is what every refusal in the tree recited at an
|
|
|
|
|
operator who had confused the two.
|
|
|
|
|
|
|
|
|
|
So the suffix names the family: `control-plane-server`, `workload-server`,
|
|
|
|
|
`runner-server`, `dev-server`, and the restored `staging-server`
|
|
|
|
|
(`class=server host=yes join=authkey` — the preset #31 retired, back under a
|
|
|
|
|
name that cannot be mistaken for its own guests). `host=yes` already installs
|
|
|
|
|
the box CLI and runs box's `setup-host`, so `staging-server` is a table row
|
|
|
|
|
rather than new machinery, and it stays **out** of the `tag:server`
|
|
|
|
|
allow-list on purpose: a host is never managed by the control plane, its
|
|
|
|
|
guests are, so mint its key with `tag:local`.
|
|
|
|
|
|
|
|
|
|
**`custom` and `workstation` keep bare names**, and that is the rule rather
|
|
|
|
|
than an exception to it. `custom` presets nothing and can be any shape — a
|
|
|
|
|
guest included — so a family claim is one it cannot make. `workstation` is
|
|
|
|
|
somebody's own device rather than fleet infrastructure: it joins by
|
|
|
|
|
interactive login, comes up user-owned and untagged, and the tailnet never
|
|
|
|
|
manages it.
|
|
|
|
|
|
|
|
|
|
**Migration — this is a hard cut, with no aliases.** The old names are
|
|
|
|
|
refused as unknown roles; a box bootstrapped under one is re-bootstrapped
|
|
|
|
|
rather than migrated, which at this fleet size costs less than four
|
|
|
|
|
deprecation paths each quietly keeping an old name alive. Two consequences
|
|
|
|
|
worth knowing before you re-run anything. `TS_HOSTNAME` defaults to the role
|
|
|
|
|
name, so a box that took the default now comes up as `control-plane-server`
|
|
|
|
|
rather than `control-plane` — pass `--hostname` to hold a name steady, and
|
|
|
|
|
check anything pinning one (ACL entries, a `cast` `environments.yaml` server
|
|
|
|
|
name, host keys). And `rig coolify install` / `rig coolify backup install`
|
|
|
|
|
match the **role name** in `/etc/rig/role`, so they now look for
|
|
|
|
|
`role=control-plane-server`; a pre-rename control plane takes their warning
|
|
|
|
|
branch until it is re-bootstrapped. That check has always been advisory and
|
|
|
|
|
never a gate, so the run still proceeds and the warning names the repair.
|
|
|
|
|
|
fix(bootstrap): stop telling operators to run bare roles; pin the migration
Two review findings from the bot round on this stack.
BLOCKING (codex-bot, claude-bot -- both, independently). bootstrap-tenant.sh
emits the staging guest's tailnet-join next step at the end of a converge
("box shell -> sudo rig bootstrap workload"), repeats it in usage, and two of
its refusals recite the old machine-role list. Fixed here rather than on the
stacked tenant PR because THIS is the branch that removes the `workload` role
-- shipping it alone would print a next step naming a role that no longer
exists.
None of those four sites is code that ACCEPTS a role, which is why the rename
missed them, and is also what makes them the worse failure. A stale flag dies
immediately with a usage error. A stale next-step is copy-pasted by a human
onto a DIFFERENT box, minutes after the run that printed it reported success,
and dies there with no thread back to the cause.
So test/cli.sh sweeps every shipped script under bin/ and commands/ for
`rig bootstrap <pre-#76 name>` rather than pinning the four known sites: the
next instance of this class will be somewhere else. Proven non-vacuous --
reintroducing the bare `workload` next-step turns the suite red (412/1),
restoring it turns it green (413/0).
NON-BLOCKING (claude-bot). The migration story was documented and untested:
every marker fixture was renamed alongside the code, so nothing asserted what
a real pre-rename box does. A `role=control-plane` fixture now pins both
halves of the promise -- such a box WARNS on the coolify verbs (its marker no
longer names a role that exists) and is never REFUSED. Both halves matter: a
rename that turned this into a refusal would break the exact boxes the
CHANGELOG promises keep working, on the command that installs the control
plane.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:21:35 +00:00
|
|
|
The rename also reaches every string that *tells an operator to run a role*,
|
|
|
|
|
not just the code that accepts one — `bootstrap-tenant.sh` emits the staging
|
|
|
|
|
guest's tailnet-join next step (`sudo rig bootstrap workload-server`), and
|
|
|
|
|
two of its refusals recite the machine-role list. A stale next-step is worse
|
|
|
|
|
than a stale flag: it fails when someone copy-pastes it, on a different box,
|
|
|
|
|
minutes after the run that printed it reported success. `test/cli.sh` sweeps
|
|
|
|
|
every shipped script for pre-rename role names rather than pinning the known
|
|
|
|
|
sites, because the next instance of this will be somewhere else.
|
|
|
|
|
|
feat(users)!: --class human|server becomes --root-door closed|open
The trait was named for who lives on a box; what it decides is whether
root SSH stays open as the control plane's automation door. Those are
different questions, and `dev-server` proved it: an unattended VM-host
appliance nobody lives on, correctly class=human because its root door
must close. After #76 gave `-server` the job of naming the machine
family, that box carried a suffix saying server and a trait saying
human. `dev-server --root-door closed` says what is true, once.
Unlike #76's role rename this field is read back on live machines, so
the compat read is mandatory rather than courteous: one resolver,
root_door_of, reads both vocabularies and every consumer goes through
it — close-root's gate, apply's note, and bootstrap-tenant's
machine-marker guard, which used the presence of `class=` as its "is
this a real fleet machine?" test and would otherwise have let a tenant
converge clobber a live box. New markers are written as `root-door=`
only. Markers carrying both fields in disagreement, or neither, fail
closed with a re-run-bootstrap repair.
Fixture markers are kept deliberately at the retired spelling (the
convention #76's pre-rename-cp fixture established) and pinned at both
consumers; deleting the compat arm turns ten checks red.
Closes #77
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 10:02:50 +00:00
|
|
|
`dev-server` was `class=human` when this landed, which read like a
|
|
|
|
|
contradiction and was not: the suffix names the family, the class named the
|
|
|
|
|
root-SSH door policy, and operators enter a dev box as themselves so
|
|
|
|
|
`close-root` shuts its door. The two axes genuinely shared the word "server",
|
|
|
|
|
which was a wart — #77, above, renames the trait to what it actually controls
|
|
|
|
|
and retires it. It stayed a separate change because it reaches markers on live
|
|
|
|
|
machines that guard root SSH, and so needed a compat read this rename did not.
|
feat(bootstrap)!: machine roles carry a -server suffix; staging-server restored
rig builds two kinds of thing on opposite sides of a trust boundary --
tailnet machines it converges, and guests a box mints -- and both families
lived in one flat namespace with nothing in a role name saying which you
meant. `staging` is where that stopped being cosmetic: the word names the
metal that hosts guests and the guests on it, only one could have it, and
#31 gave it to the guests. The VM-host shape was left nameless, spelled
`custom --class server --host yes --join authkey`, which is what every
refusal recited at an operator who had confused the two.
The suffix now names the family: control-plane-server, workload-server,
runner-server, dev-server, plus the restored staging-server (class=server
host=yes join=authkey). host=yes already installs the box CLI and runs box's
setup-host, so staging-server is a table row, not new machinery. It stays
OUT of the tag:server allow-list deliberately -- a host is never managed by
the control plane, its guests are -- so its key is minted tag:local.
custom and workstation keep bare names as the rule, not an exception to it:
custom presets nothing and can be any shape including a guest, so a family
claim is one it cannot make; a workstation is somebody's own device, joined
by interactive login, user-owned and untagged, never tailnet-managed.
Hard cut, no aliases -- old names are refused as unknown. Two consequences
this reaches beyond the CLI surface. TS_HOSTNAME defaults to the role name,
so a box taking the default now comes up control-plane-server. And the two
coolify commands match the ROLE NAME in /etc/rig/role, not the traits, so
they now look for role=control-plane-server; a pre-rename control plane
takes their warning branch, which is advisory and never a gate, so the run
proceeds and the message names the repair.
dev-server is class=human, which reads like a contradiction and is not: the
suffix names the family, the class names the root-SSH door policy. The two
axes share the word "server", which is a real wart -- #77 renames the class
trait to what it controls, kept separate because it reaches markers on live
machines that guard root SSH.
Tests cover both directions of the cut: every new name resolves, every old
name is refused as unknown, and the two deliberately-bare roles are proven
NOT to have been swept up -- the inverse error, which would otherwise only
surface at somebody's laptop.
Closes #76 (machine-role half)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:01:16 +00:00
|
|
|
|
2026-07-19 23:31:40 +00:00
|
|
|
### Fixed
|
|
|
|
|
|
|
|
|
|
- **CI's shellcheck sweep now reaches `.github/scripts/`** (#70) — the step
|
|
|
|
|
ran `shopt -s globstar` and globbed `bin/* **/*.sh`, but globs skip
|
|
|
|
|
dot-prefixed names without `dotglob`, so `**/` never descended into
|
|
|
|
|
`.github/` and two tracked scripts were linted by nothing:
|
|
|
|
|
`labels-reconcile.sh` and `release-lib.sh`. The second is the one that
|
|
|
|
|
stings — it holds `changelog_section`, the extraction `release.yml` sources
|
|
|
|
|
to build the published release body and the same function `test/release.sh`'s
|
|
|
|
|
`changelog_armed` guard (#66) calls to decide whether main is armed. The
|
|
|
|
|
script deciding both what ships and whether the changelog is safe was the
|
|
|
|
|
script CI never read. Adding `dotglob` pulls in exactly those two files and
|
|
|
|
|
nothing else; both already pass, so this closes a hole in the net rather
|
|
|
|
|
than fixing a defect behind it. Paired with a class check that fails the
|
|
|
|
|
step when any tracked `.sh` falls outside the globbed set, so the gap
|
|
|
|
|
cannot reopen quietly — including via a symlinked directory, which
|
|
|
|
|
`globstar` declines to traverse.
|
|
|
|
|
|
fix: uninstall_confirm swallows Ctrl-D — the abort was silent
uninstall_confirm() read the operator's answer unguarded:
read -r reply
case "$reply" in y|Y|yes|YES|Yes) return 0 ;; *) die "aborted." ;; esac
bin/rig runs under `set -euo pipefail`, and both call sites (the
single-version and the --all confirms) invoke the function as a plain
statement — nothing suppresses errexit. Ctrl-D makes `read` return
non-zero, so the shell died AT THE READ and the case on the next line
was never evaluated: `die "aborted."` could not fire. The operator saw
the question, pressed Ctrl-D, and got nothing — no message, exit 1, at
exactly the moment the tool had asked whether to delete their install.
It failed closed, so nothing was ever wrongly removed; the damage was
that rig went silent at the one moment silence is unreadable.
The fix is `read -r reply || reply=""` — commands/db.sh:152's spelling
for the identical [y/N] confirm one file away. Empty routes through the
existing `*)` arm, so EOF aborts through the same path a bare Enter
already does: exactly one "aborted." message, no second die to keep in
sync.
test/cli.sh gains the first drills of the interactive path, which was
structurally untested (every existing uninstall check goes through
--force or RIG_YES, which is why this survived): `y` and Ctrl-D driven
through a real pty via util-linux `script`, guarded by a command -v
skip. They assert the MESSAGE, never the exit code — the unfixed code
also exits 1, so an exit-code assertion is green against the bug.
Mutation-verified: with `|| reply=""` reverted, 403 passed / 1 failed,
the single failure being `output missing 'aborted.'`; restored, 404
passed / 0 failed.
Refs #68
2026-07-19 23:32:24 +00:00
|
|
|
- **Ctrl-D at the `rig uninstall` confirm no longer aborts in silence**
|
|
|
|
|
(#68) — `uninstall_confirm`'s `read -r reply` was unguarded. Under
|
|
|
|
|
`set -euo pipefail`, and called as a plain statement, EOF made `read`
|
|
|
|
|
return non-zero and killed the shell *at the read* — the `case` on the
|
|
|
|
|
next line never ran, so `die "aborted."` never fired. The operator saw
|
|
|
|
|
the question, pressed Ctrl-D, and got nothing back: no message, just
|
|
|
|
|
exit 1, at the exact moment the tool had asked whether to delete their
|
|
|
|
|
install. It failed closed (nothing was ever removed), but nothing said
|
|
|
|
|
so. Now `read -r reply || reply=""`, so EOF falls through to the `*)`
|
|
|
|
|
arm and aborts out loud — the spelling `commands/db.sh` already used for
|
|
|
|
|
the same `[y/N]` shape, one file away. `test/cli.sh` gains the first
|
|
|
|
|
drills of the interactive path, driving `y` and Ctrl-D through a real
|
|
|
|
|
pty (util-linux `script`, skipped where it is absent) and asserting the
|
|
|
|
|
MESSAGE rather than the exit code, which the bug also produced.
|
|
|
|
|
|
fix: gate 'users apply' on an empty file that would revoke everyone
A users file naming zero users is a valid instruction to revoke every
operator on the box, and it is indistinguishable from the file a stray '>'
produces. The per-user warnings apply already emitted arrive after the
decision and scale wrong: twenty operators is twenty lines of scrollback,
so the signal was loudest exactly where it read as noise.
The /etc/rig/users ledger draws the line apply needs. An empty file against
an empty ledger is an unambiguous no-op; against a populated one it closes
every named door. Only the second now stops, states how many operators are
at risk, and requires explicit consent: --yes, RIG_YES=1 (the
installer-family variable bin/rig's uninstall_confirm already reads), or a
y on a TTY. Without a terminal and without consent it exits 2 in that same
refusal's words, rather than assume a yes it cannot ask for or hang on a
prompt nothing can answer.
A confirmation, not bootstrap's flat refusal of the same file (#57/#59):
bootstrap asserts who lives on a box, apply converges, and converging to
zero stays a legitimate de-provisioning. Ledger entries already marked
revoked do not count toward the number, so a second identical run stays the
silent no-op convergence promises.
Mass revocation below the empty-file bright line is deliberately still
ungated — that needs a threshold someone has to justify.
Refs #65
2026-07-19 23:35:42 +00:00
|
|
|
- **`users apply` now tells "revoke everyone" apart from "I truncated the
|
|
|
|
|
file"** (#65) — a users file naming zero users is a valid instruction to
|
|
|
|
|
revoke every operator on the box, and it is indistinguishable from a file a
|
|
|
|
|
stray `>` produced. The per-user warnings apply already emitted arrive after
|
|
|
|
|
the decision and scale wrong: twenty operators is twenty lines of scrollback,
|
|
|
|
|
so the signal was loudest exactly where it read as noise. The `/etc/rig/users`
|
|
|
|
|
ledger draws the line apply needs — an empty file against an empty ledger is
|
|
|
|
|
an unambiguous no-op; against a populated one it closes every named door — so
|
|
|
|
|
only the second case now stops, states how many operators are about to be
|
|
|
|
|
revoked, and requires explicit consent: `--yes`, `RIG_YES=1` (the
|
|
|
|
|
installer-family variable `rig uninstall` already reads), or a `y` on a TTY.
|
|
|
|
|
Without a terminal and without consent it exits 2, in `uninstall_confirm`'s
|
|
|
|
|
words, rather than assume a yes it cannot ask for or hang on a prompt nothing
|
|
|
|
|
can answer. A **confirmation**, not `rig bootstrap`'s flat refusal of the same
|
|
|
|
|
file (#57/#59): bootstrap asserts who lives on a box, apply converges, and
|
|
|
|
|
converging to zero stays a legitimate de-provisioning. Ledger entries already
|
|
|
|
|
marked `revoked` don't count toward the number, so a second identical run
|
|
|
|
|
stays the silent no-op. Mass revocation below the empty-file bright line (a
|
|
|
|
|
file dropping 19 of 20) is deliberately still ungated — that needs a threshold
|
|
|
|
|
someone has to justify, and #65 stays open for it.
|
|
|
|
|
|
2026-07-19 21:31:55 +00:00
|
|
|
## 0.2.0 — 2026-07-19
|
|
|
|
|
|
feat: users apply grants the box tier, not just the socket
Role `box` resolved to exactly one action, `usermod -aG incus`. That is
the socket — step 1 of the five `box grant` performs. Without the other
four (the user-<uid> project, its narrowing to boxnet and only boxnet,
the snapshot and backup allowances clone and `box export` ride, and the
shipped box-net profile installed into that project) the user's first
`box new` refuses for want of a box-net profile, so apply's promise —
the users file is the fleet's source of truth — was not kept for this
role. Worse, until an admin arrived by hand the user held an `incus`
membership with no converged project, and incus-user would lazily hand
them a stock unhardened NAT bridge: a state box's own contract forbids.
On host=yes apply now calls `box grant <user>` per box-role user. rig
calls box's grant rather than reimplementing four fifths of it — the
"rig never installs Incus" boundary is about installation, not
invocation, and grant is already script-callable: idempotent,
root-or-sudo, stdin-pinned, with its own run-as-the-user touch.
Three decisions the code carries in comment form:
- Ordering. The call sits after `useradd` (grant opens with a getent
passwd and refuses an unknown account) and after the other groups, so
a user whose grant fails still lands with everything rig owns outright.
- Failure granularity, split the way the host= guard beside it already
splits. A missing box CLI on host=yes dies, like the missing incus
group: a broken VM host, not a per-user accident. A per-user grant
failure warns and continues — one box-role user somewhere in the fleet
must not stop apply everywhere VMs don't live. host=no and marker-less
boxes keep their existing skip-with-warning untouched.
- The group ADD is deferred to grant, while `incus` stays in the wanted
set so the exact-convergence loop never strips a box-role user's
socket. Grant's rollback only reaches a membership that run added, so
rig opening the socket first would leave a failed grant unable to
close it. And grant is the authority on whether the group belongs at
all: for an incus-admin member it deliberately does not add `incus`.
An incus-admin member is warned, never fatal: box grant refuses them
today, which heavy-duty/box#99 fixes box-side with no rig change needed.
Closes #49
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:15:55 +00:00
|
|
|
### Added
|
|
|
|
|
|
|
|
|
|
- **`users apply` grants the box *tier*, not just its socket** (#49) — role
|
|
|
|
|
`box` resolved to exactly one action, `usermod -aG incus`. That is the
|
|
|
|
|
socket; it is step 1 of the five `box grant` performs, so every box-role
|
|
|
|
|
user still needed an admin to run `box grant <user>` by hand before their
|
|
|
|
|
first `box new` would do anything but refuse ("your project has no box-net
|
|
|
|
|
profile"), and until that admin arrived they held an `incus` membership
|
|
|
|
|
with no converged project — incus-user would lazily hand them a stock
|
|
|
|
|
unhardened NAT bridge, which is worse than no grant at all. On `host=yes`
|
|
|
|
|
apply now calls `box grant` per box-role user, after `useradd` (grant
|
|
|
|
|
refuses an unknown account) and with the group ADD deferred to grant, so a
|
|
|
|
|
grant that fails partway can take the socket back with it. Failures split
|
|
|
|
|
the way the `host=` guard beside them already splits: a missing `box` CLI
|
|
|
|
|
on `host=yes` dies (a broken VM host), a per-user grant failure warns and
|
|
|
|
|
continues (one box-role user must not stop apply for the fleet). `host=no`
|
|
|
|
|
and marker-less boxes keep their existing skip-with-warning. An
|
|
|
|
|
`incus-admin` member is warned, not fatal — `box grant` refuses them today,
|
|
|
|
|
which heavy-duty/box#99 fixes box-side with no rig change needed.
|
|
|
|
|
|
feat!: bootstrap takes the users file
`rig bootstrap` already knew everything else about what a box is — class,
host, join, hostname — and wrote /etc/rig/role to say so. The users file was
the last piece of that answer it did not take, so bring-up was two commands
and the second one was the forgettable one.
--users <path> now runs the `users apply` convergence as bootstrap's final
phase: after the traits, after the verified tailnet join, after the role
marker (apply reads that marker), and after the host=yes box install (so
box-role users find the incus group box's own setup-host built). One
command, and the box has its people on it.
BREAKING: --users is required on every machine role, with --no-users as the
explicit opt-out. Omitting both is a usage error naming both flags; passing
both is a usage error too. class=server is required as well: a machine
nobody logs into routinely is exactly where shared-root access rots, and
per-human accounts keep attribution intact for the times someone does go in.
The file is never persisted — passed per invocation, read once through
apply, copied nowhere. `--users -` is refused: bootstrap's stdin belongs to
the pre-auth key prompt. The box TENANT roles take neither flag; a guest is
minted non-interactively, never joins the tailnet, and has no SSH door of
its own.
rig still never installs Incus and never calls `box setup-host` itself. The
host=yes box-role precondition refuses early only where the outcome is
already proven (RIG_SKIP_BOX_INSTALL=1); every other way that step can fail
lands in `users apply`'s existing refusal, unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:17:49 +00:00
|
|
|
### Changed
|
|
|
|
|
|
|
|
|
|
- **BREAKING: `rig bootstrap` takes the users file, and requires it** (#51) —
|
|
|
|
|
bootstrap already knew everything else about what a box *is* (class, host,
|
|
|
|
|
join, hostname) and wrote `/etc/rig/role` to say so; the users file was the
|
|
|
|
|
last piece of that answer it did not take, so bring-up was two commands and
|
|
|
|
|
the second was the forgettable one. `--users <path>` now runs the `users
|
|
|
|
|
apply` convergence as bootstrap's **final phase** — after the traits, after
|
|
|
|
|
the verified tailnet join, after the role marker (apply *reads* that
|
|
|
|
|
marker), and after the `host=yes` box install (so box-role users find the
|
|
|
|
|
`incus` group box's own `setup-host` built). One command, and the box has
|
|
|
|
|
its people on it. The file is still passed per invocation and **never
|
|
|
|
|
persisted**; `--users -` is refused, because bootstrap's stdin belongs to
|
|
|
|
|
the pre-auth key prompt.
|
|
|
|
|
|
|
|
|
|
**Migration: every existing `rig bootstrap` invocation must add `--users
|
|
|
|
|
<path>` or `--no-users`.** Omitting both is now a usage error (exit 2)
|
|
|
|
|
naming both flags, and passing both is a usage error too. Scripted
|
|
|
|
|
bring-up that already ran `rig users apply` as a separate step can either
|
|
|
|
|
fold it in (`--users ./users`, and drop the separate call) or keep the old
|
|
|
|
|
shape verbatim by adding `--no-users`. Required on `class=server` as well
|
|
|
|
|
as `class=human`: a server nobody logs into routinely is exactly where
|
|
|
|
|
shared-root access rots, and per-human accounts keep attribution intact
|
|
|
|
|
for the times someone does go in — so the complete path is the default
|
|
|
|
|
path, and skipping it is deliberate rather than an omission that looks
|
|
|
|
|
identical to forgetting. The box TENANT roles (`claude|codex|grok|
|
|
|
|
|
staging`) take neither flag: a guest is minted non-interactively by box,
|
|
|
|
|
never joins the tailnet, and has no SSH door of its own — entry is `box
|
|
|
|
|
shell`, gated by the host's `incus` grants.
|
|
|
|
|
|
|
|
|
|
A bad users file is caught **up front** now (the same parser apply uses,
|
|
|
|
|
before `apt`, the hostname change, and any spent pre-auth key), and on
|
|
|
|
|
`host=yes` with `RIG_SKIP_BOX_INSTALL=1` a box-role user with no `incus`
|
|
|
|
|
group refuses immediately instead of a hundred lines later — the one case
|
|
|
|
|
where the outcome is already certain. rig still never installs Incus and
|
|
|
|
|
never calls `box setup-host` on its own account; every other way that step
|
|
|
|
|
can fail lands in `users apply`'s existing refusal, unchanged.
|
|
|
|
|
|
2026-07-19 12:15:08 +00:00
|
|
|
### Fixed
|
|
|
|
|
|
2026-07-19 19:41:49 +00:00
|
|
|
- **A release no longer disarms the changelog under the PRs still in
|
|
|
|
|
flight** (#67) — the ceremony stamps `## Unreleased` to
|
|
|
|
|
`## X.Y.Z — YYYY-MM-DD` and stops. Every PR authored before that merge
|
|
|
|
|
wrote its entry under `## Unreleased`; with the heading gone, git files
|
|
|
|
|
the entry under whatever now occupies the position — the release that
|
|
|
|
|
already shipped. There is no conflict, because the stamped heading and
|
|
|
|
|
the incoming entry never overlap textually, so the one signal an author
|
|
|
|
|
relies on ("git told me to look") is absent exactly when the outcome is
|
|
|
|
|
wrong. It happened here: #60's #58 entry landed inside `## 0.1.0` at
|
|
|
|
|
`67386b4` and was repaired two minutes later by `0ff520c`; #54 would
|
|
|
|
|
have filed a **BREAKING** entry the same way. The published release body
|
|
|
|
|
is never affected — `release.yml` extracts it from the tree at the tag,
|
|
|
|
|
before the late merges land — so the only file that drifts is the one
|
|
|
|
|
only maintainers read, which is why it survived a whole release batch
|
|
|
|
|
unnoticed. Fixed in both halves the failure has. The ceremony now
|
|
|
|
|
**re-arms**: it adds a fresh empty `## Unreleased` above the section it
|
|
|
|
|
just stamped, so a late merge has somewhere correct to land with no
|
|
|
|
|
author action. That belongs to the ceremony step in
|
|
|
|
|
[CONTRIBUTING.md](CONTRIBUTING.md), not to `release.yml` — no workflow
|
|
|
|
|
has ever touched the heading; the stamping was always by hand, and the
|
|
|
|
|
`-dev` re-arm the workflow does perform was only ever about `VERSION`.
|
|
|
|
|
And `test/release.sh` now keys its guard to `VERSION` rather than
|
|
|
|
|
demanding a literal heading: a stamped top section is legal exactly when
|
|
|
|
|
`VERSION` is bare, and the moment it carries `-dev` — main, where
|
|
|
|
|
feature PRs merge — the top section must be `## Unreleased`. That
|
|
|
|
|
distinguishes the two states the old check collapsed into one, so it
|
|
|
|
|
catches a disarmed main **without** re-breaking the ceremony's own tree
|
|
|
|
|
the way the pre-#44 guard did. The rule is proven against seven
|
|
|
|
|
constructed `VERSION` + `CHANGELOG.md` pairs, including a re-armed
|
|
|
|
|
ceremony whose top section is legitimately empty — the state the old
|
|
|
|
|
non-empty assert would have rejected. box and cast carry the same flow
|
|
|
|
|
and the same exposure (`heavy-duty/box#96`); cast is disarmed on `main`
|
|
|
|
|
as of this writing and is getting the sibling fix.
|
|
|
|
|
|
2026-07-19 17:29:20 +00:00
|
|
|
- **A `host=no` box with an `incus` group no longer hands out the bare
|
|
|
|
|
socket** (#58) — `users apply` consulted the `host=` trait only when group
|
|
|
|
|
`incus` was ABSENT (die on `host=yes`, skip on `host=no`). When the group
|
|
|
|
|
was PRESENT the trait was never asked, so a `host=no` or marker-less box
|
|
|
|
|
that nonetheless carried the group — `box setup-host` ran, then the box was
|
|
|
|
|
re-bootstrapped with other traits — gave every box-role user a bare
|
|
|
|
|
`usermod -aG incus`: the socket with no tier behind it, which `incus-user`
|
|
|
|
|
answers by lazily building an UNHARDENED project under whoever opens it
|
|
|
|
|
(`incusbr-<uid>`, NAT on v4 and v6, no ACL, no `dns.mode=none`, no port
|
|
|
|
|
isolation). The marker now decides in BOTH directions, through one new pure
|
|
|
|
|
gate (`assert_marker_hosts_vms`, testable against fixture markers non-root
|
|
|
|
|
like `assert_marker_human`): the box role applies only where the box CLAIMS
|
|
|
|
|
to host VMs, so the verdict is identical whether or not the group exists.
|
|
|
|
|
The machine deliberately does not overrule the marker — but the skip is not
|
|
|
|
|
silent either: when the group exists and the trait disagrees, the warning
|
|
|
|
|
names the contradiction and `rig bootstrap` as the repair. On such a box
|
|
|
|
|
exact-membership convergence now strips box-role users out of `incus`, on
|
|
|
|
|
the same reasoning: a membership inherited from a previous life is the same
|
|
|
|
|
half-grant as a freshly added one.
|
fix: dropping the box role revokes through box, not behind its back
`users apply` converged group `incus` with a bare `gpasswd -d`, the same
move it makes for `rig-admin` and `rig`. Those two are rig's. `incus` is
box's, and `box revoke` does strictly more with it: it says out loud that
supplementary groups are read AT LOGIN, so a session the dropped operator
already holds keeps the Incus socket until that session dies, and it hands
over `loginctl terminate-user <user>` as the remedy.
rig logged "removed <user> from incus" and moved on. An operator who
dropped someone from the users file and watched apply succeed believed the
VM access was gone — and was wrong for as long as that user held a session.
Both removal paths — the per-user convergence loop and the dropped-user
sweep — now route the incus group through one `drop_incus` helper that
calls `box revoke`, keeping a single owner for the group. Never `--purge`:
that deletes the user's boxes, images and project, and destroying someone's
running machines is not a convergence step; it stays an explicit admin act.
The exit code is not trusted (the #12 lesson bootstrap already applies to
box's installer): a revoke that returns 0 with the membership still
standing has not closed the socket, so the effective state is checked and
rig falls back to removing the group itself — as it also does on a host
where box is not installed. Every fallback path carries the session warning
in rig's own voice, because the silence was the bug. The absent-group case
needs no new guard: `id -nG` cannot report a group that does not exist, so
the existing `in_group` test at both call sites is already false on a
host=no box or one where `box setup-host` never ran.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:18:21 +00:00
|
|
|
- **Dropping the box role revokes through `box`, not behind its back**
|
|
|
|
|
(#50) — `users apply` converged group `incus` with a bare `gpasswd -d`,
|
|
|
|
|
the same move it makes for `rig-admin` and `rig`. Those two are rig's;
|
|
|
|
|
`incus` is box's, and `box revoke` does strictly more with it: it says
|
|
|
|
|
out loud that supplementary groups are read at LOGIN, so a session the
|
|
|
|
|
dropped operator already holds keeps the Incus socket until it dies, and
|
|
|
|
|
hands over `loginctl terminate-user <user>` as the remedy. rig logged
|
|
|
|
|
`removed <user> from incus` and moved on, so an operator who dropped
|
|
|
|
|
someone from the users file and watched apply succeed believed the VM
|
|
|
|
|
access was gone — and was wrong for as long as that user held a session.
|
|
|
|
|
Both removal paths (the per-user convergence and the dropped-user sweep)
|
|
|
|
|
now call `box revoke`, which keeps one owner for the group. Never
|
|
|
|
|
`--purge`: that deletes the user's boxes, images and project, and
|
|
|
|
|
destroying someone's running machines is not a convergence step — it
|
|
|
|
|
stays an explicit admin act. The exit code is not trusted (#12's lesson):
|
|
|
|
|
a revoke that returns 0 with the membership still standing has not closed
|
|
|
|
|
the socket, and rig falls back to removing the group itself, as it also
|
|
|
|
|
does where box is not installed. Every fallback path carries the session
|
|
|
|
|
warning, because the silence was the bug.
|
fix(bootstrap): refuse a users file that names no users
An empty, comments-only or whitespace-only users file is not a parse error,
so it walked straight through the requirement #51 built: pre-flight passed,
apply converged nothing, and the box came up root-only — the exact outcome
--no-users exists to make explicit, reached by the flag added to guarantee
the opposite. `--users ./empty` and `--no-users` produced the identical box
and only one of them said so.
Catch the zero-user parse in bootstrap's pre-flight, where the file is
already parsed for validation and before apt, the hostname change, or a
spent pre-auth key. The refusal names --no-users: the root-only box is
reachable, it just has to be asked for out loud.
Deliberately narrow. This is bootstrap's contract, not the parser's and not
apply's: zero users is a legal file, and a standalone `rig users apply`
against an emptied file is a real de-provisioning operation that must stay
possible. Negative-grep tests pin both.
Closes #57
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 17:28:39 +00:00
|
|
|
- **`rig bootstrap` refuses a users file that names no users** (#57) — an
|
|
|
|
|
empty, comments-only or whitespace-only file is not a parse error, so it
|
|
|
|
|
passed pre-flight, converged nothing, and left the box root-only: the exact
|
|
|
|
|
outcome `--no-users` exists to make explicit, reached by the flag added to
|
|
|
|
|
guarantee the opposite. Bootstrap's pre-flight now catches the zero-user
|
|
|
|
|
parse — before `apt`, the hostname change, or a spent pre-auth key — and
|
|
|
|
|
refuses, naming `--no-users` as the way to ask for a root-only box out loud.
|
|
|
|
|
Scoped to `rig bootstrap`'s contract only: a standalone `rig users apply`
|
|
|
|
|
against an emptied file is a real de-provisioning operation and is
|
|
|
|
|
unchanged.
|
fix: dropping the box role revokes through box, not behind its back
`users apply` converged group `incus` with a bare `gpasswd -d`, the same
move it makes for `rig-admin` and `rig`. Those two are rig's. `incus` is
box's, and `box revoke` does strictly more with it: it says out loud that
supplementary groups are read AT LOGIN, so a session the dropped operator
already holds keeps the Incus socket until that session dies, and it hands
over `loginctl terminate-user <user>` as the remedy.
rig logged "removed <user> from incus" and moved on. An operator who
dropped someone from the users file and watched apply succeed believed the
VM access was gone — and was wrong for as long as that user held a session.
Both removal paths — the per-user convergence loop and the dropped-user
sweep — now route the incus group through one `drop_incus` helper that
calls `box revoke`, keeping a single owner for the group. Never `--purge`:
that deletes the user's boxes, images and project, and destroying someone's
running machines is not a convergence step; it stays an explicit admin act.
The exit code is not trusted (the #12 lesson bootstrap already applies to
box's installer): a revoke that returns 0 with the membership still
standing has not closed the socket, so the effective state is checked and
rig falls back to removing the group itself — as it also does on a host
where box is not installed. Every fallback path carries the session warning
in rig's own voice, because the silence was the bug. The absent-group case
needs no new guard: `id -nG` cannot report a group that does not exist, so
the existing `in_group` test at both call sites is already false on a
host=no box or one where `box setup-host` never ran.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 16:18:21 +00:00
|
|
|
|
|
|
|
|
## 0.1.0 — 2026-07-19
|
|
|
|
|
|
|
|
|
|
### Fixed
|
2026-07-19 17:29:20 +00:00
|
|
|
|
2026-07-19 13:48:41 +00:00
|
|
|
- **The release suite accepts the ceremony's own tree** (#44) —
|
|
|
|
|
`test/release.sh` demanded a literal `## Unreleased` heading in the real
|
|
|
|
|
`CHANGELOG.md`, extracting non-empty and containing `#32`. All three are
|
|
|
|
|
false by construction on the `release: X.Y.Z` tree the ceremony's own PR
|
|
|
|
|
produces (it stamps that heading into `## X.Y.Z — date`), so the first
|
|
|
|
|
real release PR turned CI red and the flow blocked itself — invisible to
|
|
|
|
|
both fork rehearsals, which tag a branch (`release.yml` runs; `ci.yml`
|
|
|
|
|
never does). The guard now asserts what it was for: whatever the TOP
|
|
|
|
|
`## ` section is — `Unreleased` between releases, the stamped version on
|
|
|
|
|
and right after one — the exact `changelog_section` the workflow runs
|
|
|
|
|
extracts it non-empty. The rotting issue-number grep is gone.
|
|
|
|
|
|
2026-07-19 14:41:43 +00:00
|
|
|
- **The installer survives an environment with no `$HOME`** (#39) —
|
|
|
|
|
cloud-init's `runcmd` runs `install.sh` with no `$HOME` set, and under
|
|
|
|
|
`set -u` the first expansion died with a bash unbound-variable stack
|
|
|
|
|
instead of an install — found live by box#88's template seed, which
|
|
|
|
|
pins `HOME=/root` as its own scar. The installer now derives the home
|
|
|
|
|
from `getent` for the effective user (root included) before any path
|
|
|
|
|
is built from `$HOME`, and when getent has no answer either it refuses
|
|
|
|
|
by name. Driven with a shim getent both ways: the derived-home install
|
|
|
|
|
lands, the no-answer refusal is pinned. (#41 — merged without its
|
|
|
|
|
entry; restored here at the release gate.)
|
2026-07-19 12:15:08 +00:00
|
|
|
- **Headless credential prompts refuse loudly instead of dying silently**
|
|
|
|
|
(#42) — the interactive credential prompts (`TS_AUTHKEY` in `bootstrap`,
|
|
|
|
|
`RUNNER_TOKEN` in `runner install`, `RUNNER_REMOVE_TOKEN` in
|
|
|
|
|
`runner remove`, and both tokens in `runner repoint` — a site the new
|
|
|
|
|
no-bare-read test caught after the issue counted three) were bare
|
|
|
|
|
`read -rsp`: with stdin not a tty (CI,
|
|
|
|
|
`box exec`, any script), `read` fails, `set -e` ends the run, and the
|
|
|
|
|
log just *stops* — exit 1, no last word, measured live in the
|
|
|
|
|
2026-07-19 release drill. Each prompt now checks for a tty first and
|
|
|
|
|
dies naming the variable that unblocks an unattended run (`runner
|
|
|
|
|
remove` also names `--local`), and every `read` is `|| die`-guarded so
|
|
|
|
|
EOF at a real prompt gets the same courtesy. `db.sh` already held the
|
|
|
|
|
line here; now all of rig does.
|
|
|
|
|
|
2026-07-18 20:57:13 +00:00
|
|
|
### Added
|
|
|
|
|
|
2026-07-19 16:34:29 +00:00
|
|
|
- **Merging a release-labeled PR IS the release — and the release re-arms
|
|
|
|
|
main itself** (#47) — the rig twin of heavy-duty/box#96, born of the
|
|
|
|
|
ceremony retro: the tag was a separate, manual, silent-when-forgotten
|
|
|
|
|
step, and a forgotten tag produces no red X. `release.yml` now fires on
|
|
|
|
|
pushes to main (fork-sourced ceremony PRs get a read-only token on
|
|
|
|
|
`pull_request` events), reading the transition from the push itself:
|
|
|
|
|
`event.before` to the pushed head. A decide step answers four states —
|
|
|
|
|
release-flow *work* merged under the `release` label (`-dev` endstates,
|
|
|
|
|
the post-release window) no-ops green with a NOTICE; the two genuinely
|
|
|
|
|
ambiguous bare states refuse loudly; a true transition then requires a
|
|
|
|
|
merged, `release`-labeled PR behind the commit (read via the API — the
|
|
|
|
|
label is the operator's declared intent). Then, in the same job, it
|
|
|
|
|
API-creates the tag at the merge commit, publishes with the extracted
|
|
|
|
|
notes — and bumps main to `X.Y.(Z+1)-dev` itself, direct push with a
|
|
|
|
|
loud open-a-PR fallback, so no follow-up bump PR exists on the paved
|
|
|
|
|
road. A `GITHUB_TOKEN`-created tag never fires the tag-push trigger, so
|
|
|
|
|
the paths cannot double-publish — and that tag-push path survives intact
|
|
|
|
|
as the documented manual fallback and backfill.
|
feat: merging a release-labeled PR is the release (#47)
The rig twin of heavy-duty/box#96, from the release-ceremony retro: the
tag was a separate, manual, silent-when-forgotten step, and a forgotten
tag produces no red X — the worst failure shape. The ship decision
already lives in the release PR; merging it is "ship". After that,
tagging is transcription, and transcription belongs to machines.
release.yml now also fires on pull_request closed into main, gated on
merged AND the `release` label. The job asserts in order, each fail-loud
and creating nothing: VERSION at the merge commit is non--dev; VERSION
changed in THIS PR (base vs merge — the interlock that fails a
mislabeled ordinary PR); the changelog section for that version extracts
non-empty via the existing changelog_section from release-lib.sh; and no
tag or release exists yet. Then, in the same job, it API-creates the tag
at the merge commit and publishes the release with the extracted notes.
Same-job is load-bearing: a GITHUB_TOKEN-created tag does not fire the
tag-push trigger, so the publish must live next to the tag and the
fallback job cannot double-publish; the nothing-exists assert covers a
manual race. The tag-push path survives verbatim as the documented
manual fallback and backfill, and CONTRIBUTING's Releasing section now
reads merge-is-ship with the manual tag as fallback.
test/release.sh pins the merge path in the house grep-pin style: the
merged+labeled gate, the four asserts, the same-job tag+publish (awk
from release-on-merge: to EOF), the asserts-precede-the-tag ordering,
and the surviving tag-push trigger.
Fixes #47
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 15:19:47 +00:00
|
|
|
|
2026-07-18 20:57:13 +00:00
|
|
|
- **Tagged releases, and an installer that installs them** (#32) — the rig
|
|
|
|
|
half of the flow designed in heavy-duty/box#83, near-verbatim. A release
|
|
|
|
|
is a PR, then a tag: the `release: X.Y.Z` PR bumps `VERSION` and stamps
|
|
|
|
|
this file's Unreleased section with version + date; the merge commit is
|
|
|
|
|
tagged bare `X.Y.Z` (box's tag scheme — no `v` prefix). `release.yml`
|
|
|
|
|
turns the tag into the GitHub release — after asserting tag == `VERSION`
|
|
|
|
|
(mismatch fails loudly and creates nothing) — with that version's section
|
|
|
|
|
of this file as the body, extracted by the same `changelog_section` the
|
|
|
|
|
test harness drives. No assets: for a pure-bash tree, GitHub's source
|
|
|
|
|
tarball for the tag IS the package. `install.sh` now defaults to the
|
|
|
|
|
**latest release**: the tag is resolved by following the
|
|
|
|
|
`releases/latest` redirect and reading the `Location` header — no API, no
|
|
|
|
|
token — and the download is `archive/refs/tags/<tag>.tar.gz`. `RIG_REF`
|
|
|
|
|
picks the other two channels: a tag pins (`refs/tags` outranks a
|
|
|
|
|
same-named branch), a branch (`RIG_REF=main`) tracks the development
|
|
|
|
|
tree. Until 0.1.0 is cut the default channel has nothing to resolve and
|
|
|
|
|
dies saying exactly that, naming `RIG_REF=main` as the way to install
|
|
|
|
|
today — it never falls back to main silently, because "I installed the
|
|
|
|
|
latest release" must not quietly mean "I installed whatever main was that
|
|
|
|
|
second". Step 5 of #32 — pinning `BOX_REF` in the host-installs-box path
|
|
|
|
|
— stays open until box cuts its next tagged release.
|