-
0.8.0 Stable
released this
2026-07-19 23:04:02 +00:00 | 79 commits to main since this releaseAdded
- Merging the release PR IS the release — and the release re-arms main
itself (#96) — the 0.7.0 ceremony ended in an absence: the release PR
merged with four approvals and nothing happened, correctly, because
publishing hung off a separate, manual, silent-when-forgotten tag push —
a failure shape with no error and no red X. The ship decision already
lives in the release PR (the one PR whose whole diff is "the version
leaves-dev"), sorelease.ymlnow fires on pushes to main
(fork-sourced ceremony PRs get a read-only token onpull_request
events), reading the transition from the push itself:event.beforeto
the pushed head. A decide step answers four states — release-flow work
merged under thereleaselabel (-devendstates, the post-release
window) no-ops green with a NOTICE; the two genuinely ambiguous bare
states refuse loudly; a true transition then requires a merged,
release-labeled PR behind the commit (read via the API — the label is
the operator's declared intent) before anything is created. Then, in the
same job, it tags the merge commit via the API, publishes — and bumps
main toX.Y.(Z+1)-devitself, direct push with a loud open-a-PR
fallback, so no follow-up bump PR exists on the paved road. Same-job on
purpose: aGITHUB_TOKEN-created tag triggers no workflows, which is
also what makes double-publish impossible. The tag-push path stays
unchanged as the documented manual fallback and backfill (it shipped
0.7.0 itself).test/release.shgrep-pins the gate, every decide
verdict, the singleon.pushkey, and the same-job tag+publish+re-arm
in the same daemon-free, fail-closed style.
Fixed
-
The release ceremony re-arms
CHANGELOG.md, and CI refuses to let
mainsit disarmed (#108) — the ceremony stamps## Unreleasedinto
## X.Y.Z — DATEby hand, and nothing put the heading back, somain
sat with no## Unreleasedfrom the release until the next PR that
happened to re-create one. A PR authored before the release wrote its
entry under## Unreleased; with that heading gone, git lands the entry
under whatever now occupies the position — the section that just
shipped — and it merges cleanly. No conflict, no error, no red X:
the one signal an author would trust is absent exactly when the outcome
is wrong, and the changelog credits a released version with a change it
does not contain until a human reads the file. Confirmed in the sibling
repo (heavy-duty/rig#66); box has not drifted yet, and the reason is
luck rather than design — 0.6.0's ceremony (77599ab) added its heading
without removing## Unreleased, so main was never disarmed, while
0.7.0 did disarm it and left a window that nothing happened to cross.
Two halves land together. The ceremony step inCONTRIBUTING.mdis now
explicitly two edits: stamp, then put an empty## Unreleasedback
above the section just stamped — it belongs there and not in
release.yml, which only ever touchesVERSION. And
.github/scripts/changelog-armed.shenforces it in CI, keyed on
VERSIONbecause the two states are genuinely different: a-devtree
must carry## Unreleasedon top, a bare-VERSIONtree (the ceremony
PR, and the merge that publishes it) may carry either that or its own
stamped section. The keying is the whole design and not an
over-complication — box previously had no top-section guard at all,
and the obvious one, an unconditional## Unreleasedrequirement, is
false by construction on the ceremony PR's own tree, which is why rig#44
and heavy-duty/cast#108 both had to revert it. So a forgotten re-arm
does not block the release; it turnsmainred on the very next push,
the automatic-devbump the release itself makes. Leaving the bare
branch's top heading unconstrained is what keeps both ceremony shapes
legal, and a review round on the sibling fix (heavy-duty/cast#114) found
the gap that asymmetry leaves: a half-ceremony tree —VERSION
bumped,## Unreleasedstill populated on top, and the section for that
version never stamped — makes the wrong-number test false on its first
clause, short-circuits, and passes. Nothing then refuses until
release.ymlextracts the notes, which is after the merge, onmain,
with the release already half-shipped. So the bare branch now also
requires that the section it is about to publish exists and is non-empty,
and it asserts that by runningrelease-notes.sh— the very script
release.ymlruns — so the guard and the publisher cannot drift apart
over what a section is. The message is its own: a missing stamp is not a
misnumbered one, and an operator sent to correct a version number that is
already right will not find the real problem. Matches
heavy-duty/rig#67, so the three repos agree. -
Ctrl-D at a confirmation prompt aborts out loud, instead of exiting
in silence (#111) —confirm()anduninstall_confirm()both took
the operator's answer with a bareread -r reply. Every answer a
human can type routes through thecasebelow it and ends at a
returnor atdie "aborted."— every answer except EOF. Ctrl-D
makesreadreturn non-zero,set -euo pipefailends the run on that
line, and thecaseis never reached: box exits 1 having printed
nothing at all after the question it just asked. It fails closed,
which is why this is a small fix and not an incident — nothing is
destroyed, the abort is real. The damage is that the tool goes mute at
the one moment it had the operator's full attention, and someone who
Ctrl-Ds out ofbox rm workcannot tell from the output whether the
box is still there. The cure is one token in each function,
read -r reply || die "aborted.", the same one heavy-duty/rig#43
applied to rig's credential prompts so the two repos read alike. The
bug predates everything it touches —rmhas carried a confirm gate
for as long as the verb has existed — but #105 took the number of
verbs reaching that line from one to two, and both are irreversible,
which is the argument for closing it now rather than the next time
someone notices. The three answers a human can actually give (y,
n, and Ctrl-D) are now driven for real on a pty via util-linux
script: they were structurally untested before, because[ -t 0 ]
sends a terminal-less suite to the refusal branch and every existing
check stopped there — which is exactly how this survived four
releases. Review caught that the first pass fixed the bug where it was
reported and stopped there, while the same defect sat at two more
destructive gates in this repo:host/revoke-user.sh:50, the prompt
guardingbox revoke --purge— the one whose own text says "this
cannot be undone" — andhost/teardown-host.sh:31, guarding a full
host teardown. Both run underset -euo pipefail, both died mute on
EOF with theirabortedline never reached; both now carry the guard
in their own script's wording. The threedrill/prompts are
deliberately left alone — they run underset -uonly, so EOF falls
through to the*)arm and already aborts out loud — and
install.sh:65was already guarded. What keeps the class closed is a
repo-wide sweep intest/cli.sh: every statement-initialreadfed
from stdin, in any file that turns on errexit, must carry a||
guard, withwhile readloops and<<<herestrings excluded because
neither is a prompt. The sweep flags all four sites when their guards
are removed and nothing else across the tree's fifteen shell files —
the absence of exactly this check is why thehost/pair was missed
in the first place. -
box restoreasks before it destroys — and the confirmation prompt is
now the row's, not rm's (#105) —restoreandrmboth irreversibly
discard user state, and only one of them asked. The table gaverestore
the preconditionsbox,arg2: the instance is ours, a snapshot name is
present, go. Sobox restore work stale-labelsilently threw away
everything done in the box since that snapshot, with no prompt, no
--force, and no way to take it back — a warning in--helpis not a
gate. It has been that way since the verb shipped, and it is about to
become routine rather than rare (heavy-duty/rig#62's pristine snapshot),
which is the wrong time to still be relying on the operator typing the
right label. The reason it stayed ungated is worth recording, because it
is the actual bug:confirmwas already a precondition token, but the
dispatch line hardcoded the words —confirm "delete $inst and all its snapshots"— so the one-token fix would have gated restore behind a
prompt offering to DELETE the box the operator was trying to rescue. A
gate that names the wrong act is worse than no gate; it is how people
learn to answerywithout reading. So the prompt moved into the table
as a seventh field, each row saying what it is about to do in its own
words, andrestorenow asks to "roll<box>back to snapshot
<label>and discard everything in the box since it was taken" — naming
the label, because picking the wrong one is the whole risk.rm's
wording is unchanged and pinned verbatim by a test, since rewording the
one verb that already worked would be a regression shipped as a
refactor. A row markedconfirmwith no words is now a hard internal
error rather than a blank question.--forceand the no-TTY refusal come
free —confirm()already had both. The one automated caller had to
consent explicitly:drill/multiuser.shdrives restore unattended on real
Incus and now passes--force, which is the rehearsal proving the gate
rather than working around it — the CI run of this very PR failed there
first, which is the shape a gate is supposed to have. Coverage went from two
argument-validation checks that never reached dispatch to the destructive
path itself, driven against a fake incus: refusing leaves the call log
empty,--forceproduces exactly oneincus snapshot restore. Not
changed, deliberately:restorestill does not require the box stopped
(#105 makes that case separately and it deserves its own call), and
--helpnow says plainly that a rollback of a running box is
crash-consistent, because these snapshots are stateless. -
box-firewallcould hand a UFW host the no-UFW firewall, ~2% of the
time (#102) — filed as an intermittent test flake (test/cli.sh's
fresh-UFW block going four-assertions-red on an unmodifiedmain,
measured here at 5 failing runs in 40), it was not one. The branch that
decides the host's entire firewall stance read
ufw status | grep -q "Status: active", andStatus: activeis the FIRST
line ufw prints:grep -qmatches it and exits immediately, closing the
pipe while ufw is still writing the rest of the table, so ufw dies of
SIGPIPE.grepreturned 0, but under this script'sset -o pipefailthe
PIPELINE returns 141 — theifreads false and a host with UFW plainly
active takes the nft-fallback branch, never building the DNS carve-out its
persisted rules depend on. A pure scheduling race, isolated at ~2% per
invocation (PIPESTATUS=141 0; a draining reader flakes 0/2000, a
reader whose match is on the last line flakes 0/2000). Real ufw is a
slower, longer writer than the test shim, so production had no reason to
be safer.ufw statusis now read ONCE into a variable and matched with
[[ ]]— no reader, no race — and the stale-rule scan reads that same
snapshot, so the branch decision and the converge loop can no longer
disagree.host/teardown-host.shcarried the same live defect and is
fixed with it: that file does setpipefail(line 12), so its UFW
crumb-removal branch could read a plainly-active UFW as inactive and skip
silently, leaving staleboxnet/claudenetrules on a host the operator
was told is clean — and its numbered-delete loop had the same early-exit
reader as its condition, so it could end while rules remained. Both now
read captures. The sibling calls indrill/wipe.shanddrill/doctor.sh
are the same shape but set onlyset -u, so the SIGPIPE is discarded
there and the branch holds — latent, not live, until either gains
pipefail. -
A missing firewall log now diagnoses itself (#102) — the four greps
reading$WFW/*.logused to fail together with empty output when the
driving run took the wrong branch, a signature that looks specific and
says nothing (#102 was filed reading it as "the log is not written";
the log existed, the mutations did not, and that distinction was the
diagnosis).test/cli.shnow asserts the precondition explicitly before
the content greps and, on failure, prints the contents of$WFW, the log
itself, and the stderr of the run that should have written it. It also
keepsan agreeing UFW host deletes nothinghonest: that check asserts an
absence, which a run that did nothing at all passes for the wrong reason. -
box grantprovisions anincus-adminmember instead of refusing them
(#99) — the refusal read "they already have the admin tier; there is
nothing tighter to grant", which is true about permission and silent
about provisioning: theincusgroup is indeed a strict subset of what
incus-adminopens at the daemon API, but theuser-<uid>project, the
boxnet narrowing, the snapshot and backup allowances, and thebox-net
profile installed into that project are none of them permissions, and an
incus-adminmember had none of them —box_tier()resolves them to
admin, so they worked in the shared default project next to root and every
other admin, with no world of their own and no supported way to get one.
box grantnow runs the full convergence for them.The group step is part of that convergence, not an exception to it: an
incus-adminmember is added toincuslike anyone else. The subset
argument holds for the API and fails at the filesystem, which is where
it matters here — the two sockets are two files with two owning groups
(Debian 13 / Incus 6.0.4, measured):socket group mode /var/lib/incus/unix.socketincus-admin0660 /var/lib/incus/unix.socket.userincus0660 incus-adminopens the first and not the second, and only the second
provisions auser-<uid>project. Without the membership the provisioning
touch takesEACCES, the swallowing|| truehides it, no project appears,
and the grant dies blaming a perfectly healthy incus-user — the exact
incus-admin-only user #99 is about, left no better off. So the membership is
granted, and the grant says out loud why: it is the key to a file, not a new
privilege (box_tier()still reads them asadmin, both-groups →admin).The touch itself is pinned at incus-user's socket: the incus client picks
by writability (client/connection.go— the daemon socket when writable,
unix.socket.useronly otherwise), so for anincus-adminmember an
unpinned touch sails past incus-user and provisions nothing. The user-side
proof that closes the grant names their project for the same reason, since an
unqualifiedprofile showwould have answered from the shared default
project and proved nothing. The socket's existence is probed through$SUDO,
not a bare[ -e ]—/var/lib/incusis not traversable by a non-root
admin, so an unprivileged stat reports a present socket as absent, and this
probe exits on absent (the disciplinebox revokealready documents).On success the grant prints the caveat the hard exit was gesturing at, in the
two forms it actually takes: the restrictions are a default placement, not
a confinement (admin membership still wins at the socket — the default
project and other users' instances stay one flag away), and until
incus-admingoes their ownboxcommands keep landing in the default
project. Droppingincus-adminthen lands them in their ready project with
no re-grant — a promise that is only true because they keepincus;
without it that drop would leave them in neither group,box_tier()none,
and a converged project they could not open. The failure path follows: the
membership this run added is rolled back and verified, while the backout
refuses to call that a lockout —incus-adminis untouched and still opens
every project.box revokemirrors it. A bare revoke of a grantedincus-adminmember now
takes theincusmembership back and reportspartial:— the socket key
box grantadded is gone, their project is kept, and they are explicitly
not locked out. Anincus-adminmember who was never granted is still a
named no-op that makes no privileged call at all.--purgeunmakes the
provisioning while refusing to call them "out". Every path names
gpasswd -d <user> incus-adminas the only thing that ends their access.Unblocks rig's
users apply(heavy-duty/rig#49), which had to callbox grantfor a user who is bothincus-adminby hand and roleboxin the
fleet file. Driven end to end intest/cli.shunder logging incus/sudo shims
— every assertion is made against what the run did, not what the source says
it would — and, because those shims model neitherINCUS_SOCKETnor file
permissions and so cannot reproduce theEACCES, measured on real Incus in
CI by a newdrill/multiuser.shcriterion (o): anincus-admin-only member
is granted, the membership lands, the project appears,unix.socket.user
opens as them, and droppingincus-adminleaves them in their own project
with no re-grant.
Downloads
-
Source code (ZIP)
1 download
-
Source code (TAR.GZ)
0 downloads
- Merging the release PR IS the release — and the release re-arms main