Fleet scope and cross-repo discovery — the two guards, the runner hole, the roster question (epic) #56
Labels
No labels
attention
blocked
blocker:ci-red
blocker:conflict
blocker:drill-pending
blocker:unrequested
bug
claimed
documentation
enhancement
epic
merge-next
needs-ruling
needs-triage
offsite
post-merge
ready
release
scope:docs
scope:guards
scope:labels
scope:release-flow
stale
state:addressing
state:bots-reviewing
state:building
state:needs-human
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference: heavy-duty/ceremony#56
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Accepted from discussion #55 — "Where the fleet may work — three scopes, two guards, one registry", filed at @danmt's request. This epic carries the decisions triage made in accepting it, the corrections that verification forced on the proposal's own evidence, and the one question that is not triage's to answer.
All line references pinned at
2f0d3c6.Context — the incidents, re-verified
The proposal named two failures. Both are real; neither is quite the failure it was reported as, and the difference changes what the guards must say. Every claim below was checked against the API, not taken from the report.
dan-claude-bot/incubator#89"has sat since 01:20Z with zero reviews and zero requested reviewers" — a discovery failureheavy-duty/rig#112because its poll list isheavy-duty/ceremonyalonepanel=claude-bot-andresmgsl codex-bot-andresmgsl grok-bot-andresmgsl: kimi is not on rig's panel at all. Even with perfect discovery, no verdict of kimi's would ever have been required thereready_for_review+codex+grokby the builder at 00:34Z, thenkimirequested by @danmt at 01:24Z, 50 minutes later. But by rig'spanel=line the builder was correct: panel minus author = codex + grokNine hours on, rig#112 still has no verdict from kimi, and both required verdicts (codex, grok) approve the current headClosed out (triage, 2026-07-23): kimi approved3c72c1b.3c72c1bat 10:42Z — 9h18m after @danmt's off-roster request at 01:24Z. All three verdicts now approve the head, and rig#112 has carriedstate:needs-humansince 10:47Z. The incident is over; what it evidenced is not — the delay is measured, not hypothetical, and neither guard nor ruling exists yet.So: one closed discovery failure that took 9h18m to produce an off-roster verdict (the delay is the evidence; D1–D4 stand on it), one live ambiguity (which roster is "the panel" — D7, unruled), one live hole with no incident yet (the self-hosted runner), and one proposal artifact — the registry — that has exactly one true row and no readers.
Decisions — children must not reopen these
REQUIRED_BOTS=panel=minus author, labels-reconcile.sh L38-L47). An off-panel reviewer's verdict is advisory and says so in its body. This is not a new rule — it is what the reconciler already computes; doctrine has simply never said it, which is why rig#112 now reads as if it were waiting on a verdict no machine will ever requirepull_request-triggered work must never reach a self-hosted runner. Verified live:heavy-duty/incubator'sdeploy.ymljob runs on[self-hosted, ci-runner]and ispush-triggered;pr-checks.ymlispull_request-triggered andubuntu-latest. Nothing is wrong today — the guard is what keeps it that way once fork PRs run workflows there (#16's ruling)FLEET.ymlis not minted now. Today it would carry oneadopted: truerow (ceremony), four rows that restate #13–#16, and no reader — every box's duty script is operator-owned and would need changing to read it. Revisit when #13 lands and there are two adopted repos to disagree about. A registry minted before it has readers is a file that goes stale in a weekoffsite, named by @danmt. The failure it fixes is real and mis-attributed in the proposal: the reclaim that runs on this board is the in-repo sweep'sclaim_decision, not the triage box'shygiene.sh. The box script is the operator's and gets a spec, never a claim of done — the same constraint FLEET.md carries above. #68offsitedoes not touch the epic-completion nudge — not as a preference but because there is no interaction to have. The nudge fires on children being closed; an offsite child is open. Recorded so nobody builds the suppression @danmt left to triage's judgementConstraints the children must preserve
FLEET.mdis descriptive, not doctrine (its own status header). The duty scripts that implement wake conditions live inside each box and are the operator's to change. A FLEET.md edit is a spec for that change, never a substitute for it — and a child that claims a wake condition works because a document says so is lying..ceremony/set; this repo is the source, so no re-sync happens here — consumers get the change at their next pin bump (CONTRIBUTING, "How the other repos use this").actions/<name>/action.yml+<name>.sh+test/<name>.test.sh, mawk-compatible,set -euo pipefail, shellcheck- and actionlint-clean.Task list
5e283ec), closed 11:37Z. The box-side implementation of the reviewer wake condition remains the operator's, as the constraint above requires — FLEET.md specifies it, it does not claim it done.actions/runner-isolated: nopull_request-triggered job on a self-hosted runner. Landed 2026-07-23 in #60 (merged 12:24Z), closed the same minute. Red-run evidence on a scratch fork branch (run 30002684920), and ceremony runs the guard on its own tree (ci.yml L77).offsite: the label, the doctrine, and the claim-reclaim exemption. The machine-readable half of #57's cross-repo linkage rule — a claim whose PR lives in another repo has no local PR by construction, so the reclaim clock reads it as abandoned. Landed 2026-07-23 in #70 (e9928f1), closed 13:07Z. Verified onmain:claim_clock_exemptpauses only the reclaim clock —test/issueflow-reconcile.test.shpins "a 10-day-quiet offsite claim is not reclaimed", "an unassigned offsite claim is still flagged", "offsite alone still needs triage", and "no reconciler mutation names offsite (#68 D4)". Accepted from discussion 67 (D8–D10 below).offsitestale-flag nudge: when every cross-referenced PR the sweep can see is closed, say so once. Reads the timeline the sweep already fetches; comments and never acts. Landed 2026-07-23 in #71 (553409c), closed 13:43Z.offsite_resolved_decisionnudges on all-closed and stays quiet on an open or unreadable one ("one open offsite PR keeps quiet", "an unreadable offsite PR keeps quiet"); the flag is never cleared and the claim never changed.offsitelabel does not exist on this board yet — a maintainer must dispatch the labels workflow. The taxonomy row landed with #70 (labels-reconcile.shL394,offsite|CFD3D7|…), but the bootstrap only writes labels when dispatched. Triage cannot do it:POST /repos/heavy-duty/ceremony/labelsreturns 404 attriagepermission (attempted 2026-07-23). Until it runs, the three live offsite claims — #14 (box#164), #15 (cast#143), #16 (incubator#25) — carry the exemption's code and not its flag, so the 48-hour reclaim clock still reads them as local claims with no PR. Same shape as #50's post-merge dispatch item, same one-dispatch fix. @danmt.FLEET.yml) — deliberately not minted (D6), and gated on D7's ruling besides. If the ruling enumerates build targets rather than openingheavy-duty/*, the registry becomes load-bearing and gets minted then; if the boundary is the org, it may never be needed. Tracked here so the omission is a decision on the board rather than a gap in it.Flag status (triage, 2026-07-23): this epic now carries
needs-ruling. D7's escalation has been live and unruled since 10:46Z; the label was unappliable until the bootstrap dispatch ran at 11:48Z, which is why it arrives an hour late rather than with the escalation. Nothing is blocked on the ruling — the two discovery children have landed, the twooffsitechildren are queued behind #52, and D6 holds independently of all four.Definition of done
5e283ec): BUILDER.md names the PR repo'spanel=as the answer, CONTRIBUTING as its human-readable form,panel=governing on disagreement, and ask triage, do not guess when the PR repo names no roster; the linkage rule (Part of <owner>/<repo>#N+ comment the draft link on the authorizing issue, triage closes by hand) is at L26-L35.5e283ec): REVIEWER.md L46-L55, carrying D3 verbatim and citing rig#112 as the case that proved authorization and membership must not be conflated.5e283ec): FLEET.md L57-L68 — request-first ordering with its reason (it reaches repos the list does not name), dedup against the reviewer's own latest review SHA rather than the lagging search index, and the explicit "until an operator makes them, the request trigger exists on paper only".pull_request-triggered job that runs on a self-hosted runner fails CI in any repo carrying the guard, proven red once on a scratch fixture, and ceremony runs the guard on itself. — done in #58 (#60, merged 12:24Z): the guard is red on claude-bot-andresmgsl/ceremony#2's scratchruns-on: self-hostedworkflow (run 30002684920), caught twice — once by theself-guardsstep, once independently by the unit suite's own-tree case — and.github/workflows/ci.ymlruns./actions/runner-isolatedon ceremony itself.claimedissue whose deliverable is a PR in another repo is not reclaimed as abandoned, and a flag that outlives its PR is surfaced rather than left standing — done in #68 (#70,e9928f1) and #69 (#71,553409c), verified onmainagainst the suite rather than off the PR bodies. One caveat, tracked as an open task above: the behaviour is real in the code and unreachable on this board until a maintainer dispatch creates theoffsitelabel, so today's three cross-repo claims are still on the ordinary reclaim clock.Escalation — @danmt, three rulings. This is D7 of the epic above, raised under the escalation contract (LABELS.md): the question, the options, a recommendation, and the decider named. Nothing is blocked on you — #57 and #58 are
readyand a builder can start either right now. R2 is the one worth reading first: a live PR in rig is waiting on it.R1 — Where may a builder build? (discussion #55, open question 1)
Today there is no rule; each box answers from its own duty script.
heavy-duty/*repo, authorised by a ceremony-minted issue that names the target repo. The issue is the token — the pattern #13 and #16 already ran.Recommendation: (a). The authorisation already exists per target, in a reviewable artifact with a spec attached; (b) copies it into a second place that can disagree with the first, and the copy is the thing that goes stale. What actually went wrong in incubator was not authority — it would have been granted in ten seconds — it was a whole build cycle spent against a dead namespace. An enumeration does not check for informedness; an issue that names the target does.
R2 — Is the reviewer bench fleet-wide, or per-repo? ⏳ rig#112 is waiting on this
Not in the proposal — it surfaced while verifying it, and it is the sharper half of the kimi incident. ceremony's
panel=is four bots; rig's is three — kimi is not on it. You requested kimi on rig#112 by hand at 01:24Z, which reads like you expect the bench to be the same everywhere. The machine disagrees: required verdicts arepanel=minus the author, so kimi's verdict there is not required and never was.panel=carries the same four bot identities; the panel is that list minus the author. Applied as a one-line change per repo, at its next pin bump.Recommendation: (a). CONTRIBUTING claims convergence "always means three cross-vendor approvals of the current head" — that claim is only true if the bench is the same everywhere. rig's three-name panel is a pre-conversion artifact, not a decision anyone made.
The consequence is immediate, which is why this is first: under (a), kimi's verdict becomes required on rig#112 and that PR is not converged — it is waiting on a reviewer that cannot see it until #57's wake condition reaches kimi's box. Under (b), rig#112 is converged right now on codex + grok (both approving head
3c72c1b, no blocker standing), and your kimi request is advisory. One word from you unsticks it either way.R3 — Fork-PR workflow settings: org-wide, or per repo?
You ruled this for incubator on #16 — option (a), enable, with "send write tokens" and "send secrets" both off. The proposal asks to make that org-wide, and adds "require approval off".
Recommendation: (b). Approval-off is safe because the other two are off and nothing
pull_request-triggered touches the self-hosted runner — and that second half is currently a sentence inpr-checks.yml's header, kept true by whoever remembers it. #58 turns it into a gate. The cost of (b) is one approval click per fork PR in incubator until #58 lands and is adopted there; the cost of getting it wrong is unreviewed fork code executing on the tailnet runner. If you prefer (a), say so and it is done — the exposure is small and the org is six accounts — but it should be a decision, not a default.Answer inline, in any order — no format required. I judge when agreement is reached, record each ruling as a decision in the epic's D-table, and return the epic to its flow in the same comment (D6/D7 of #50).
Triage:
needs-rulingis now on this epic — the flag, not a new question.D7's escalation to @danmt has been live and unruled since 10:46Z: three questions (the build-scope boundary, roster uniformity, org-wide fork-PR settings), each with options and a recommendation. It carried the full D4 contract from the moment it was posted; what it did not carry was the label, because
needs-rulingdid not exist on this repo until the bootstrap dispatch ran at 11:48Z. That gap was tracked on #50's task list, and applying the flag here the moment the dispatch landed was the item's stated consequence.So: nothing new is being asked, and the ball has not moved — it was already yours. The board now says so, which is the whole point of the label (LABELS.md). Per D5/D6 it stays up until agreement is reached, not until you reply, and triage — as the setter — closes it out: the ruling recorded here as a decision in one comment, the label removed, the epic returned to flow, all in that same comment.
Nothing is blocked on it. Both children are in flight (#57 landed in #62; #58 is claimed with #60 open), and D6 keeps
FLEET.ymlunminted independent of the ruling. What the rulings gate is whether the registry ever becomes load-bearing, and whether the bench roster is uniform across governed repos.Every issue referenced by this epic's task list is closed. Please close the epic or extend its task list.
Triage: nudge answered — the epic is extended, not complete.
The sweep was right that every issue on the task list had closed. #58 landed in #60 at 12:24Z, joining #57; both definition-of-done items they owned are now checked with their evidence.
Two new children take their place, accepted from discussion 67:
offsite: the label, the doctrine, the claim-reclaim exemption.blockedby #52.blockedby #68.They belong here rather than under #50 because they finish what #57 started: #57 gave cross-repo work its human-readable linkage, and the sweep still cannot read it. #13 and #16 are both claimed with no local PR by construction right now, and only today's comment traffic is keeping them off the reclaim clock.
Three decisions recorded in the D-table above, D8–D10 — the label is exempt on trust alone, verification only ever nudges, and the epic-completion nudge has no interaction with
offsiteto build.needs-rulingstays up. D7's three questions (10:46Z) are still unanswered, and the flag clears on agreement, not on activity elsewhere. @danmt — R2 is the one with a queue behind it.Triage, epic hygiene — the task list's #68 line was stale, no scope change.
It still read "
blockedby #52". #52 landed at 12:41Z in #64 (dfcfd45) — the rewrite of the same call site — so the sweep cleared #68 toreadythat minute and @codex-bot-andresmgsl claimed it at 12:46Z; PR #70 is open and in review. The line now says that.#69 stays
blocked— verified, not assumed: its blocker is #68, which is open. The D7 escalation to @danmt from 10:46Z is still unruled, soneeds-rulingstays up on this epic.Triage, epic hygiene — both
offsitechildren have landed; one operator action now blocks the flag from working on this board.Checked against
mainand the live board, not off the PR bodies.e9928f1).claim_clock_exemptpauses only the reclaim clock — missing assignees, queue conflicts, epic completion and the PR side are untouched, pinned by "an unassigned offsite claim is still flagged", "offsite alone still needs triage" and "no reconciler mutation names offsite (#68 D4)".553409c). The nudge fires only when every cross-referenced PR the sweep can see is closed; an open or unreadable one keeps it quiet, and it never clears the flag or touches the claim.Both task-list lines and the fourth definition-of-done item are checked with that evidence.
@danmt — one dispatch is owed, and it is the same one #50 needed. The
offsiterow is in the taxonomy onmain(labels-reconcile.shL394) but the label does not exist on heavy-duty/ceremony — the bootstrap only writes labels on manual dispatch, and nothing has dispatched since #70 merged. I cannot create it:POST /repos/heavy-duty/ceremony/labels404s attriagepermission (attempted today).The cost of waiting is concrete. Three claims are offsite right now — #14 (box#164), #15 (cast#143), #16 (incubator#25) — and none can carry the flag, so all three sit on the ordinary 48-hour reclaim clock with no local PR to save them. Every one is active today, so nothing is at risk this minute; a quiet night is what would cost a builder their claim. This is not a ruling and it does not touch D7 — it is a button.
D7 stays unruled and
needs-rulingstays up, unchanged since 10:46Z.what needs ruling? whats the choice?
Closing: ceremony's half of this epic is done. What remains is not ceremony's
to answer, and it moves to the
crewrepo @danmt is standing up.And answering the question directly, @danmt — what needs ruling, what's the
choice? Two of the three no longer need you at all.
What landed
5e283ec) — BUILDER.md's panel + linkage rules, REVIEWER.md's "a request is authorisation, not membership", FLEET.md's request-before-poll orderingactions/runner-isolated, proven red on a scratchruns-on: self-hostedfixture, and ceremony runs it on itselfoffsitein the taxonomy, the staleness skip atissueflow-reconcile.shL109, and the trust-first verify pass at L195–316 that nudges once when every referenced PR has closed and never clears the flag itselfThe
offsitelabel is now live on this board — @danmt dispatched the labelsworkflow at ~14:15Z, and it exists as
offsite | CFD3D7 | Issue deliverable is a PR in another repository — claim clock paused. That was the last open taskhere. The three live cross-repo claims — #14 (box#164),
#15 (cast#143), #16
(incubator#25, merged) —
still need the label applied; that is a reconciler tick, not work, and it is
triage's.
Nothing is missing from ceremony's side. No follow-up issue filed. The gap
check was the point of this pass and it came back clean: the exemption is real
in the code, not just declared in a table.
The three rulings, as they stand now
R2 — bench fleet-wide or per-repo? Overtaken by events. Its forcing case was
rig#112 waiting on kimi. Kimi approved at 10:42:29Z once the request trigger
reached its box, and @danmt merged at 11:58:33Z with all three verdicts in. The
PR converged under either answer, so nothing is stuck. The general question
survives — is
panel=uniform across governed repos — but it is now a fleetquestion with no live casualty, which is exactly the kind that belongs in crew
alongside the roster it describes.
R3 — fork-PR settings. Its precondition landed. The recommendation was (b),
keep require approval ON until #58's guard exists, because "nothing
pull_request-triggered touches the self-hosted runner" was a sentence in aheader rather than a gate. #58 merged at 12:24Z and is that gate. So (a) —
org-wide, write tokens off, secrets off, approval off — now costs what (b) was
protecting against, which is nothing. Both repos in question are flipped
already; canonical turned out never to have been off. This needs recording, not
deciding.
R1 — where may a builder build? Genuinely open, and better answered in crew.
The choice is unchanged: (a) any
heavy-duty/*repo authorised by an issue thatnames the target, or (b) an enumerated registry. Triage recommended (a) and I
agree — an enumeration is a copy of the authorisation that can disagree with it.
But R1 and D6 (
FLEET.yml, deliberately not minted because "today it'd haveone real row and no reader") are the same question wearing two hats, and crew
gives both an answer: it is the roster, and the duty loops that would read
FLEET.yml live in it.
There is a second reason to move rather than force it. R1 and R2 are questions
about how the fleet works, and the fleet's actual behaviour currently exists as
five untracked
~/dutydirectories on five disposable boxes that nobody hascompared. Ruling now means ruling with less information than a week from now,
when those are in git and readable side by side.
Handoff to crew
Carried over, and nothing else: R1 (build boundary), R2 (bench
uniformity), D6 (
FLEET.yml— the registry, its readers, its schema).needs-rulingcomes off this epic with this comment: the flag marked a decisionblocking this board, and this board is no longer where the decision lives.
What ceremony keeps is what passes the test crew's split is drawn on — could a
team with no agents adopt this? The release workflow, the guard family
including
runner-isolated,labels-reconcileand its taxonomy, the doctrinein LABELS.md and CONTRIBUTING: yes. The roster, the wake conditions, the duty
loops: no, and those go.