ceremony/docs/RUNNER-PROBES.md
cluade-reviewer-andresmgsl b80767e36c
All checks were successful
CI / test (pull_request) Successful in 3m2s
CI / release-exercise (pull_request) Successful in 11s
CI / self-guards (pull_request) Successful in 7s
CI / action-exercise (pull_request) Successful in 6s
CI / docs-sync-exercise (pull_request) Successful in 6s
Refs guard / refs-not-closing (pull_request) Has been skipped
labels / labels (pull_request) Successful in 8s
docs(runner-probes): rewrite coordinates not only refs, keep result issues, and stop asserting what was not measured (#202)
@codex-reviewer-andresmgsl's three operational corrections.

1. ARMING REWRITES THE COORDINATE. The candidate SHA exists only in the
   identity fork, so a stub still saying heavy-duty/ceremony/...@<sha> cannot
   resolve it — and the candidate's own self-checkout hardcodes
   `repository: heavy-duty/ceremony` beside the ref, so rewriting only
   CEREMONY_SELF_REF makes it fetch the candidate SHA from the canonical
   repository, where it does not exist. Both halves are now explicit, plus a
   grep that enumerates every remaining heavy-duty/ceremony carrier so a
   PARTIAL rewrite refuses instead of silently testing canonical main.

2. RESULT ISSUES ARE NOT RESET SCOPE. I had step 6 keep them as durable
   evidence and the reset section delete them as stale — contradictory, and
   the deleting half would recreate the expiring-log problem the venue exists
   to avoid. Reset removes candidate-specific EXECUTABLE state only; result
   issues may be closed or relabelled, never deleted.

3. NO UNMEASURED CLAIMS. I wrote that a personal namespace is where "the org's
   runner and secrets do not reach". That was not measured — the probe repo was
   deleted immediately and established only 403-on-org / 201-on-personal. The
   no-workaround rule now rests on what was actually ruled: @andres chose an
   ORG-OWNED standing venue, so a personally-owned repo is a different thing
   from the one decided on and cannot satisfy #202's acceptance target. If
   runner reach matters, it gets measured once the venue exists.

test/run.sh 28/28; shellcheck 0.10.0, changelog-armed clean.

Refs #202
2026-08-05 13:46:03 +00:00

8 KiB

Runner probes

Not a drill. A drill rehearses the release doors on a disposable repo and ends. This is the opposite shape: one standing repo that exists so that runner-only facts can be measured on demand, and it is never archived.

heavy-duty/ceremony-runner-probe — private, standing, reset between probes. Ruled by the operator as option (A) of ceremony#202 (#5631).

Why a standing repo, when drills are disposable

Some facts are only true inside Actions, under the token Actions injects, and no local harness or PAT can reproduce them. The worked example is ceremony#192:

DELETE /issues/{n}/labels/{id}   ->  500   under ${{ github.token }} in a workflow
DELETE /issues/{n}/labels/{id}   ->  204   under a maintainer PAT, same call

A probe that runs anywhere else passes and proves nothing. Before this venue existed the answer was "un-archive a drill repo", which was requested three times in two days across two issues and never became anything — the three drill repos (ceremony-drill-0.4.1, -0.4.1-final, -191) are all archived, and each was minted for one probe and then wanted again.

The disposal rule above does NOT apply here

The rehearsal section says the builder archives the scratch repo and the operator deletes it. That rule is for drills. Archiving this repo defeats its entire purpose, and it is the failure mode the three archived drill repos demonstrate — each was archived correctly, by the rule, and each then had to be un-archived or replaced.

So: never archive it, never delete it, and if you find it archived, un-archive it rather than minting a fourth one.

Standing it up is the operator's step

Bot identities cannot create repositories in heavy-duty. Measured 2026-08-05 with a fleet identity holding the repo scope:

POST /api/v1/orgs/heavy-duty/repos   ->  403  "not allowed to create repository in organization"
POST /api/v1/user/repos              ->  201  (personal namespace only)

This is the same shape as the drill delete: a deliberate permission boundary, not a misconfiguration. Do not retry it, and do not work around it by putting the venue in a personal namespace — not because a personal namespace is proven unable to reach the org's runner (that was not measured; the probe repository above was deleted immediately, so nothing about runner or secret reach was established), but because @andres ruled an org-owned standing venue (#5631). A personally-owned repo is a different thing from the one that was decided on, and cannot satisfy #202's named acceptance target.

If runner or secret reach turns out to matter, measure it once the venue exists rather than assuming it here.

Running a probe

  1. Reset the repo to a clean state — the probe's own fixtures only, no leftovers from the last one. A probe that inherits state is a probe whose result you cannot attribute.
  2. Arm it against the candidate ref (below), if the probe is about ceremony's own code rather than about a bare API call.
  3. Run it as an Actions job under ${{ github.token }}. This is the whole point of the venue and the one step that cannot be shortcut. A curl from a laptop with a PAT answers a different question — see the 204/500 split above — and a probe run that way is worse than no probe, because it produces a confident wrong answer.
  4. The job writes its raw results into an issue in the PROBE repoheavy-duty/ceremony-runner-probe — not into ceremony. Logs age out; ceremony#192's run 701 survived only because the job wrote its findings into an issue it created.
  5. A human then records the probe issue's URL and the Actions run number on the ceremony issue the probe serves. That hop is deliberate and is the whole of the boundary: the probe workflow holds no credential and no code path that can write to heavy-duty/ceremony, so "the probe reports its findings" and "the probe cannot touch the live board" stay compatible rather than contradicting each other (@codex-reviewer-andresmgsl, #202 review).

Arming a candidate ref

A probe that exercises ceremony's own machinery needs the candidate tree reachable from a uses: line. The shape is the drill rehearsal's, reused rather than reinvented (drills/README.md step 2):

  1. The candidate is a commit SHA, on a fork ref. Push the candidate tree to a fork under the identity running the probe — <identity>/ceremony@probe-<issue> — and take its canonical SHA. Never create a branch on heavy-duty/ceremony named like a tag: it shadows that tag for every consumer until somebody remembers to delete it.

  2. Rewrite the COORDINATE, not only the ref. The candidate SHA exists only in your fork, so a stub still saying heavy-duty/ceremony/...@<sha> cannot resolve it — it would fail, or worse, resolve something else. In the probe repo, every caller uses: becomes <identity>/ceremony/<path>@<canonical-sha>.

  3. Rewrite the candidate's own self-checkout, both halves. labels.yml and the release doors hardcode repository: heavy-duty/ceremony beside ref: ${{ env.CEREMONY_SELF_REF }}. Changing only the ref makes the candidate fetch your SHA from the canonical repository, where it does not exist. So: every repository: carrier becomes <identity>/ceremony, and both CEREMONY_SELF_REF values become the canonical SHA.

    Then prove the rewrite was total, because a partial one silently tests canonical main instead of the candidate:

    grep -rn 'heavy-duty/ceremony' .github/ docs/CONSUMERS.md
    

    Every remaining hit must be prose. A uses:, a repository: or a pin among them means the probe is not armed, and its result is about the wrong tree.

  4. Invoke the probe by the event it is about, and record which: a workflow_dispatch of the caller, or the real board event the probe is testing. A probe that fires a different event than the one under test proves something else.

  5. Record the fork ref, the rewritten pin, the workflow invoked and the run number in the probe repo's result issue. Those four are what make the result reproducible; without the pin especially, a later reader cannot tell which tree answered.

  6. Reset removes the candidate-specific EXECUTABLE state: the caller stubs, the probe workflow, the candidate branch — so the next probe cannot inherit a pin it did not choose. Result issues are never deleted. They may be closed or relabelled; deleting them would recreate the expiring-log problem this venue exists to avoid.

Who may reset it

Operator-owned until ruled otherwise. #202's task 4 asks who may reset the venue, and creating the repo is the operator's step, so the access policy is his to set at the same time (@codex-reviewer-andresmgsl, #202 review).

Two levels, deliberately separated:

  • content reset — removing probe branches, workflows and fixtures; the ordinary between-probes operation. It does not include deleting result issues, which are the evidence and are immutable once written (@codex-reviewer-andresmgsl, #202 review);
  • archive / delete / admin — which is where the drill rule's damage came from, and which no bot identity should hold here.

If fleet identities are given push access for content reset, this section records that; until then, ask.

What must never happen here

No probe touches heavy-duty/ceremony's board. No labels, no comments, no runs attributable to a probe. The venue exists so that the live board does not have to be the test fixture.

The probes this venue owes

  • ceremony#192 — that the repaired sweep actually lifts a label under the workflow token, which is the half its acceptance criteria cannot get from the hermetic contract tests.
  • ceremony#205 — whether POST /actions/workflows/{file}/dispatches works on this instance with a valid ref and inputs. Measured so far: GET /actions/workflows 404s and the dispatch route answers 500 rather than a 4xx, which is not enough to port against.
  • A 0.6.0 consumer exercise once ceremony#198 has merged.