ceremony/docs/RUNNER-PROBES.md
cluade-reviewer-andresmgsl 20f4b287f7 docs(runner-probes): the generator survives a one-layer probe, callers carry full coordinates, one domain (#202)
@codex-reviewer-andresmgsl drove the published commands again and found three.

1. THE GENERATOR ABORTED ON AN ABSENT CALLER CLASS — the same `set -e` +
   `git grep` no-match bug I had just fixed in the CHECKER, in the generator I
   wrote in the same commit and did not apply the lesson to. A probe that
   exercises one layer produced no manifest and no diagnostic. `|| true` on
   every extraction, plus an explicit count so ZERO ceremony callers refuses by
   name while workflow-only and action-only probes generate valid manifests.

   That count check was itself broken on its first write: `grep -E '\t…'` reads
   a literal `t`, not a tab, so it counted zero on a perfectly good manifest
   and refused it. Found by running it.

2. CALLERS CARRY THE COMPLETE COORDINATE. The manifest stored only the sha and
   the checker compared owner and suffix separately, so
   `<fork>/actions/WRONG-ONE@<right-sha>` passed. The manifest now records
   `<fork>/<path>@<sha>` and every kind is one exact comparison — which also
   removes the per-kind branch that made the omission possible.

3. GENERATOR AND CHECKER SHARE ONE DOMAIN. `actual` extracted every `uses:`
   while the generator manifested only ceremony patterns, so a legitimate
   `actions/checkout` was always an unrecognised carrier. Both are restricted
   to ceremony callers; a wrong OWNER is still caught because
   `wrong-owner/ceremony/...` is still a ceremony caller.

And the stale fragment wording, which glm flagged and codex re-flagged:
"both CEREMONY_SELF_REF values" -> "every".

DRIVEN, all of it:

  generator: both / workflow-only / action-only  -> valid manifests
  generator: zero ceremony callers               -> refuses by name
  deletion, role swap x2, wrong owner, wrong sha,
  wrong path, deleted caller class, extra carrier -> all refuse
  armed control, third-party actions/checkout present -> passes

test/run.sh 28/28; shellcheck 0.10.0 and changelog-armed clean.

Refs #202
2026-08-05 15:07:11 +00:00

16 KiB

Runner probes

Not a drill. A drill rehearses the release doors on a disposable repo and ends. This is the opposite shape: one standing repo that exists so that runner-only facts can be measured on demand, and it is never archived.

heavy-duty/ceremony-runner-probe — private, standing, reset between probes. Ruled by the operator as option (A) of ceremony#202 (#5631).

Why a standing repo, when drills are disposable

Some facts are only true inside Actions, under the token Actions injects, and no local harness or PAT can reproduce them. The worked example is ceremony#192:

DELETE /issues/{n}/labels/{id}   ->  500   under ${{ github.token }} in a workflow
DELETE /issues/{n}/labels/{id}   ->  204   under a maintainer PAT, same call

A probe that runs anywhere else passes and proves nothing. Before this venue existed the answer was "un-archive a drill repo", which was requested three times in two days across two issues and never became anything — the three drill repos (ceremony-drill-0.4.1, -0.4.1-final, -191) are all archived, and each was minted for one probe and then wanted again.

The disposal rule above does NOT apply here

The rehearsal section says the builder archives the scratch repo and the operator deletes it. That rule is for drills. Archiving this repo defeats its entire purpose, and it is the failure mode the three archived drill repos demonstrate — each was archived correctly, by the rule, and each then had to be un-archived or replaced.

So: never archive it, never delete it, and if you find it archived, un-archive it rather than minting a fourth one.

Standing it up is the operator's step

Bot identities cannot create repositories in heavy-duty. Measured 2026-08-05 with a fleet identity holding the repo scope:

POST /api/v1/orgs/heavy-duty/repos   ->  403  "not allowed to create repository in organization"
POST /api/v1/user/repos              ->  201  (personal namespace only)

This is the same shape as the drill delete: a deliberate permission boundary, not a misconfiguration. Do not retry it, and do not work around it by putting the venue in a personal namespace — not because a personal namespace is proven unable to reach the org's runner (that was not measured; the probe repository above was deleted immediately, so nothing about runner or secret reach was established), but because @andres ruled an org-owned standing venue (#5631). A personally-owned repo is a different thing from the one that was decided on, and cannot satisfy #202's named acceptance target.

If runner or secret reach turns out to matter, measure it once the venue exists rather than assuming it here.

Running a probe

  1. Reset the repo to a clean state — the probe's own fixtures only, no leftovers from the last one. A probe that inherits state is a probe whose result you cannot attribute.
  2. Arm it against the candidate (below) — two layers, candidate code and armed workflow — if the probe is about ceremony's own machinery rather than about a bare API call.
  3. Run it as an Actions job under ${{ github.token }}. This is the whole point of the venue and the one step that cannot be shortcut. A curl from a laptop with a PAT answers a different question — see the 204/500 split above — and a probe run that way is worse than no probe, because it produces a confident wrong answer.
  4. The job writes its raw results into an issue in the PROBE repoheavy-duty/ceremony-runner-probe — not into ceremony. Logs age out; ceremony#192's run 701 survived only because the job wrote its findings into an issue it created.
  5. A human then records the probe issue's URL and the Actions run number on the ceremony issue the probe serves. That hop is deliberate and is the whole of the boundary: the probe workflow holds no credential and no code path that can write to heavy-duty/ceremony, so "the probe reports its findings" and "the probe cannot touch the live board" stay compatible rather than contradicting each other (@codex-reviewer-andresmgsl, #202 review).

Arming a candidate ref

A probe that exercises ceremony's own machinery needs the candidate tree reachable from a uses: line. This is the fork-ref shape drills/README.md step 2 points at, written out — and it has two layers, which is the part that is easy to get wrong and impossible to fix afterwards.

Why two. The candidate's own workflows contain repository: heavy-duty/ceremony beside ref: ${{ env.CEREMONY_SELF_REF }}, so they must be rewritten to point at the fork and at the candidate. But rewriting them creates a new commit, and a commit cannot contain its own object ID. A single-layer arming is therefore self-referential: pin the callers to the pre-rewrite SHA and they load the unarmed workflows; pin them to the post-rewrite one and you are asking a commit to embed itself (@codex-reviewer-andresmgsl, #202 review).

So:

layer what it is what it carries
candidate code SHA the immutable tree under test actions/, lib/ — untouched
armed workflow SHA a small child commit on top of it workflows rewritten to the fork + CEREMONY_SELF_REF = the candidate code SHA

The procedure

  1. Push the candidate tree to a fork under the identity that will run the probe — one branch, <identity>/ceremony@probe-<issue> — and record its SHA. Steps 1 and 2 advance the tip of that same branch; there are two commits, not two branches. That is the candidate code SHA. Never create a branch on heavy-duty/ceremony named like a tag: it shadows that tag for every consumer until somebody remembers to delete it.

  2. Commit the arming on top of it, and write the manifest. In that same fork branch rewrite, for every carrier the manifest below enumerates: ceremony's own internal repository: checkouts → <identity>/ceremony, and every CEREMONY_SELF_REF value → the candidate code SHA from step 1. There were three self-ref carriers on main at the time of writing and the count is not a constant — derive it, do not remember it (@glm-reviewer-andresmgsl, @codex-reviewer-andresmgsl, #202 review). The consumer checkouts (${{ github.repository }}) are left alone. Record the resulting SHA: that is the armed workflow SHA.

  3. Pin the probe repo's callers by layer, because they are not the same thing:

    • composite-action callers → <identity>/ceremony/actions/<name>@<candidate-code-sha>;
    • reusable-workflow callers → <identity>/ceremony/.github/workflows/<file>@<armed-workflow-sha>, since that is the only revision whose inner checkout is rewritten.
  4. Gate the arming against a MANIFEST, byte for byte. Every weaker shape has a hole, and each of these was found in a published draft of this file (@codex-reviewer-andresmgsl, #202 review):

    weaker check what slips through
    "the old literal is absent" a carrier rewritten to the wrong fork, or to the armed SHA
    "every extracted value equals X" a carrier that vanished — nothing to compare
    "each value is one of {fork, dynamic}" a role swap: an internal checkout made dynamic, a consumer checkout pointed at the fork
    "the SHA suffix matches" wrong-owner/ceremony/actions/foo@<right-sha>
    "known callers match" an unrecognised caller, or none at all

    So the arming step writes a manifest — one line per carrier, path, kind, full expected value — and the gate compares the tree's actual carriers against it as a set. A deletion, a role swap, a wrong fork, a wrong SHA, an extra carrier and a missing caller are then all the same kind of failure: the sets differ.

    Generate it while arming, from the tree you are arming, so the manifest cannot drift from the repository:

    #!/usr/bin/env bash
    # write-manifest <armed-checkout> <probe-checkout> <fork> <code-sha> <armed-sha>
    #
    # `|| true` on every extraction, for the same reason the checker needs it:
    # git grep exits 1 on no-match and `set -e` would abort BEFORE the manifest
    # is written — silently, which is how the first version of this generator
    # produced no file and no diagnostic when a probe exercised only one layer
    # (@codex-reviewer-andresmgsl). A probe need not use both.
    set -euo pipefail
    armed="$1"; probe="$2"; fork="$3"; code_sha="$4"; armed_sha="$5"
    {
      git -C "$armed" grep -n 'CEREMONY_SELF_REF:' -- .github/workflows \
        | cut -d: -f1,2 | sed "s|$|\tself_ref\t$code_sha|" || true
      git -C "$armed" grep -n 'repository: heavy-duty/ceremony' -- .github/workflows \
        | cut -d: -f1,2 | sed "s|$|\tinternal_repo\t$fork|" || true
      git -C "$armed" grep -n 'repository: ${{ github.repository }}' -- .github/workflows \
        | cut -d: -f1,2 | sed 's|$|\tconsumer_repo\t${{ github.repository }}|' || true
      # Callers record the COMPLETE expected coordinate, not just the sha: the
      # path is as rewritable as the owner, and a manifest that stores only the
      # suffix cannot notice `…/actions/wrong-one@<right-sha>`.
      git -C "$probe" grep -nE 'uses:[[:space:]]*[^[:space:]]*/ceremony/\.github/workflows/' -- .github \
        | sed -E "s|^([^:]+):([0-9]+):.*/ceremony/(\.github/workflows/[^@[:space:]]+)@.*|\\1:\\2\\tworkflow_caller\\t$fork/\\3@$armed_sha|" || true
      git -C "$probe" grep -nE 'uses:[[:space:]]*[^[:space:]]*/ceremony/actions/' -- .github \
        | sed -E "s|^([^:]+):([0-9]+):.*/ceremony/(actions/[^@[:space:]]+)@.*|\\1:\\2\\taction_caller\\t$fork/\\3@$code_sha|" || true
    } | sort >manifest.tsv
    
    # Zero ceremony callers is a refusal by name; one layer only is fine.
    callers="$(grep -cE '(workflow|action)_caller' manifest.tsv || true)"
    [ "$callers" -gt 0 ] || { echo "manifest: no ceremony callers found in $probe" >&2; exit 1; }
    

    Run it against the pre-arming tree — that is what enumerates the carriers that must change — then arm, then check:

    #!/usr/bin/env bash
    # check-arming <armed-checkout> <probe-checkout> <fork> <code-sha> <armed-sha> <manifest.tsv>
    set -euo pipefail
    armed="$1"; probe="$2"; fork="$3"; code_sha="$4"; armed_sha="$5"; manifest="$6"
    fail() { echo "arming incomplete: $*" >&2; exit 1; }
    
    # `|| true` on every extraction: git grep exits 1 when nothing matches, and
    # under `set -e` that would kill this script BEFORE the comparison — so a
    # carrier class that vanished ENTIRELY produced silence instead of a
    # refusal. Silence is the worst of the three outcomes; the comparison below
    # is what must report it.
    [ "$(grep -cE '(workflow|action)_caller' "$manifest" || true)" -gt 0 ] \
      || fail "manifest names no ceremony callers — it cannot prove an arming"
    
    actual="$(mktemp)"
    {
      git -C "$armed" grep -nP '(?<=CEREMONY_SELF_REF: ")[^"]+' -- .github/workflows \
        | sed -E 's/^([^:]+):([0-9]+):.*CEREMONY_SELF_REF: "([^"]*)".*/\1:\2\tself_ref\t\3/' || true
      git -C "$armed" grep -nE 'repository: .+' -- .github/workflows \
        | sed -E 's|^([^:]+):([0-9]+):[[:space:]]*repository:[[:space:]]*(.*)$|\1:\2\t__repo__\t\3|' || true
      # Only CEREMONY callers, matching the generator's domain exactly — a
      # third-party `actions/checkout` is not this gate's business, and
      # extracting it here while the generator ignores it made every probe fail
      # as an "unrecognised carrier" (@codex-reviewer-andresmgsl). A wrong OWNER
      # is still caught: `wrong-owner/ceremony/...` matches this pattern.
      git -C "$probe" grep -nE 'uses:[[:space:]]*[^[:space:]]*/ceremony/' -- .github \
        | sed -E 's|^([^:]+):([0-9]+):[[:space:]]*-?[[:space:]]*uses:[[:space:]]*(.*)$|\1:\2\t__uses__\t\3|' || true
    } | sort >"$actual"
    
    # every manifest line must be present with its EXACT expected value, and the
    # kinds must match — a role swap changes the kind, not just the value.
    while IFS=$'\t' read -r loc kind want; do
      case "$kind" in
        self_ref)        have="$(awk -F'\t' -v l="$loc" '$1==l && $2=="self_ref"{print $3}' "$actual")" ;;
        internal_repo|consumer_repo)
                         have="$(awk -F'\t' -v l="$loc" '$1==l && $2=="__repo__"{print $3}' "$actual")" ;;
        workflow_caller|action_caller)
                         have="$(awk -F'\t' -v l="$loc" '$1==l && $2=="__uses__"{print $3}' "$actual")" ;;
      esac
      [ -n "$have" ] || fail "carrier vanished: $loc ($kind)"
      # ONE comparison for every kind: the manifest already carries the complete
      # expected value, so owner, path AND sha are checked at once. Checking the
      # owner and the sha separately let `…/actions/wrong-one@<right-sha>`
      # through (@codex-reviewer-andresmgsl).
      [ "$have" = "$want" ] || fail "$loc ($kind): expected '$want', found '$have'"
    done <"$manifest"
    
    # and nothing UNRECOGNISED: every uses:/repository: in the trees must appear
    # in the manifest, so an added carrier is a failure rather than a silence.
    while IFS=$'\t' read -r loc _ _; do
      grep -qF "$loc"$'\t' "$manifest" || fail "carrier not in manifest: $loc"
    done <"$actual"
    

    Why a manifest rather than a longer list of assertions. The carrier set is a property of the tree at the moment of arming; any list written into this document is stale the next time a workflow is added. The manifest is generated from the tree, recorded in the result issue (step 6), and is the thing a later reader compares against — so "what was armed" is evidence rather than recollection.

  5. Invoke the probe by the event it is about, and record which: a workflow_dispatch, or the real board event under test. A probe that fires a different event than the one under test proves something else.

  6. The result issue records all of it: the fork repository, the candidate code SHA, the armed workflow SHA, every rewritten carrier, the workflow invoked and the run number. Those are what make the result reproducible; without the two SHAs distinguished, a later reader cannot tell which tree answered.

  7. Reset removes the candidate-specific EXECUTABLE state: the caller stubs, the probe workflow, and the fork's probe branch — whose tip carries both the candidate commit and the armed commit on top of it — so the next probe cannot inherit a pin it did not choose. Result issues are never deleted. They may be closed or relabelled; deleting them would recreate the expiring-log problem this venue exists to avoid.

Who may reset it

Operator-owned until ruled otherwise. #202's task 4 asks who may reset the venue, and creating the repo is the operator's step, so the access policy is his to set at the same time (@codex-reviewer-andresmgsl, #202 review).

Two levels, deliberately separated:

  • content reset — removing probe branches, workflows and fixtures; the ordinary between-probes operation. It does not include deleting result issues, which are the evidence and are immutable once written (@codex-reviewer-andresmgsl, #202 review);
  • archive / delete / admin — which is where the drill rule's damage came from, and which no bot identity should hold here.

If fleet identities are given push access for content reset, this section records that; until then, ask.

What must never happen here

No probe touches heavy-duty/ceremony's board. No labels, no comments, no runs attributable to a probe. The venue exists so that the live board does not have to be the test fixture.

The probes this venue owes

  • ceremony#192 — that the repaired sweep actually lifts a label under the workflow token, which is the half its acceptance criteria cannot get from the hermetic contract tests.
  • ceremony#205 — whether POST /actions/workflows/{file}/dispatches works on this instance with a valid ref and inputs. Measured so far: GET /actions/workflows 404s and the dispatch route answers 500 rather than a 4xx, which is not enough to port against.
  • A 0.6.0 consumer exercise once ceremony#198 has merged.