diff --git a/changelog.d/129.md b/changelog.d/129.md new file mode 100644 index 0000000..c6ff707 --- /dev/null +++ b/changelog.d/129.md @@ -0,0 +1,3 @@ +### Added + +- The release drill exercises `rig forgejo-runner` beside `rig runner`, so both shipped runner families carry evidence (#129) diff --git a/drill/README.md b/drill/README.md index efbbc62..fd54cf7 100644 --- a/drill/README.md +++ b/drill/README.md @@ -11,7 +11,7 @@ release (#105, and #107's debt). - **A throwaway Debian 13 machine** you can format, reached as root. The drill hardens its sshd, renames it, joins it to a tailnet, and installs - box/Incus, Coolify and a GitHub runner on it. It is not coming back. + box/Incus, Coolify and Actions runners on it. It is not coming back. The machine is its own reset — there is no teardown script and no need for one. - **The pinned candidate refs, both of them.** `--rig-ref` and @@ -34,6 +34,18 @@ release (#105, and #107's debt). and does something trivial (`echo drilled`). Tokens come from an authenticated `gh`, or from `RUNNER_TOKEN` / `RUNNER_REMOVE_TOKEN`. Without a fork the leg **skips, loudly, into the record**. +- **For leg 3's Forgejo half** (#129): `--forgejo-instance ` and + `--forgejo-runner-repo /`, where that repo carries the same + `workflow_dispatch` workflow — but with `runs-on: drill`, because a + Forgejo runner matches the bare label it registered with. Tokens come from + `FORGEJO_RUNNER_TOKEN` (a registration token) or `FORGEJO_API_TOKEN`, which + mints one and is also what dispatches the job. `--forgejo-ref` names the + branch to dispatch (default `main`): Forgejo's dispatch endpoint requires a + ref in the body, where GitHub's defaults to the repo's default branch. + Without an instance and a repo this half **skips, loudly and separately**. + Note there is no removal token — Forgejo has no deregistration endpoint, so + the leg removes locally and the record tells you to delete the stale runner + row by hand. - **For leg 4** (coolify): a version pin, `--coolify-version 4.1.2`. No pin, no leg — rig's own `coolify install` refuses to default a version and so does its drill. The skip is recorded. @@ -80,7 +92,11 @@ passes, failures and skips separately. skip, exit 0) survives into the record as a SKIP, never a pass. 3. **Runner lifecycle** — register against the fork, dispatch the drill workflow and watch the runner take it, deregister, and assert the - box's registration is actually gone. + box's registration is actually gone. Runs **once per forge**: `rig runner` + against GitHub, then `rig forgejo-runner` against a Forgejo instance + (#129). Both families ship, so a release that evidences only one + evidences half of what it ships; each half skips separately, so a record + can honestly show one forge drilled and the other not. 4. **Coolify** — installed at the pin, `AUTOUPDATE=false` landed in the effective `.env`, container running. diff --git a/drill/drill.sh b/drill/drill.sh index a6f161d..2e7412b 100644 --- a/drill/drill.sh +++ b/drill/drill.sh @@ -4,13 +4,16 @@ # ⚠ DESTRUCTIVE, AND MEANT TO BE. Run it on a THROWAWAY Debian machine you # can format. It wipes any installed rig and reinstalls from the pinned # ref, hardens sshd, sets the hostname, joins the tailnet, installs box -# and its Incus stack, installs Coolify and a GitHub Actions runner. +# and its Incus stack, installs Coolify and Actions runners (GitHub, and +# Forgejo when --forgejo-instance is given). # Never run it on a machine you care about. # # TS_AUTHKEY=tskey-... bash drill/drill.sh \ # --rig-ref release/0.4.0 --box-ref 0.9.0 \ # --users ./drill-users --run-id drill-2026-07-24-a \ -# --coolify-version 4.1.2 --runner-repo you/rig --yes +# --coolify-version 4.1.2 --runner-repo you/rig \ +# --forgejo-instance https://forgejo.example.com \ +# --forgejo-runner-repo you/drill-probe --yes # (--box-ref is a tag: since #103 the box that ships is the BOX_RELEASE tag.) # rig's drill asserts CONVERGENCE — a machine reaches its role, idempotently. # The legs (drills/README.md, issue #105): @@ -22,6 +25,10 @@ # the isolation boundary is box's drill's assertion, not this one's). # 2. db — the real dump/restore round-trip, test/db-integration.sh. # 3. runner lifecycle — register, take a job, deregister, against a fork. +# Runs once per forge: `rig runner` against GitHub (--runner-repo), and +# `rig forgejo-runner` against a Forgejo instance (--forgejo-instance + +# --forgejo-runner-repo). Both forges ship, so both need evidence; each +# skips loudly and separately when its inputs are absent (#129). # 4. coolify install — at a pinned version, AUTOUPDATE=false. # # Execution order is 1, 4, 2, 3 — coolify's installer is what puts Docker on @@ -67,6 +74,11 @@ RECORD="${DRILL_RECORD:-}" COOLIFY_VERSION="${DRILL_COOLIFY_VERSION:-}" RUNNER_REPO="${DRILL_RUNNER_REPO:-}" RUNNER_WORKFLOW="${DRILL_RUNNER_WORKFLOW:-drill.yml}" +FJ_INSTANCE="${DRILL_FORGEJO_INSTANCE:-}" +FJ_RUNNER_REPO="${DRILL_FORGEJO_RUNNER_REPO:-}" +# The branch the dispatch names. Forgejo's dispatch endpoint requires a ref in +# the body — unlike GitHub's, which defaults to the repo's default branch. +FJ_REF="${DRILL_FORGEJO_REF:-main}" YES=0 while [ $# -gt 0 ]; do @@ -83,7 +95,10 @@ while [ $# -gt 0 ]; do --coolify-version) COOLIFY_VERSION="$2"; shift 2 ;; --runner-repo) RUNNER_REPO="$2"; shift 2 ;; --runner-workflow) RUNNER_WORKFLOW="$2"; shift 2 ;; - -h|--help) sed -n '2,33p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;; + --forgejo-instance) FJ_INSTANCE="$2"; shift 2 ;; + --forgejo-runner-repo) FJ_RUNNER_REPO="$2"; shift 2 ;; + --forgejo-ref) FJ_REF="$2"; shift 2 ;; + -h|--help) sed -n '2,40p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;; *) echo "drill: unknown option: $1 (see --help)" >&2; exit 2 ;; esac done @@ -108,6 +123,50 @@ phase(){ printf '\n\033[1m══ %s\033[0m\n' "$*"; } LEG_NAMES=(); LEG_RESULTS=() leg() { LEG_NAMES+=("$1"); LEG_RESULTS+=("$2"); } +# forgejo_run_verdict — the verdict for OUR dispatch: +# success | failed | pending. Reads GET /repos/{o}/{r}/actions/tasks, whose +# shape is NOT GitHub's and was measured against forgejo.heavyduty.builders +# (8.0.3+gitea-1.22.0) on 2026-07-30 rather than read from the docs: +# +# * There is no `conclusion` field. `status` carries the terminal outcome +# directly ("success"), where GitHub splits status:completed + +# conclusion:success. Reading `conclusion` here gets an empty string on +# every run, which would grade a green job as failed. +# * `id` is a GLOBAL task id; the run's own URL ends in `run_number`. The +# pre-dispatch guard therefore compares `id`, exactly as the GitHub leg +# compares databaseId — an old run must never be read as this one. +# * The payload lists ASSIGNED tasks only. A run sitting queued is simply +# absent (measured: 200s of total_count 0 while the web UI showed the run +# as "job is not started"). So "no new id" is the ONLY signal that the +# runner never took the job — which is the verdict this leg exists for. +# * A task appears when ASSIGNED, so it can be seen mid-flight. A +# non-terminal status is pending, not failed; grading a running job as a +# failure would make the leg flaky inside its own watch window. +# +# grep-and-sed, not jq: a throwaway drill machine has neither jq nor an +# authenticated forge CLI, the same constraint json_field() carries in +# commands/lib/runner-config.sh. `id` is a bare number, which json_field's +# quoted-value shape cannot read, so this reads both forms itself. +forgejo_run_verdict() { + # `run_status`, not `status`: the latter is read-only in zsh, and while this + # file is bash, the function is awk-extracted and sourced by other shells in + # test harnesses — where the assignment fails silently and grades every run + # as pending. + local pre="$1" file="$2" id run_status + id="$(grep -o '"id"[[:space:]]*:[[:space:]]*[0-9][0-9]*' "$file" 2>/dev/null \ + | head -n1 | sed 's/.*:[[:space:]]*//')" + [ -n "$id" ] || { echo pending; return 0; } + [ "$id" != "$pre" ] || { echo pending; return 0; } + run_status="$(grep -o '"status"[[:space:]]*:[[:space:]]*"[^"]*"' "$file" 2>/dev/null \ + | head -n1 | sed 's/.*:[[:space:]]*"//; s/"$//')" + case "$run_status" in + success) echo success ;; + failure | cancelled | skipped | timedout) echo failed ;; + *) echo pending ;; + esac +} + + # run_logged — run a long command with its narration in a file # and a dot every 5s on the terminal: a silent multi-minute apt/install run is # indistinguishable from a wedge, and that ambiguity has cost box whole @@ -690,6 +749,123 @@ else fi fi +# ============================================================================= +phase "Leg 3 — forgejo runner lifecycle against an instance" +# ============================================================================= +# The same leg as above, for the other forge. Both forges ship a runner family +# (#109 added `rig forgejo-runner` beside `rig runner`), so a release that +# evidences only GitHub evidences half of what it ships (#129). +# +# Three things differ from the GitHub leg, all measured against +# forgejo.heavyduty.builders (8.0.3+gitea-1.22.0) on 2026-07-30, not read: +# +# * Scope is the TOKEN's, never a flag — `rig forgejo-runner install` refuses +# --repo on purpose (commands/forgejo-runner-install.sh:160). The repo here +# is only where the registration token is minted from, and where the +# dispatched workflow lives. +# * The mint path is /repos///actions/runners/registration-token. The +# instance's own swagger documents /repos///runners/registration-token +# — WITHOUT /actions/ — and that path 404s. Do not "fix" this to match the +# published API reference. +# * There is no deregistration endpoint, so there is no removal token and no +# remote deregistration: `rig forgejo-runner remove` is local-only by +# design (commands/forgejo-runner-remove.sh:7-11) and the runner row +# survives in the UI until a human deletes it. The record says so rather +# than implying a clean remote teardown the way the GitHub leg can. +# +# Tokens: FORGEJO_RUNNER_TOKEN (a registration token) is used directly; else +# FORGEJO_API_TOKEN mints one from FJ_RUNNER_REPO. Without an instance, a repo, +# or a token source the leg SKIPS loudly and the record says it did not run. +if [ -z "$FJ_INSTANCE" ] || [ -z "$FJ_RUNNER_REPO" ]; then + skip "forgejo runner lifecycle: no --forgejo-instance/--forgejo-runner-repo given — the leg did not run" + leg "forgejo runner lifecycle" "SKIPPED — no instance/repo provided" +else + fj_reg="${FORGEJO_RUNNER_TOKEN:-}" + if [ -z "$fj_reg" ] && [ -n "${FORGEJO_API_TOKEN:-}" ]; then + fj_reg="$(curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \ + "${FJ_INSTANCE%/}/api/v1/repos/${FJ_RUNNER_REPO}/actions/runners/registration-token" 2>/dev/null \ + | grep -o '"token"[[:space:]]*:[[:space:]]*"[^"]*"' | head -n1 \ + | sed 's/.*:[[:space:]]*"//; s/"$//')" + fi + if [ -z "$fj_reg" ]; then + skip "forgejo runner lifecycle: no FORGEJO_RUNNER_TOKEN and no FORGEJO_API_TOKEN to mint one — the leg did not run" + leg "forgejo runner lifecycle ($FJ_RUNNER_REPO)" "SKIPPED — no registration token source" + else + FJ_NAME="drill-$(hostname)-$$" + # The label MUST carry a docker:// image: forgejo-runner runs jobs in + # containers, and a bare label leaves runs-on matched but unrunnable. + if FORGEJO_RUNNER_TOKEN="$fj_reg" run_logged /tmp/drill-forgejo-runner-install.log \ + rig forgejo-runner install --instance "$FJ_INSTANCE" --name "$FJ_NAME" \ + --labels 'drill:docker://node:22-bookworm'; then + ok "rig forgejo-runner install --instance $FJ_INSTANCE exited 0 (registered as $FJ_NAME)" + else + no "forgejo-runner install FAILED — tail: $(tail -3 /tmp/drill-forgejo-runner-install.log | tr '\n' ' ')" + fi + rig forgejo-runner status 2>/dev/null | grep -qF "${FJ_INSTANCE%/}" \ + && ok "forgejo-runner status names the instance: $FJ_INSTANCE" \ + || no "forgejo-runner status does not name ${FJ_INSTANCE}" + + fj_took=none + if [ -n "${FORGEJO_API_TOKEN:-}" ]; then + fj_api="${FJ_INSTANCE%/}/api/v1/repos/${FJ_RUNNER_REPO}" + # Read the newest ASSIGNED task id BEFORE dispatching, same guard as the + # GitHub leg: an already-completed run must never be read as ours. + fj_pre_body="$(curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \ + "$fj_api/actions/tasks" 2>/dev/null || echo '{}')" + printf '%s' "$fj_pre_body" > /tmp/drill-forgejo-pre.json + fj_pre="$(grep -o '"id"[[:space:]]*:[[:space:]]*[0-9][0-9]*' /tmp/drill-forgejo-pre.json \ + 2>/dev/null | head -n1 | sed 's/.*:[[:space:]]*//')" + if curl -fsSL -o /dev/null -X POST -H "Authorization: token ${FORGEJO_API_TOKEN}" \ + -H "Content-Type: application/json" -d "{\"ref\":\"${FJ_REF}\"}" \ + "$fj_api/actions/workflows/${RUNNER_WORKFLOW}/dispatches" 2>/dev/null; then + inf "dispatched $RUNNER_WORKFLOW on $FJ_RUNNER_REPO — waiting for the runner to take it (≤5 min)…" + fj_took=timeout + for _i in $(seq 1 30); do + sleep 10 + curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \ + "$fj_api/actions/tasks" -o /tmp/drill-forgejo-tasks.json 2>/dev/null || continue + case "$(forgejo_run_verdict "${fj_pre:-}" /tmp/drill-forgejo-tasks.json)" in + success) fj_took=success; break ;; + failed) fj_took=failed; break ;; + *) : ;; # pending — queued, or assigned and still running + esac + done + else + fj_took=nodispatch + fi + case "$fj_took" in + success) ok "the forgejo runner took a job and it succeeded ($RUNNER_WORKFLOW)" ;; + failed) no "the dispatched job completed UNSUCCESSFULLY — the runner ran it, the workflow failed; read the run on $FJ_RUNNER_REPO" ;; + # A queued task is INVISIBLE in this API until a runner claims it, so a + # timeout means "nothing ever took it". Two causes, and the second one + # is not rig's: the runs-on label may not match, or the daemon's poller + # can go quiet — a restarted daemon claims a minutes-old backlog in + # about a second. Check 'systemctl restart forgejo-runner' before + # reading this as a rig defect. + timeout) no "the dispatched job was never taken within 5 min — check the workflow's runs-on is 'drill', then restart forgejo-runner and re-read (a quiet poller looks exactly like this)" ;; + nodispatch) no "could not dispatch $RUNNER_WORKFLOW on $FJ_RUNNER_REPO — does it carry that workflow, with workflow_dispatch, on its default branch?" ;; + esac + else + skip "took a job: not attempted — no FORGEJO_API_TOKEN to dispatch $RUNNER_WORKFLOW with" + fi + + # No removal token exists on this forge — remove is local by design. + rig forgejo-runner remove >/dev/null 2>&1 \ + && note "forgejo-runner removed locally — Forgejo has no deregistration endpoint, so DELETE the stale '$FJ_NAME' row under $FJ_RUNNER_REPO > Settings > Actions > Runners by hand" \ + || no "forgejo-runner remove FAILED" + rig forgejo-runner status >/dev/null 2>&1 \ + && no "forgejo-runner status still answers after remove — the removal did not take" \ + || ok "forgejo-runner status confirms: nothing registered" + + leg "forgejo runner lifecycle ($FJ_RUNNER_REPO)" \ + "$(case "$fj_took" in + success) echo "PASS — registered, took a job, removed (stale row needs deleting by hand)" ;; + none) echo "PARTIAL — registered and removed; took a job: not attempted (no FORGEJO_API_TOKEN)" ;; + *) echo "FAIL — see Failed below" ;; + esac)" + fi +fi + # ============================================================================= phase "Summary" # ============================================================================= diff --git a/drills/README.md b/drills/README.md index 5b2927a..581e0db 100644 --- a/drills/README.md +++ b/drills/README.md @@ -113,6 +113,7 @@ Candidate refs: box@1a2b3c4 (BOX_REF=release/0.4.0), rig@5d6e7f8, cast@9a0b1c2. | --host yes: pinned box installed, host stack up | PASS — box doctor clean | | `test/db-integration.sh` | PASS — 14 passed, 0 failed | | runner lifecycle against a fork | PASS — registered, took a job, deregistered clean | +| forgejo runner lifecycle (you/drill-probe) | PASS — registered, took a job, removed (stale row needs deleting by hand) | | coolify install (4.1.2) | PASS (6 min) | Failed: `rig users apply` left one revoked key in `authorized_keys` diff --git a/test/drill.sh b/test/drill.sh index 8c01a08..04b8f38 100644 --- a/test/drill.sh +++ b/test/drill.sh @@ -50,10 +50,10 @@ trap 'rm -rf "$WORK"' EXIT # --- the functions under test, extracted ------------------------------------- FNS="$WORK/drill-fns.sh" -for fn in tree_of assert_installed_from classify_leg capture_state emit_record; do +for fn in tree_of assert_installed_from classify_leg capture_state emit_record forgejo_run_verdict; do awk "/^${fn}\(\) \{/,/^\}/" "$ROOT/drill/drill.sh" >> "$FNS" done -for fn in tree_of assert_installed_from classify_leg capture_state emit_record; do +for fn in tree_of assert_installed_from classify_leg capture_state emit_record forgejo_run_verdict; do check "extraction guards the awk: ${fn}() landed" 0 "${fn}() {" grep -F "${fn}() {" "$FNS" done # shellcheck source=/dev/null @@ -201,6 +201,54 @@ check "an all-green record says every leg ran and passed" 0 "Every leg ran and e # ============================================================================= # the shipped script itself # ============================================================================= +# ============================================================================= +# forgejo_run_verdict — did OUR dispatched run land, and how (#129) +# ============================================================================= +# The Forgejo half of the runner leg cannot reuse the GitHub reader. Measured +# against forgejo.heavyduty.builders (8.0.3+gitea-1.22.0) on 2026-07-30, a +# completed run in GET /repos/{o}/{r}/actions/tasks carries NO `conclusion` +# field at all — `status` holds the terminal outcome directly, where GitHub +# splits status:completed + conclusion:success. And `id` is a global task id +# (25) while the run's own URL ends in run_number (1), so the pre-dispatch +# guard has to compare `id`. +# +# The payload is also an ASSIGNED-task view: it reads total_count 0 for as +# long as a run sits queued (measured: 200s), so "no new id" is the ONLY +# signal that the runner never took the job. That is the verdict this leg +# exists to produce, which is why it gets its own function and its own tests. +FJ="$WORK/fj"; mkdir -p "$FJ" +printf '%s' '{"workflow_runs":[],"total_count":0}' > "$FJ/empty.json" +printf '%s' '{"workflow_runs":[{"id":25,"status":"success","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-success.json" +printf '%s' '{"workflow_runs":[{"id":25,"status":"failure","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-failure.json" +printf '%s' '{"workflow_runs":[{"id":25,"status":"cancelled","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-cancelled.json" +printf '%s' '{"workflow_runs":[{"id":24,"status":"success","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/stale-only.json" + +check "verdict: an empty task list is PENDING, never a pass" 0 "pending" \ + forgejo_run_verdict "" "$FJ/empty.json" +check "verdict: a queued run the runner never took stays PENDING" 0 "pending" \ + forgejo_run_verdict "24" "$FJ/stale-only.json" +check "verdict: OUR new run, status success, is SUCCESS" 0 "success" \ + forgejo_run_verdict "24" "$FJ/new-success.json" +check "verdict: status carries the outcome — failure is FAILED, not success" 0 "failed" \ + forgejo_run_verdict "24" "$FJ/new-failure.json" +check "verdict: a cancelled run is FAILED, not silently passed" 0 "failed" \ + forgejo_run_verdict "24" "$FJ/new-cancelled.json" +check "verdict: the first run ever (no pre-id) still resolves" 0 "success" \ + forgejo_run_verdict "" "$FJ/new-success.json" +# A task appears in this payload the moment it is ASSIGNED, which can be before +# it finishes — so a non-terminal status must read as pending, not as a failure. +# Calling a still-running job "failed" would make the leg flaky in exactly the +# window the leg is watching. +printf '%s' '{"workflow_runs":[{"id":25,"status":"running","run_number":1}],"total_count":1}' > "$FJ/new-running.json" +check "verdict: an assigned-but-running task is PENDING, not FAILED" 0 "pending" \ + forgejo_run_verdict "24" "$FJ/new-running.json" + +# The anti-false-positive guard, stated as its own case: an OLD completed run +# with the SAME id as pre_id must never be read as this dispatch's result. +check "verdict: a pre-existing success with the pre-id is NOT our run" 0 "pending" \ + forgejo_run_verdict "25" "$FJ/new-success.json" + + # Arg refusals fire before the root check (repo doctrine, bootstrap.sh:114), # which is what makes them provable here without a throwaway machine. check "drill.sh refuses to run without BOTH refs pinned (#103)" 2 "--box-ref" \ @@ -217,6 +265,14 @@ check "an unknown flag dies loudly, exit 2" 2 "unknown option" \ bash "$ROOT/drill/drill.sh" --frobnicate check "--help prints the header and exits 0" 0 "THROWAWAY" \ bash "$ROOT/drill/drill.sh" --help +check "--forgejo-instance is a known flag (the leg's opt-in)" 2 "--users is required" \ + bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-instance https://f.example.com --yes +check "--forgejo-runner-repo is a known flag" 2 "--users is required" \ + bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-runner-repo o/r --yes +check "--help names the forgejo runner leg's flags" 0 "--forgejo-instance" \ + bash "$ROOT/drill/drill.sh" --help +check "--forgejo-ref is a known flag (Forgejo's dispatch needs a ref)" 2 "--users is required" \ + bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-ref dev --yes echo "---" echo "$PASS passed, $FAIL failed"