Compare commits

..

No commits in common. "main" and "main" have entirely different histories.
main ... main

23 changed files with 33 additions and 1041 deletions

4
.github/labels.conf vendored
View file

@ -1,5 +1,5 @@
panel=cluade-reviewer-andresmgsl codex-reviewer-andresmgsl grok-reviewer-andresmgsl kimi-reviewer-andresmgsl
triage-actors=dan-claude-bot cluade-reviewer-andresmgsl
panel=claude-bot-andresmgsl codex-bot-andresmgsl grok-bot-andresmgsl kimi-bot-andresmgsl
triage-actors=dan-claude-bot
scope:bootstrap|C5DEF5|bootstrap — hardening a pristine server into a node
scope:users|C5DEF5|users-* — class model, apply/status, close-root
scope:runner|C5DEF5|runner-* / forgejo-runner-* — CI runner lifecycle, either forge

View file

@ -26,21 +26,7 @@ jobs:
# The file list is printed so under-coverage shows up in the log, and
# the comm below turns under-coverage into a failure rather than a
# thing someone has to notice: every tracked `.sh` must be in the set.
#
# Install shellcheck when missing (#144). rig's default Forgejo label
# maps ubuntu-latest to catthehacker's act-22.04 (slim), which does not
# ship shellcheck; GitHub-hosted ubuntu-latest does. The conditional
# keeps each forge from paying for the other.
#
# sudo: load-bearing on GitHub (job runs as `runner` with passwordless
# sudo) and a no-op on act-22.04 (jobs run as uid 0; the image has no
# `runner` account). Do not delete it as "dead weight" — that breaks
# the GitHub half the day that image stops preinstalling shellcheck.
run: |
if ! command -v shellcheck >/dev/null 2>&1; then
sudo apt-get update
sudo apt-get install -y shellcheck
fi
shopt -s globstar dotglob
files=(bin/* **/*.sh)
printf 'shellcheck: %s\n' "${files[@]}"

View file

@ -16,8 +16,8 @@ genuinely rig's.
1. **Fork and branch.** Contributors work from forks; upstream branches are
for maintainers. Title the PR conventionally (`feat:`, `fix:`, `docs:`).
2. **The review panel** (`.github/labels.conf`'s `panel=` line):
`cluade-reviewer-andresmgsl`, `codex-reviewer-andresmgsl`,
`grok-reviewer-andresmgsl`, `kimi-reviewer-andresmgsl` —
`claude-bot-andresmgsl`, `codex-bot-andresmgsl`, `grok-bot-andresmgsl`,
`kimi-bot-andresmgsl` —
the required verdicts for a PR are the panel minus its author. The
maintainer (`danmt`) takes the last word and merges.
3. **Checks must be green**: `shellcheck`, `bash test/cli.sh` and

View file

@ -19,11 +19,6 @@ takes arguments, does its work, and stores no credential, ever.
curl -fsSL https://raw.githubusercontent.com/heavy-duty/rig/main/install.sh | RIG_REF=main bash
# the latest release:
curl -fsSL https://raw.githubusercontent.com/heavy-duty/rig/main/install.sh | bash
# the same two channels from the Forgejo mirror (RIG_HOST picks the forge; #111):
curl -fsSL https://forgejo.heavyduty.builders/heavy-duty/rig/raw/branch/main/install.sh \
| RIG_HOST=https://forgejo.heavyduty.builders RIG_REF=main bash
curl -fsSL https://forgejo.heavyduty.builders/heavy-duty/rig/raw/branch/main/install.sh \
| RIG_HOST=https://forgejo.heavyduty.builders bash
```
This README tracks `main`, so the quick start installs that same development
@ -42,11 +37,6 @@ curl -fsSL .../install.sh | RIG_REF=main bash # the development tree
A tag outranks a branch of the same name (the pin must win); anything that
is not a tag falls back to `refs/heads/<ref>`.
`RIG_HOST` picks the forge the channel reads — `https://github.com` by
default (#111) — and it is chosen independently of the URL the script itself
came from, so installing *from* the Forgejo mirror means naming the mirror
twice, as the quick start's third line does.
The layout, under the install root (`~/.local/share/rig`):
```

View file

@ -1,3 +0,0 @@
### Fixed
- Ceremony recognizes the Forgejo review panel and both forges' triage actors (#116)

View file

@ -1,3 +0,0 @@
### Added
- The release drill exercises `rig forgejo-runner` beside `rig runner`, so both shipped runner families carry evidence (#129)

View file

@ -1,3 +0,0 @@
### Fixed
- The README's install quick start documents the Forgejo channel — `RIG_HOST` shipped in #111 but was documented nowhere (#131)

View file

@ -1,3 +0,0 @@
### Fixed
- `rig forgejo-runner status` no longer lets a service's `(active)` stand as proof the runner is fetching jobs (#133)

View file

@ -1,3 +0,0 @@
### Fixed
- Forgejo runners installed by rig can start their cache server — `$HOME/.cache` is created and punched through `ProtectHome` (#135)

View file

@ -1,4 +0,0 @@
### Fixed
- `rig forgejo-runner install`, `rig runner install`, `rig users apply` and `rig bootstrap <tenant>` refuse with a named remedy when root's `PATH` carries no `/usr/sbin`, instead of dying on `useradd: command not found` after prompting for a token (#139)
- `rig users apply` no longer reports success having silently skipped the sudoers drop-in when `visudo` is off `PATH` (#139)

View file

@ -1,8 +0,0 @@
### Added
- Default Forgejo runner labels include opt-in `ubuntu-latest-full` for the GitHub-parity image (#144)
### Fixed
- `ci.yml` installs `shellcheck` when the runner image lacks it, so Forgejo's slim `ubuntu-latest` can run `check` (#144)
- Plain `rig forgejo-runner install` warns when recorded labels are a retired rig default, without nagging custom `--labels` (#144)

View file

@ -31,8 +31,6 @@ HERE="$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)"
. "$HERE/lib/templates.sh" # templates_resolve / template_parse_env / render_tenant_context
# shellcheck source=SCRIPTDIR/lib/users-config.sh
. "$HERE/lib/users-config.sh" # read_role_marker / root_door_of
# shellcheck source=SCRIPTDIR/lib/admin-path.sh
. "$HERE/lib/admin-path.sh" # require_admin_bins
# shellcheck source=SCRIPTDIR/lib/sshd.sh
. "$HERE/lib/sshd.sh" # harden_sshd (the staging-box tenant)
# shellcheck source=SCRIPTDIR/lib/manifest.sh
@ -202,15 +200,6 @@ else
fi
[ "$(id -u)" -eq 0 ] || die "must run as root"
# Root is not enough here either (#139). This mint adds the tenant user to the
# docker group with `usermod` far below — AFTER installing docker and node,
# which is what makes it the worst-placed of the four call sites: on a
# PATH-shorn root it dies mid-convergence with a bare `usermod: command not
# found`, having already changed the machine, rather than before touching it.
#
# Unconditional because the docker block below is unconditional — "every tenant
# gets docker" is the stated rule there, so every tenant reaches the usermod.
require_admin_bins usermod
if [ -r /etc/os-release ]; then
# Sourced in a subshell: os-release defines VERSION, NAME, ID, etc. —
# sourcing it in the main shell silently clobbers same-named script vars.

View file

@ -17,8 +17,6 @@ set -euo pipefail
HERE="$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)"
# shellcheck source=SCRIPTDIR/lib/forgejo-runner-config.sh
. "$HERE/lib/forgejo-runner-config.sh"
# shellcheck source=SCRIPTDIR/lib/admin-path.sh
. "$HERE/lib/admin-path.sh"
log() { printf 'rig-forgejo-runner: %s\n' "$*"; }
warn() { printf 'rig-forgejo-runner: WARNING: %s\n' "$*" >&2; }
@ -33,32 +31,7 @@ die() { printf 'rig-forgejo-runner: ERROR: %s\n' "$1" >&2; exit "${2:-1}"; }
# on the box itself. No docker-in-docker: the guide this came from stacks a
# privileged dind sidecar with a plaintext tcp://…:2375 daemon to isolate jobs
# from a shared CI server, and inside a box that boundary is already paid for.
#
# WHY act-22.04 (slim) for ubuntu-latest, not full-22.04 — measured 2026-08-01
# against ghcr manifests (#144):
# act-22.04: ~0.55 GB compressed / ~2.2 GB on disk — no shellcheck
# full-22.04: ~18.67 GB compressed / ~54.5 GB on disk — has shellcheck 0.8.0
# A normal box-class ci tenant cannot hold 54.5 GB (typical free space ~34 GB).
# So ubuntu-latest stays slim, and workflows must not assume GitHub-image tools
# (rig's own ci.yml installs shellcheck when missing). Operators who need the
# full tool surface opt in with runs-on: ubuntu-latest-full — that label is
# inert until matched, so boxes that never ask pay nothing.
DEFAULT_LABELS='ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04,ubuntu-latest-full:docker://ghcr.io/catthehacker/ubuntu:full-22.04,docker:docker://node:22-bookworm'
# labels_are_a_retired_default <recorded>
#
# True when <recorded> is a past DEFAULT_LABELS value rig has shipped — the
# only plain-converge case that should warn about re-registration (#144).
# Custom operator maps (drill's Leg 3, any --labels) must return false so a
# bare re-run stays quiet. One pattern per past default; add when the string
# changes. Extracted and driven by test/cli.sh — a grep pin alone cannot prove
# the match is exact.
labels_are_a_retired_default() {
case "${1:-}" in
'ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04,docker:docker://node:22-bookworm') return 0 ;;
*) return 1 ;;
esac
}
DEFAULT_LABELS='ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04,docker:docker://node:22-bookworm'
# fetch_and_verify_sha256 <asset-url> <file> <sumfile> <label>
#
@ -120,10 +93,8 @@ usage: rig forgejo-runner install --instance <url> [options]
time). Pin it for a deterministic, auditable install.
--name <name> runner name (default: this host's hostname)
--labels <csv> runner labels; replaces the default. The default maps
ubuntu-latest (slim act image), ubuntu-latest-full
(opt-in parity image), and docker onto containers so
a workflow written for GitHub runs; full tools need
runs-on: ubuntu-latest-full or an install step.
ubuntu-latest and docker onto container images, so a
workflow written for GitHub runs unchanged.
--user <name> unprivileged service user (default: the tenant user
`ci` when it exists, else forgejo-runner; created if
absent; never root)
@ -231,17 +202,6 @@ fi
# --- guards ----------------------------------------------------------------
[ "$(id -u)" -eq 0 ] || die "must run as root"
# Root is not enough: the admin binaries must also be reachable (#139). This
# sits beside the root check so identity and capability are asserted together,
# and BEFORE any prompt or download — a token typed for a doomed run is waste.
#
# BOTH binaries this command goes on to call: useradd at the user-create below,
# usermod at the docker-group add further down. Naming only the first would
# still consume the token on a PATH that happened to resolve useradd but not
# usermod — the same failure one step later, which is the shape #75 exists to
# refuse (a sweep that covers most of its call sites is the hole the next bug
# arrives through).
require_admin_bins useradd usermod
if [ -r /etc/os-release ]; then
# Sourced in a subshell: os-release defines VERSION (e.g. "13 (trixie)"),
# which would clobber this script's $VERSION.
@ -437,40 +397,17 @@ fi
# holds the UUID. assert_runner_instance's contract survives either way — it
# asks about the instance, which both spellings record.
install -d -m 0755 -o "$RUNNER_USER" -g "$RUNNER_GROUP" "$RUNNER_DIR"
# forgejo-runner's cache server writes to $HOME/.cache, which ProtectHome makes
# read-only below. Create it HERE, before the unit can reference it: a
# ReadWritePaths entry naming a path that does not exist makes systemd refuse
# to start the unit at all ("Failed to set up mount namespacing"), which is
# worse than the disabled cache it was meant to fix. Measured on a live runner,
# 2026-07-31 (#135). Root-owned would fail the same way under User=, so it
# carries the runner's own ownership like RUNNER_DIR above.
install -d -m 0755 -o "$RUNNER_USER" -g "$RUNNER_GROUP" "$USER_HOME/.cache"
if [ -e "$RUNNER_DIR/.runner" ]; then
log "already registered; skipping registration"
# Registration was skipped, so the labels on the instance are the ones it was
# registered with — NOT whatever this invocation was passed. Forgejo owns
# labels from registration time; a re-run never rewrites them.
#
# Two warn paths (#144):
# EXPLICIT --labels that differs → operator asked and it was not applied.
# Plain converge whose recorded labels match a known *retired* default →
# rig's default map moved (e.g. added ubuntu-latest-full). Without this
# the operator re-runs install, sees "already registered", and believes
# they have the new default while Forgejo still holds the old set.
#
# Do NOT warn on every RECORDED != current default: that fires forever for
# any runner the operator deliberately gave custom --labels (drill's Leg 3
# registers drill:docker://node:22-bookworm). The old LABELS_EXPLICIT-only
# gate existed to avoid that noise; retired-default matching keeps the
# silence for intentional maps and still catches silent drift off a past
# rig default. Re-register only to pick up labels the old set never had —
# nothing matching the recorded set is broken by the map change alone.
if [ -r "$RUNNER_DIR/.rig-labels" ]; then
# registered with — NOT whatever this invocation was passed. Say so when the
# operator explicitly asked for different ones, rather than letting the
# request evaporate. Only when EXPLICIT: comparing the default against a
# runner registered with custom labels would warn on every plain converge.
if [ "$LABELS_EXPLICIT" -eq 1 ] && [ -r "$RUNNER_DIR/.rig-labels" ]; then
RECORDED="$(cat "$RUNNER_DIR/.rig-labels")"
if [ "$LABELS_EXPLICIT" -eq 1 ] && [ "$RECORDED" != "$LABELS" ]; then
if [ "$RECORDED" != "$LABELS" ]; then
warn "--labels was not applied: this runner is already registered, and Forgejo owns its labels from registration time. It still has: ${RECORDED}. Labels are what 'runs-on' matches, so changing them means re-registering: 'rig forgejo-runner remove' then install again with the labels you want."
elif [ "$LABELS_EXPLICIT" -eq 0 ] && labels_are_a_retired_default "$RECORDED"; then
warn "this runner was registered with an older rig default label set. The current default adds ubuntu-latest-full (the GitHub-parity image). Labels are fixed at registration, so picking it up means re-registering: 'rig forgejo-runner remove' then install again. Nothing you run today is affected — re-register only if you want the new label."
fi
fi
else
@ -524,7 +461,7 @@ NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=full
ProtectHome=read-only
ReadWritePaths=${RUNNER_DIR} ${USER_HOME}/.cache
ReadWritePaths=${RUNNER_DIR}
[Install]
WantedBy=multi-user.target

View file

@ -11,31 +11,6 @@ log() { printf 'rig-forgejo-runner: %s\n' "$*"; }
warn() { printf 'rig-forgejo-runner: WARNING: %s\n' "$*" >&2; }
die() { printf 'rig-forgejo-runner: ERROR: %s\n' "$1" >&2; exit "${2:-1}"; }
# forgejo_runner_liveness_note <systemctl-state> — what `active` does not cover.
#
# `active` is the strongest health signal this command has, and it proves only
# that a process exists — not that the runner is still asking Forgejo for work.
# A poller can go quiet while the daemon stays up: measured 2026-07-30 (#129,
# #133), a daemon logged "[poller] launched" and never fetched a job dispatched
# four minutes later, while a daemon started fresh claimed that same queued task
# in one second. Both times it read as a label-mapping bug on the forge, which
# is the wrong place to look.
#
# A FUNCTION rather than an inline `if`, because the state boundary is the part
# worth pinning: an absent or inactive unit must say nothing, and a grep over
# the source cannot tell the difference (codex/kimi, !134).
#
# log, not warn: nothing has been DETECTED here. An idle runner with no queued
# jobs is silent in exactly the same way a stalled one is, so there is no signal
# separating them — a warning on every status run would be crying wolf, and warn
# in this file means a drift actually measured (the .runner mode below).
forgejo_runner_liveness_note() {
[ "${1:-}" = active ] || return 0
log " note: 'active' is not proof the runner is fetching jobs — only that the process is up."
log " If a job stays queued and its run page says it never started, run"
log " 'systemctl restart forgejo-runner' and re-read before suspecting the labels."
}
usage() {
cat <<'EOF'
usage: rig forgejo-runner status [--user <name>]
@ -105,8 +80,6 @@ log "name: ${RUNNER_NAME:-unknown}"
log "labels: ${LABELS}"
log "dir: ${RUNNER_DIR}"
log "service: ${SERVICE}"
forgejo_runner_liveness_note "${STATE:-}"
# status is the only command an operator runs when nothing is obviously wrong,
# which makes it the right place to notice a mode that drifted. It reports and

View file

@ -1,47 +0,0 @@
#!/usr/bin/env bash
# admin-path.sh — assert the admin binaries are REACHABLE, not merely that we
# are root.
#
# Being uid 0 and being able to find useradd are different facts, and rig
# asserted only the first. `su` without `-`, sudo with a sanitised secure_path,
# and several container images all hand you a root shell whose PATH carries no
# /usr/sbin — which is where useradd, usermod and groupadd live on Debian. The
# result was a bare `useradd: command not found` naming a line number inside a
# versioned install root, emitted AFTER a registration token had been read off
# the operator's terminal (#139).
#
# Which binaries this covers is measured, not assumed (Debian 13, 2026-08-01):
#
# useradd usermod groupadd userdel groupdel /usr/sbin package: passwd
# visudo /usr/sbin package: sudo
# gpasswd /usr/bin package: passwd
#
# Two consequences worth keeping written down. `gpasswd` is in the same PACKAGE
# as useradd but a different DIRECTORY, so it is reachable on a PATH-shorn root
# and does not belong in any of these preflights — do not add it for symmetry.
# And `visudo` shares the directory but not the package, so its absence has a
# second, innocent cause (sudo simply not installed) that the others do not, so
# `rig users apply` checks it separately, after the point where that cause is
# ruled out. See the comment there.
#
# (Spelled without the `.sh` on purpose: test/cli.sh pins that exactly one file
# under commands/ names that script, to catch a second caller appearing. A
# comment is not a caller, but the pin is deliberately blunt and cheap.)
#
# It REFUSES rather than repairing PATH itself. A command that quietly prepends
# /usr/sbin teaches the operator nothing and leaves a misconfigured host
# misconfigured; the same reason bootstrap refuses rather than guessing. The
# message carries the fix so the refusal costs one paste, not an investigation.
# require_admin_bins <bin>... — die unless every one resolves on PATH.
require_admin_bins() {
local missing=() b
for b in "$@"; do
command -v "$b" >/dev/null 2>&1 || missing+=("$b")
done
[ "${#missing[@]}" -eq 0 ] && return 0
# Names the REMEDY, not this script: the operator typed a `rig ...` command,
# and echoing the internal path back at them is the unhelpful half of the
# original `useradd: command not found`.
die "cannot find ${missing[*]} on PATH — it lives in /usr/sbin, which this root shell does not carry (a 'su' without '-' does this, and so do some container images). Re-run the same rig command with: PATH=/usr/sbin:/sbin:\$PATH"
}

View file

@ -9,8 +9,6 @@ set -euo pipefail
HERE="$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)"
# shellcheck source=SCRIPTDIR/lib/runner-config.sh
. "$HERE/lib/runner-config.sh"
# shellcheck source=SCRIPTDIR/lib/admin-path.sh
. "$HERE/lib/admin-path.sh"
log() { printf 'rig-runner: %s\n' "$*"; }
warn() { printf 'rig-runner: WARNING: %s\n' "$*" >&2; }
@ -87,10 +85,6 @@ VERSION="${VERSION#v}"
# --- guards ----------------------------------------------------------------
[ "$(id -u)" -eq 0 ] || die "must run as root"
# Root is not enough: the admin binaries must also be reachable (#139). This
# sits beside the root check so identity and capability are asserted together,
# and BEFORE any prompt or download — a token typed for a doomed run is waste.
require_admin_bins useradd
if [ -r /etc/os-release ]; then
# Sourced in a subshell: os-release defines VERSION (e.g. "13 (trixie)"),
# which would clobber this script's $VERSION.

View file

@ -11,8 +11,6 @@ set -euo pipefail
HERE="$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)"
# shellcheck source=SCRIPTDIR/lib/users-config.sh
. "$HERE/lib/users-config.sh"
# shellcheck source=SCRIPTDIR/lib/admin-path.sh
. "$HERE/lib/admin-path.sh"
log() { printf 'rig-users: %s\n' "$*"; }
warn() { printf 'rig-users: WARNING: %s\n' "$*" >&2; }
@ -138,10 +136,6 @@ done <<< "$PARSED"
# --- guards ------------------------------------------------------------------
[ "$(id -u)" -eq 0 ] || die "must run as root"
# Root is not enough: the admin binaries must also be reachable (#139). This
# sits beside the root check so identity and capability are asserted together,
# and BEFORE any prompt or download — a token typed for a doomed run is waste.
require_admin_bins useradd usermod groupadd
# Identity management gates its INVOKER, not just its uid: %rig's sudoers rule
# is binary-scoped but not argument-scoped, so without this gate a rig-role
@ -201,26 +195,6 @@ if [ "$NEED_SUDO" -eq 1 ] && ! command -v sudo >/dev/null 2>&1; then
DEBIAN_FRONTEND=noninteractive apt-get install -y -qq sudo
CHANGED=1
fi
# visudo is checked HERE and not beside the root check, because until the block
# above has run there is a legitimate reason for it to be absent: sudo is not
# installed yet, and apply is what installs it. Above, a missing visudo would
# be indistinguishable from that, so the refusal would fire on a healthy box.
#
# Below, it is unambiguous. sudo is present, so `visudo` missing means only one
# thing: /usr/sbin is off PATH. And it MUST refuse here rather than be left to
# the sudoers block further down, because that block asks `command -v visudo`
# and treats false as "no sudo on the box means no role needed it" — which on a
# PATH-shorn root is FALSE TWICE. A role does need it, sudo is installed, and
# apply would finish reporting success having silently never written the
# sudoers drop-in: the users get their roles and not the escalation the roles
# are FOR. That is the failure this whole issue is about (a wrong effective
# state reported as success, #12), in its quietest form — the other three sites
# at least crash. Refusing before the first mutation is what keeps it loud.
#
# It sits before `groupadd` below, so nothing has been converged when it fires.
if [ "$NEED_SUDO" -eq 1 ]; then
require_admin_bins visudo
fi
# --- groups ------------------------------------------------------------------
groupadd -f rig-admin

View file

@ -77,12 +77,7 @@ fetch_and_verify_sha256() {
printf 'checksum verified (%s)\n' "$got"
}
# CIBOX_BIN is a TEST-ONLY override, in the same spirit as bootstrap-undo.sh's
# RIG_FORGEJO_RUNNER_DIR: the production default is the only path the mechanism
# ever uses, but test/cli.sh must be able to drive this script on a box that
# already has a real runner installed. Without it the early-exit below fires
# against the host and the checksum checks silently test nothing (#136).
BIN="${CIBOX_BIN:-/usr/local/bin/forgejo-runner}"
BIN=/usr/local/bin/forgejo-runner
if [ -x "$BIN" ]; then
exit 0

View file

@ -11,7 +11,7 @@ release (#105, and #107's debt).
- **A throwaway Debian 13 machine** you can format, reached as root. The
drill hardens its sshd, renames it, joins it to a tailnet, and installs
box/Incus, Coolify and Actions runners on it. It is not coming back.
box/Incus, Coolify and a GitHub runner on it. It is not coming back.
The machine is its own reset — there is no teardown script and no need
for one.
- **The pinned candidate refs, both of them.** `--rig-ref` and
@ -34,18 +34,6 @@ release (#105, and #107's debt).
and does something trivial (`echo drilled`). Tokens come from an
authenticated `gh`, or from `RUNNER_TOKEN` / `RUNNER_REMOVE_TOKEN`.
Without a fork the leg **skips, loudly, into the record**.
- **For leg 3's Forgejo half** (#129): `--forgejo-instance <url>` and
`--forgejo-runner-repo <owner>/<repo>`, where that repo carries the same
`workflow_dispatch` workflow — but with `runs-on: drill`, because a
Forgejo runner matches the bare label it registered with. Tokens come from
`FORGEJO_RUNNER_TOKEN` (a registration token) or `FORGEJO_API_TOKEN`, which
mints one and is also what dispatches the job. `--forgejo-ref` names the
branch to dispatch (default `main`): Forgejo's dispatch endpoint requires a
ref in the body, where GitHub's defaults to the repo's default branch.
Without an instance and a repo this half **skips, loudly and separately**.
Note there is no removal token — Forgejo has no deregistration endpoint, so
the leg removes locally and the record tells you to delete the stale runner
row by hand.
- **For leg 4** (coolify): a version pin, `--coolify-version 4.1.2`.
No pin, no leg — rig's own `coolify install` refuses to default a
version and so does its drill. The skip is recorded.
@ -92,11 +80,7 @@ passes, failures and skips separately.
skip, exit 0) survives into the record as a SKIP, never a pass.
3. **Runner lifecycle** — register against the fork, dispatch the drill
workflow and watch the runner take it, deregister, and assert the
box's registration is actually gone. Runs **once per forge**: `rig runner`
against GitHub, then `rig forgejo-runner` against a Forgejo instance
(#129). Both families ship, so a release that evidences only one
evidences half of what it ships; each half skips separately, so a record
can honestly show one forge drilled and the other not.
box's registration is actually gone.
4. **Coolify** — installed at the pin, `AUTOUPDATE=false` landed in the
effective `.env`, container running.

View file

@ -4,16 +4,13 @@
# ⚠ DESTRUCTIVE, AND MEANT TO BE. Run it on a THROWAWAY Debian machine you
# can format. It wipes any installed rig and reinstalls from the pinned
# ref, hardens sshd, sets the hostname, joins the tailnet, installs box
# and its Incus stack, installs Coolify and Actions runners (GitHub, and
# Forgejo when --forgejo-instance is given).
# and its Incus stack, installs Coolify and a GitHub Actions runner.
# Never run it on a machine you care about.
#
# TS_AUTHKEY=tskey-... bash drill/drill.sh \
# --rig-ref release/0.4.0 --box-ref 0.9.0 \
# --users ./drill-users --run-id drill-2026-07-24-a \
# --coolify-version 4.1.2 --runner-repo you/rig \
# --forgejo-instance https://forgejo.example.com \
# --forgejo-runner-repo you/drill-probe --yes
# --coolify-version 4.1.2 --runner-repo you/rig --yes
# (--box-ref is a tag: since #103 the box that ships is the BOX_RELEASE tag.)
# rig's drill asserts CONVERGENCE — a machine reaches its role, idempotently.
# The legs (drills/README.md, issue #105):
@ -25,10 +22,6 @@
# the isolation boundary is box's drill's assertion, not this one's).
# 2. db — the real dump/restore round-trip, test/db-integration.sh.
# 3. runner lifecycle — register, take a job, deregister, against a fork.
# Runs once per forge: `rig runner` against GitHub (--runner-repo), and
# `rig forgejo-runner` against a Forgejo instance (--forgejo-instance +
# --forgejo-runner-repo). Both forges ship, so both need evidence; each
# skips loudly and separately when its inputs are absent (#129).
# 4. coolify install — at a pinned version, AUTOUPDATE=false.
#
# Execution order is 1, 4, 2, 3 — coolify's installer is what puts Docker on
@ -74,11 +67,6 @@ RECORD="${DRILL_RECORD:-}"
COOLIFY_VERSION="${DRILL_COOLIFY_VERSION:-}"
RUNNER_REPO="${DRILL_RUNNER_REPO:-}"
RUNNER_WORKFLOW="${DRILL_RUNNER_WORKFLOW:-drill.yml}"
FJ_INSTANCE="${DRILL_FORGEJO_INSTANCE:-}"
FJ_RUNNER_REPO="${DRILL_FORGEJO_RUNNER_REPO:-}"
# The branch the dispatch names. Forgejo's dispatch endpoint requires a ref in
# the body — unlike GitHub's, which defaults to the repo's default branch.
FJ_REF="${DRILL_FORGEJO_REF:-main}"
YES=0
while [ $# -gt 0 ]; do
@ -95,10 +83,7 @@ while [ $# -gt 0 ]; do
--coolify-version) COOLIFY_VERSION="$2"; shift 2 ;;
--runner-repo) RUNNER_REPO="$2"; shift 2 ;;
--runner-workflow) RUNNER_WORKFLOW="$2"; shift 2 ;;
--forgejo-instance) FJ_INSTANCE="$2"; shift 2 ;;
--forgejo-runner-repo) FJ_RUNNER_REPO="$2"; shift 2 ;;
--forgejo-ref) FJ_REF="$2"; shift 2 ;;
-h|--help) sed -n '2,40p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
-h|--help) sed -n '2,33p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) echo "drill: unknown option: $1 (see --help)" >&2; exit 2 ;;
esac
done
@ -123,133 +108,6 @@ phase(){ printf '\n\033[1m══ %s\033[0m\n' "$*"; }
LEG_NAMES=(); LEG_RESULTS=()
leg() { LEG_NAMES+=("$1"); LEG_RESULTS+=("$2"); }
# forgejo_run_verdict <pre_id> <tasks-json> — the verdict for OUR dispatch:
# success | failed | pending. Reads GET /repos/{o}/{r}/actions/tasks, whose
# shape is NOT GitHub's and was measured against forgejo.heavyduty.builders
# (8.0.3+gitea-1.22.0) on 2026-07-30 rather than read from the docs:
#
# * There is no `conclusion` field. `status` carries the terminal outcome
# directly ("success"), where GitHub splits status:completed +
# conclusion:success. Reading `conclusion` here gets an empty string on
# every run, which would grade a green job as failed.
# * `id` is a GLOBAL task id; the run's own URL ends in `run_number`. The
# pre-dispatch guard therefore compares `id`, exactly as the GitHub leg
# compares databaseId — an old run must never be read as this one.
# * The payload lists ASSIGNED tasks only. A run sitting queued is simply
# absent (measured: 200s of total_count 0 while the web UI showed the run
# as "job is not started"). So "no new id" is the ONLY signal that the
# runner never took the job — which is the verdict this leg exists for.
# * A task appears when ASSIGNED, so it can be seen mid-flight. A
# non-terminal status is pending, not failed; grading a running job as a
# failure would make the leg flaky inside its own watch window.
#
# EVERY entry is inspected, and the NEWEST id above pre_id decides — never
# entry[0]. actions/tasks accumulates, so the moment a repo is drilled twice
# our run shares the payload with older ones, and nothing documents the sort
# order. Reading the first entry made a green job report as a timeout, a false
# FAILURE on the gate this leg exists to provide (grok/kimi on !130).
#
# grep-and-sed, not jq: a throwaway drill machine has neither jq nor an
# authenticated forge CLI, the same constraint json_field() carries in
# commands/lib/runner-config.sh. `id` is a bare number, which json_field's
# quoted-value shape cannot read, so this reads both forms itself. Newlines are
# stripped first so a pretty-printed payload parses identically to a compact
# one — the instance documents neither.
forgejo_run_verdict() {
local pre="$1" file="$2" obj id best_id="" best_st=""
[ -r "$file" ] || { echo pending; return 0; }
while IFS= read -r obj; do
[ -n "$obj" ] || continue
id="$(printf '%s' "$obj" | grep -o '"id"[[:space:]]*:[[:space:]]*[0-9][0-9]*' \
| head -n1 | sed 's/.*:[[:space:]]*//')"
[ -n "$id" ] || continue
# Strictly newer than the pre-dispatch id. Equal is the run that was
# already there; lower is older still.
if [ -n "$pre" ]; then
[ "$id" -gt "$pre" ] 2>/dev/null || continue
fi
if [ -z "$best_id" ] || [ "$id" -gt "$best_id" ] 2>/dev/null; then
best_id="$id"
best_st="$(printf '%s' "$obj" | grep -o '"status"[[:space:]]*:[[:space:]]*"[^"]*"' \
| head -n1 | sed 's/.*:[[:space:]]*"//; s/"$//')"
fi
done <<EOF
$(tr -d '\n' < "$file" 2>/dev/null | grep -o '{[^{}]*}')
EOF
[ -n "$best_id" ] || { echo pending; return 0; }
case "$best_st" in
success) echo success ;;
failure | cancelled | skipped | timedout) echo failed ;;
*) echo pending ;;
esac
}
# forgejo_leg_row <install_ok> <status_ok> <took> <remove_ok> <absent_ok>
# The record row for the Forgejo runner leg. PASS requires the WHOLE lifecycle,
# not just the take-a-job outcome.
#
# Keying the row on <took> alone let it read "PASS — registered, took a job,
# removed" when install had failed, because the dispatched job only needs
# SOMETHING answering runs-on: drill — and this leg removes locally, telling the
# operator to delete the stale runner by hand, so a leftover drill-labeled
# runner from the previous drill is the designed-for aftermath rather than a
# contrived case (codex/grok/kimi on !130). drills/<v>.md is the release's
# durable evidence; a row claiming a lifecycle that did not happen is exactly
# what the gate exists to refuse.
#
# The drill's exit code was never wrong here — every one of those failures also
# called `no`. What was wrong is the row, and the row is what outlives the run.
forgejo_leg_row() {
local install_ok="$1" status_ok="$2" took="$3" remove_ok="$4" absent_ok="$5"
if [ "$install_ok" != 1 ] || [ "$status_ok" != 1 ] \
|| [ "$remove_ok" != 1 ] || [ "$absent_ok" != 1 ]; then
echo "FAIL — see Failed below"
return 0
fi
case "$took" in
success) echo "PASS — registered, took a job, removed (stale row needs deleting by hand)" ;;
none) echo "PARTIAL — registered and removed; took a job: not attempted (no FORGEJO_API_TOKEN)" ;;
*) echo "FAIL — see Failed below" ;;
esac
}
# forgejo_max_task_id <tasks-json> — the highest numeric task id in the payload,
# empty when there is none. This is the PRE-DISPATCH baseline, and it must fold
# max exactly as forgejo_run_verdict does: taking the first id instead names an
# OLD run as the baseline whenever the payload is not newest-first (the order is
# undocumented). A later poll that finds the same body then reads the PREVIOUS
# drill's run as this dispatch's result — a false PASS on the take-a-job
# assertion, which is worse than the false FAIL the same mistake caused inside
# the verdict (grok/kimi, !130). The GitHub leg is safe from this only because
# `gh run list --limit 1` contracts newest-first; this API contracts nothing.
forgejo_max_task_id() {
local file="$1" id best=""
[ -r "$file" ] || return 0
while IFS= read -r id; do
[ -n "$id" ] || continue
if [ -z "$best" ] || [ "$id" -gt "$best" ] 2>/dev/null; then best="$id"; fi
done <<EOF
$(tr -d '\n' < "$file" 2>/dev/null \
| grep -o '"id"[[:space:]]*:[[:space:]]*[0-9][0-9]*' | sed 's/.*:[[:space:]]*//')
EOF
printf '%s\n' "$best"
}
# forgejo_token_verdict <resolved_token> <api_token> — ok | mint-failed | no-source.
# #129's acceptance: "Token source present but the instance is unreachable ->
# the leg FAILS; it must not skip and must not pass". A mint that yields
# nothing — unreachable instance, under-scoped token, wrong repo — is a
# CONFIGURED leg failing, and reporting it as "no token source" both writes
# SKIPPED where the record owes a FAIL and sends the operator to check an env
# var they already set. Absent inputs are the only honest skip.
forgejo_token_verdict() {
if [ -n "$1" ]; then echo ok
elif [ -n "$2" ]; then echo mint-failed
else echo no-source
fi
}
# run_logged <log> <cmd...> — run a long command with its narration in a file
# and a dot every 5s on the terminal: a silent multi-minute apt/install run is
# indistinguishable from a wedge, and that ambiguity has cost box whole
@ -494,7 +352,7 @@ This will, ON THIS HOST ($(hostname)):
· run 'rig bootstrap $ROLE --users $USERS_FILE' — sshd hardening, hostname
change, tailnet join, box ($BOXREPO@$BOXREF) + its Incus stack — TWICE
(the second run is the idempotence assertion)
· install Coolify${COOLIFY_VERSION:+ $COOLIFY_VERSION}, a GitHub runner${RUNNER_REPO:+ against $RUNNER_REPO} and a Forgejo runner${FJ_INSTANCE:+ against $FJ_INSTANCE}${FJ_RUNNER_REPO:+ ($FJ_RUNNER_REPO)}
· install Coolify${COOLIFY_VERSION:+ $COOLIFY_VERSION} and a GitHub runner${RUNNER_REPO:+ against $RUNNER_REPO}
Only do this on a THROWAWAY machine you can format.
EOF
[ -t 0 ] || { echo "drill: no TTY to confirm on — pass --yes if you mean it." >&2; exit 2; }
@ -832,139 +690,6 @@ else
fi
fi
# =============================================================================
phase "Leg 3 — forgejo runner lifecycle against an instance"
# =============================================================================
# The same leg as above, for the other forge. Both forges ship a runner family
# (#109 added `rig forgejo-runner` beside `rig runner`), so a release that
# evidences only GitHub evidences half of what it ships (#129).
#
# Three things differ from the GitHub leg, all measured against
# forgejo.heavyduty.builders (8.0.3+gitea-1.22.0) on 2026-07-30, not read:
#
# * Scope is the TOKEN's, never a flag — `rig forgejo-runner install` refuses
# --repo on purpose (commands/forgejo-runner-install.sh:160). The repo here
# is only where the registration token is minted from, and where the
# dispatched workflow lives.
# * The mint path is /repos/<o>/<r>/actions/runners/registration-token. The
# instance's own swagger documents /repos/<o>/<r>/runners/registration-token
# — WITHOUT /actions/ — and that path 404s. Do not "fix" this to match the
# published API reference.
# * There is no deregistration endpoint, so there is no removal token and no
# remote deregistration: `rig forgejo-runner remove` is local-only by
# design (commands/forgejo-runner-remove.sh:7-11) and the runner row
# survives in the UI until a human deletes it. The record says so rather
# than implying a clean remote teardown the way the GitHub leg can.
#
# Tokens: FORGEJO_RUNNER_TOKEN (a registration token) is used directly; else
# FORGEJO_API_TOKEN mints one from FJ_RUNNER_REPO. Without an instance, a repo,
# or a token source the leg SKIPS loudly and the record says it did not run.
if [ -z "$FJ_INSTANCE" ] || [ -z "$FJ_RUNNER_REPO" ]; then
skip "forgejo runner lifecycle: no --forgejo-instance/--forgejo-runner-repo given — the leg did not run"
leg "forgejo runner lifecycle" "SKIPPED — no instance/repo provided"
else
fj_reg="${FORGEJO_RUNNER_TOKEN:-}"
if [ -z "$fj_reg" ] && [ -n "${FORGEJO_API_TOKEN:-}" ]; then
fj_reg="$(curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \
"${FJ_INSTANCE%/}/api/v1/repos/${FJ_RUNNER_REPO}/actions/runners/registration-token" 2>/dev/null \
| grep -o '"token"[[:space:]]*:[[:space:]]*"[^"]*"' | head -n1 \
| sed 's/.*:[[:space:]]*"//; s/"$//')"
fi
fj_tok_verdict="$(forgejo_token_verdict "$fj_reg" "${FORGEJO_API_TOKEN:-}")"
if [ "$fj_tok_verdict" = mint-failed ]; then
# Configured, and it did not work. Never a skip: see forgejo_token_verdict.
no "registration-token mint FAILED against ${FJ_INSTANCE} — is it reachable, and does FORGEJO_API_TOKEN own ${FJ_RUNNER_REPO}? (the token is never printed)"
leg "forgejo runner lifecycle ($FJ_RUNNER_REPO)" "FAIL — registration-token mint failed"
elif [ "$fj_tok_verdict" = no-source ]; then
skip "forgejo runner lifecycle: no FORGEJO_RUNNER_TOKEN and no FORGEJO_API_TOKEN to mint one — the leg did not run"
leg "forgejo runner lifecycle ($FJ_RUNNER_REPO)" "SKIPPED — no registration token source"
else
FJ_NAME="drill-$(hostname)-$$"
fj_install_ok=0 fj_status_ok=0 fj_remove_ok=0 fj_absent_ok=0
# The label MUST carry a docker:// image: forgejo-runner runs jobs in
# containers, and a bare label leaves runs-on matched but unrunnable.
if FORGEJO_RUNNER_TOKEN="$fj_reg" run_logged /tmp/drill-forgejo-runner-install.log \
rig forgejo-runner install --instance "$FJ_INSTANCE" --name "$FJ_NAME" \
--labels 'drill:docker://node:22-bookworm'; then
ok "rig forgejo-runner install --instance $FJ_INSTANCE exited 0 (registered as $FJ_NAME)"
fj_install_ok=1
else
no "forgejo-runner install FAILED — tail: $(tail -3 /tmp/drill-forgejo-runner-install.log | tr '\n' ' ')"
fi
if rig forgejo-runner status 2>/dev/null | grep -qF "${FJ_INSTANCE%/}"; then
ok "forgejo-runner status names the instance: $FJ_INSTANCE"; fj_status_ok=1
else
no "forgejo-runner status does not name ${FJ_INSTANCE}"
fi
fj_took=none
# Do NOT dispatch once install or status has failed. The job would be taken
# by whatever else answers runs-on: drill — a stale runner this leg's own
# hand-delete caveat leaves behind — and its success would be evidence about
# someone else's runner (codex/grok/kimi, !130).
if [ "$fj_install_ok" != 1 ] || [ "$fj_status_ok" != 1 ]; then
skip "took a job: not attempted — install or status failed, and a foreign runner answering 'drill' could only manufacture a false pass"
elif [ -n "${FORGEJO_API_TOKEN:-}" ]; then
fj_api="${FJ_INSTANCE%/}/api/v1/repos/${FJ_RUNNER_REPO}"
# Read the newest ASSIGNED task id BEFORE dispatching, same guard as the
# GitHub leg: an already-completed run must never be read as ours.
fj_pre_body="$(curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \
"$fj_api/actions/tasks" 2>/dev/null || echo '{}')"
printf '%s' "$fj_pre_body" > /tmp/drill-forgejo-pre.json
fj_pre="$(forgejo_max_task_id /tmp/drill-forgejo-pre.json)"
if curl -fsSL -o /dev/null -X POST -H "Authorization: token ${FORGEJO_API_TOKEN}" \
-H "Content-Type: application/json" -d "{\"ref\":\"${FJ_REF}\"}" \
"$fj_api/actions/workflows/${RUNNER_WORKFLOW}/dispatches" 2>/dev/null; then
inf "dispatched $RUNNER_WORKFLOW on $FJ_RUNNER_REPO — waiting for the runner to take it (≤5 min)…"
fj_took=timeout
for _i in $(seq 1 30); do
sleep 10
curl -fsSL -H "Authorization: token ${FORGEJO_API_TOKEN}" \
"$fj_api/actions/tasks" -o /tmp/drill-forgejo-tasks.json 2>/dev/null || continue
case "$(forgejo_run_verdict "${fj_pre:-}" /tmp/drill-forgejo-tasks.json)" in
success) fj_took=success; break ;;
failed) fj_took=failed; break ;;
*) : ;; # pending — queued, or assigned and still running
esac
done
else
fj_took=nodispatch
fi
case "$fj_took" in
success) ok "the forgejo runner took a job and it succeeded ($RUNNER_WORKFLOW)" ;;
failed) no "the dispatched job completed UNSUCCESSFULLY — the runner ran it, the workflow failed; read the run on $FJ_RUNNER_REPO" ;;
# A queued task is INVISIBLE in this API until a runner claims it, so a
# timeout means "nothing ever took it". Two causes, and the second one
# is not rig's: the runs-on label may not match, or the daemon's poller
# can go quiet — a restarted daemon claims a minutes-old backlog in
# about a second. Check 'systemctl restart forgejo-runner' before
# reading this as a rig defect.
timeout) no "the dispatched job was never taken within 5 min — check the workflow's runs-on is 'drill', then restart forgejo-runner and re-read (a quiet poller looks exactly like this)" ;;
nodispatch) no "could not dispatch $RUNNER_WORKFLOW on $FJ_RUNNER_REPO — does it carry that workflow, with workflow_dispatch, on its default branch?" ;;
esac
else
skip "took a job: not attempted — no FORGEJO_API_TOKEN to dispatch $RUNNER_WORKFLOW with"
fi
# No removal token exists on this forge — remove is local by design.
if rig forgejo-runner remove >/dev/null 2>&1; then
note "forgejo-runner removed locally — Forgejo has no deregistration endpoint, so DELETE the stale '$FJ_NAME' row under $FJ_RUNNER_REPO > Settings > Actions > Runners by hand"
fj_remove_ok=1
else
no "forgejo-runner remove FAILED"
fi
if rig forgejo-runner status >/dev/null 2>&1; then
no "forgejo-runner status still answers after remove — the removal did not take"
else
ok "forgejo-runner status confirms: nothing registered"; fj_absent_ok=1
fi
leg "forgejo runner lifecycle ($FJ_RUNNER_REPO)" \
"$(forgejo_leg_row "$fj_install_ok" "$fj_status_ok" "$fj_took" \
"$fj_remove_ok" "$fj_absent_ok")"
fi
fi
# =============================================================================
phase "Summary"
# =============================================================================

View file

@ -113,7 +113,6 @@ Candidate refs: box@1a2b3c4 (BOX_REF=release/0.4.0), rig@5d6e7f8, cast@9a0b1c2.
| --host yes: pinned box installed, host stack up | PASS — box doctor clean |
| `test/db-integration.sh` | PASS — 14 passed, 0 failed |
| runner lifecycle against a fork | PASS — registered, took a job, deregistered clean |
| forgejo runner lifecycle (you/drill-probe) | PASS — registered, took a job, removed (stale row needs deleting by hand) |
| coolify install (4.1.2) | PASS (6 min) |
Failed: `rig users apply` left one revoked key in `authorized_keys`

View file

@ -156,9 +156,8 @@ UNDO_FIX="$(mktemp -d)"
UNDO_BIN="$UNDO_FIX/bin"
UNDO_MARKER="$UNDO_FIX/role"
UNDO_RUNNER="$UNDO_FIX/runner"
UNDO_FJRUNNER="$UNDO_FIX/fjrunner"
UNDO_CALLS="$UNDO_FIX/tailscale.calls"
mkdir -p "$UNDO_BIN" "$UNDO_RUNNER" "$UNDO_FJRUNNER"
mkdir -p "$UNDO_BIN" "$UNDO_RUNNER"
cat > "$UNDO_BIN/tailscale" <<'SH'
#!/usr/bin/env bash
printf '%s\n' "$*" >> "$UNDO_CALLS"
@ -170,13 +169,8 @@ if [ "${1:-}" = -u ]; then printf '0\n'; else exec /usr/bin/id "$@"; fi
SH
chmod +x "$UNDO_BIN/tailscale" "$UNDO_BIN/id"
undo() {
# RIG_FORGEJO_RUNNER_DIR is as load-bearing as RIG_RUNNER_DIR: without it
# bootstrap-undo.sh scans /home/*/forgejo-runner/.runner and the systemd unit
# on the REAL box, so these checks fail on any machine that has actually been
# drilled or used as a ci-box — which is the machine that matters (#136).
env PATH="$UNDO_BIN:$PATH" UNDO_CALLS="$UNDO_CALLS" \
RIG_ROLE_MARKER="$UNDO_MARKER" RIG_RUNNER_DIR="$UNDO_RUNNER" \
RIG_FORGEJO_RUNNER_DIR="$UNDO_FJRUNNER" \
"$ROOT/bin/rig" bootstrap --undo
}
undo_untouched() {
@ -202,8 +196,7 @@ rm -f "$UNDO_RUNNER/.runner"
check "bootstrap --undo: failed logout is loud" \
1 "role marker kept" env TAILSCALE_LOGOUT_FAIL=1 PATH="$UNDO_BIN:$PATH" \
UNDO_CALLS="$UNDO_CALLS" RIG_ROLE_MARKER="$UNDO_MARKER" \
RIG_RUNNER_DIR="$UNDO_RUNNER" RIG_FORGEJO_RUNNER_DIR="$UNDO_FJRUNNER" \
"$ROOT/bin/rig" bootstrap --undo
RIG_RUNNER_DIR="$UNDO_RUNNER" "$ROOT/bin/rig" bootstrap --undo
check "bootstrap --undo: failed logout preserves the marker" 0 "" test -e "$UNDO_MARKER"
: > "$UNDO_CALLS"
check "bootstrap --undo: proven rig join succeeds" 0 "tailnet join removed" undo
@ -3283,45 +3276,10 @@ cibox_src_matches_install() {
# shellcheck source=/dev/null
. "$ROOT/commands/lib/templates.sh"
template_parse_env "$dir/template.env" >/dev/null || return 2
# BIN carries a test-only override (#136), so the agreement this asserts is
# with the DEFAULT — the only path the mechanism itself ever uses.
grep -qF "CIBOX_BIN:-$TPL_CLI_SRC" "$dir/install.sh"
grep -qF "BIN=$TPL_CLI_SRC" "$dir/install.sh"
}
check "ci-box: CLI_SRC is the path its install.sh installs" 0 "" cibox_src_matches_install
# codex/kimi on !137: the two wiring lines this change exists to add could be
# deleted tomorrow and the suite stayed 786/786 on any host without a real
# Forgejo runner — hermetic today, unpinned. #136's task list names the guard
# verbatim: "a check that fails if either group can see host state".
#
# These assert on the SUITE's own helpers, not the production knobs — the knobs
# are already covered above. What must go red is a deletion on the test side,
# because that is the regression that silently reintroduces host dependence.
undo_is_sealed() {
sed -n '/^undo() {/,/^}/p' "$0" | grep -q 'RIG_FORGEJO_RUNNER_DIR='
}
cibox_run_is_sealed() {
sed -n '/^cibox_run() {/,/^}/p' "$0" | grep -q 'CIBOX_BIN='
}
check "hermetic: undo() seals the Forgejo-runner host scan" 0 "" undo_is_sealed
check "hermetic: cibox_run() seals the real /usr/local/bin lookup" 0 "" cibox_run_is_sealed
# The failed-logout check builds its own env rather than calling undo(), so it
# needs the same seal — and it is the one that was missed first time round.
inline_undo_is_sealed() {
# Locate the REAL check by line number and read only its own block. Anchoring
# on a string and grepping the whole file cannot work here: any pattern this
# function searches for necessarily appears inside this function, so the
# search matches itself and can never fail. kimi caught the first version of
# that on !137; the anchored second version had the identical flaw for the
# identical reason. head -1 takes the real check (~:202), never this body.
local start end
start="$(grep -n 'failed logout is loud' "$0" | head -1 | cut -d: -f1)"
[ -n "$start" ] || return 1
end=$((start + 5))
sed -n "${start},${end}p" "$0" | grep -q RIG_FORGEJO_RUNNER_DIR
}
check "hermetic: the hand-rolled undo invocation is sealed too" 0 "" inline_undo_is_sealed
# Registration holds a credential, so it must NOT be in the definition: a
# tenant install is creds-free by contract — box auto-runs it at mint, holding
# nothing. Registration is the operator's separate, out-loud act.
@ -3334,121 +3292,6 @@ check "ci-box: its install.sh does not register" 1 "" \
check "ci-box: bootstrap-tenant does not read the staging dir" 1 "" \
grep -q 'docs/templates' "$ROOT/commands/bootstrap-tenant.sh"
# --- admin binaries must be reachable, not just root (#139) ------------------
# Being root and being able to FIND the admin binaries are different facts, and
# rig asserted only the first. `su` without `-`, sudo with a sanitised
# secure_path, and several container images all give a root shell whose PATH
# carries no /usr/sbin — where useradd lives. Reported from a real ci-box:
#
# root@ci-forgejo-box:/home/dev# rig forgejo-runner install --instance …
# forgejo runner registration token:
# …/forgejo-runner-install.sh: line 250: useradd: command not found
#
# Note where it died: AFTER reading a registration token off the operator's
# terminal. A secret typed for a run that could never succeed is the avoidable
# half of the bug, so the refusal has to come before the prompt.
#
# The root check fires first and correctly, so these stub `id -u` to 0 — the
# idiom the bootstrap --undo block above already uses — to reach the preflight.
ADMPATH_DIR="$(mktemp -d)"
mkdir -p "$ADMPATH_DIR/bin"
# shellcheck disable=SC2016 # the body is shell source being written, not expanded
printf '#!/usr/bin/env bash\nif [ "${1:-}" = -u ]; then printf "0\\n"; else exec /usr/bin/id "$@"; fi\n' \
> "$ADMPATH_DIR/bin/id"
chmod +x "$ADMPATH_DIR/bin/id"
# A PATH with the stub and the ordinary bindirs, but deliberately no /usr/sbin.
SBINLESS="$ADMPATH_DIR/bin:/usr/local/bin:/usr/bin:/bin"
adm_run() { env PATH="$SBINLESS" "$@" 2>&1; }
adm_prompted() { # did it read a token before refusing?
adm_run "$@" | grep -qi 'registration token:'
}
check "preflight: forgejo-runner install refuses a PATH with no /usr/sbin" 1 "useradd" \
adm_run "$ROOT/commands/forgejo-runner-install.sh" --instance https://f.example.com
check "preflight: …and names PATH as the cause, not just the missing binary" 1 "PATH" \
adm_run "$ROOT/commands/forgejo-runner-install.sh" --instance https://f.example.com
check "preflight: …and refuses BEFORE prompting for a token" 1 "" \
adm_prompted "$ROOT/commands/forgejo-runner-install.sh" --instance https://f.example.com
check "preflight: the GitHub runner installer refuses too" 1 "useradd" \
adm_run "$ROOT/commands/runner-install.sh" --repo o/r
check "preflight: users apply refuses before it converges anything" 1 "useradd" \
adm_run "$ROOT/commands/users-apply.sh" --file /dev/null
# A sbin-less PATH loses ALL of /usr/sbin at once, so the checks above can only
# ever prove the FIRST binary is named — `useradd` wins every race and would
# hide a preflight that forgot the others. These fixtures resolve the earlier
# binaries and withhold exactly one, which is the only way to show the sweep
# covers what each command actually calls (the #75 lesson: a sweep that misses
# one call site is how the next bug gets in).
#
# The withheld binary is real: these are PATHs, not stubs of the tools
# themselves — a stub that silently succeeded would let convergence run on.
admstub() { # admstub <dir> <bin>... — a bindir resolving id(0) plus <bin>...
local d="$1"; shift
mkdir -p "$d"
# shellcheck disable=SC2016 # shell source being written, not expanded
printf '#!/usr/bin/env bash\nif [ "${1:-}" = -u ]; then printf "0\\n"; else exec /usr/bin/id "$@"; fi\n' > "$d/id"
chmod +x "$d/id"
local b
for b in "$@"; do printf '#!/usr/bin/env bash\nexit 0\n' > "$d/$b"; chmod +x "$d/$b"; done
}
adm_with() { # adm_with <bindir> <cmd>...
local d="$1"; shift
env PATH="$d:/usr/local/bin:/usr/bin:/bin" "$@" 2>&1
}
# adm_saw <bindir> <pattern> <cmd>... — exit 0 if the run printed <pattern>.
# Used with `check ... 1 ""` to assert a thing did NOT happen (a token prompt,
# a group creation), the same shape as adm_prompted above.
adm_saw() {
local d="$1" pat="$2"; shift 2
adm_with "$d" "$@" | grep -qi -- "$pat"
}
# useradd resolves, usermod does not — the forgejo installer calls both, and
# only reaches usermod after the token has been spent.
admstub "$ADMPATH_DIR/no-usermod" useradd
check "preflight: forgejo-runner install names usermod when only useradd resolves" 1 "usermod" \
adm_with "$ADMPATH_DIR/no-usermod" "$ROOT/commands/forgejo-runner-install.sh" --instance https://f.example.com
check "preflight: …and still refuses before the token prompt" 1 "" \
adm_saw "$ADMPATH_DIR/no-usermod" 'registration token:' \
"$ROOT/commands/forgejo-runner-install.sh" --instance https://f.example.com
# users apply calls groupadd unconditionally, two lines into its convergence.
admstub "$ADMPATH_DIR/no-groupadd" useradd usermod
check "preflight: users apply names groupadd when useradd and usermod resolve" 1 "groupadd" \
adm_with "$ADMPATH_DIR/no-groupadd" "$ROOT/commands/users-apply.sh" --file /dev/null
# visudo is the quiet one. With sudo INSTALLED and /usr/sbin off PATH, the
# sudoers block reads `command -v visudo` as "no sudo on the box" and skips the
# drop-in — apply then reports success having granted roles without the
# escalation those roles exist for. So this fixture needs a role that wants
# sudo, and asserts the refusal by name rather than a silent success.
ADM_USERS="$ADMPATH_DIR/users"
printf '%s\n' 'dan admin ssh-ed25519 AAAAC3fixture dan@laptop' > "$ADM_USERS"
admstub "$ADMPATH_DIR/no-visudo" useradd usermod groupadd
check "preflight: users apply refuses a sudo-needing role when visudo is unreachable" 1 "visudo" \
adm_with "$ADMPATH_DIR/no-visudo" "$ROOT/commands/users-apply.sh" --file "$ADM_USERS" --yes
# …and the refusal must land BEFORE the groups are converged, or the machine is
# already half-changed when the operator reads it.
check "preflight: …before any group is created" 1 "" \
adm_saw "$ADMPATH_DIR/no-visudo" 'rig-admin' \
"$ROOT/commands/users-apply.sh" --file "$ADM_USERS" --yes
# A users file needing no sudo must NOT be refused for a missing visudo: the
# guard has to be as narrow as the need, or it refuses healthy boxes.
printf '%s\n' 'maria ops ssh-ed25519 AAAAC3fixture maria@mac' > "$ADMPATH_DIR/users-nosudo"
check "preflight: …but a file with no sudo-backed role is not refused for visudo" 1 "" \
adm_saw "$ADMPATH_DIR/no-visudo" 'visudo' \
"$ROOT/commands/users-apply.sh" --file "$ADMPATH_DIR/users-nosudo" --yes
# The fourth call site (#139 review): bootstrap-tenant adds the tenant user to
# the docker group AFTER installing docker and node, so an unguarded PATH fails
# it mid-convergence on a machine it has already changed.
# staging-box because it is the one tenant role defined in rig's own tree: the
# others resolve through the template registry, and a preflight test must not
# depend on the network to reach the guard it is testing.
admstub "$ADMPATH_DIR/no-any"
check "preflight: bootstrap-tenant refuses a sbin-less PATH before it converges" 1 "usermod" \
adm_with "$ADMPATH_DIR/no-any" "$ROOT/commands/bootstrap-tenant.sh" staging-box
rm -rf "$ADMPATH_DIR"
# --- rig forgejo-runner (#109) ----------------------------------------------
FR="$ROOT/commands/forgejo-runner-install.sh"
check "forgejo-runner: bare subcommand shows usage, exit 2" 2 "usage:" "$ROOT/bin/rig" forgejo-runner
@ -3464,22 +3307,12 @@ check "forgejo-runner: --version refuses a path, not a release number" 2 "releas
"$FR" --instance https://f.example.com --version ../../etc/passwd
check "forgejo-runner: --version refuses a non-numeric pin" 2 "release number like" \
"$FR" --instance https://f.example.com --version latest
# Reaching a gate AFTER --version parsing is the proof a good pin got THROUGH
# validation. Which gate depends on the uid: non-root hits "must run as root";
# act/Forgejo jobs run as uid 0 (no `runner` account — #144), so they sail past
# the root check and hit the unattended-token refuse instead. Both prove the
# same thing. GitHub-hosted ubuntu-latest is non-root and takes the first arm.
if [ "$(id -u)" -ne 0 ]; then
check "forgejo-runner: a plain release number passes validation" 1 "must run as root" \
"$FR" --instance https://f.example.com --version 12.13.2
check "forgejo-runner: a leading v is stripped before that check" 1 "must run as root" \
"$FR" --instance https://f.example.com --version v12.13.2
else
check "forgejo-runner: a plain release number passes validation" 1 "FORGEJO_RUNNER_TOKEN is unset" \
env -u FORGEJO_RUNNER_TOKEN "$FR" --instance https://f.example.com --version 12.13.2
check "forgejo-runner: a leading v is stripped before that check" 1 "FORGEJO_RUNNER_TOKEN is unset" \
env -u FORGEJO_RUNNER_TOKEN "$FR" --instance https://f.example.com --version v12.13.2
fi
# Reaching the root check is the proof a good pin got THROUGH validation: this
# runs as a normal user in CI, so "must run as root" is the next gate down.
check "forgejo-runner: a plain release number passes validation" 1 "must run as root" \
"$FR" --instance https://f.example.com --version 12.13.2
check "forgejo-runner: a leading v is stripped before that check" 1 "must run as root" \
"$FR" --instance https://f.example.com --version v12.13.2
# A schemeless host and a repo URL are the two ways an operator mis-states the
# instance, and only one of them would fail loudly on its own — a repo URL
# registers somewhere subtly wrong instead. Both refuse by name.
@ -3517,27 +3350,6 @@ check "forgejo-runner: remove warns about an orphaned unit" 0 "orphaned unit" \
check "forgejo-runner: remove never rm's an unguarded \$RUNNER_DIR path" 1 "" \
grep -qE '^rm -f "\$RUNNER_DIR' "$ROOT/commands/forgejo-runner-remove.sh"
# #135: ProtectHome=read-only made the whole home read-only and only RUNNER_DIR
# was punched back through, so forgejo-runner could not create $HOME/.cache and
# disabled its cache server on every install — actions/cache silently off, one
# error line in the journal, and `status` reporting a healthy runner.
#
# BOTH halves are asserted because the obvious one-line version is WORSE than
# the bug: listing the path in ReadWritePaths without creating it makes systemd
# refuse to start the unit at all ("Failed to set up mount namespacing"),
# measured on a live box. The directory must exist first.
FRI="$ROOT/commands/forgejo-runner-install.sh"
grep_cache_install() {
# shellcheck disable=SC2016 # $RUNNER_USER is literal text in the shipped file
grep -qE 'install -d .*-o "\$RUNNER_USER".*\.cache' "$FRI"
}
check "forgejo-runner: the cache dir is punched through ProtectHome" 0 "" \
grep -qE 'ReadWritePaths=.*\.cache' "$FRI"
check "forgejo-runner: …and the install creates it, owned by the runner user" 0 "" \
grep_cache_install
check "forgejo-runner: ProtectHome stays read-only (the cache is not an excuse to widen it)" 0 "" \
grep -qF 'ProtectHome=read-only' "$FRI"
check "forgejo-runner: status --help exits 0" 0 "usage:" "$ROOT/commands/forgejo-runner-status.sh" --help
check "forgejo-runner: remove --help exits 0" 0 "usage:" "$ROOT/commands/forgejo-runner-remove.sh" --help
@ -3595,47 +3407,6 @@ check "forgejo-runner: install converges the mode on EVERY run, not only at regi
fr_secure_every_run
check "forgejo-runner: status warns on a drifted mode" 0 "FORGEJO_RUNNER_FILE_MODE" \
grep -o "FORGEJO_RUNNER_FILE_MODE" "$ROOT/commands/forgejo-runner-status.sh"
# #133: `active` is the strongest signal this command has, and it proves only
# that a process exists. A poller can go quiet while the process stays up —
# measured for #129: a daemon logged "[poller] launched" and never fetched a
# job dispatched four minutes later, while a fresh daemon claimed the same
# queued task in one second. status is where an operator looks when nothing
# is obviously wrong, so it says so there.
FJS="$ROOT/commands/forgejo-runner-status.sh"
check "forgejo-runner: status says 'active' is not proof the runner is fetching" 0 "not proof" \
grep -o "not proof" "$FJS"
check "forgejo-runner: …and names the remedy, so the line is actionable" 0 "restart" \
grep -oi "systemctl restart forgejo-runner" "$FJS"
# It must NOT be a warning: nothing has been detected. An idle-but-healthy
# runner logs nothing either, so there is no signal that separates it from a
# stalled one — a WARNING on every status run would be crying wolf, and this
# file reserves warn for drift it has actually measured (the .runner mode).
check "forgejo-runner: the liveness note is informational, never a WARNING" 1 "" \
grep -nE 'warn ".*not proof' "$FJS"
# codex/kimi on !134: the three greps above prove the LINES EXIST; nothing
# proved they fire only when the unit is active. Deleting the state guard left
# the suite 790/790 green, so the acceptance boundary #133 cares about most —
# no misleading liveness note on an absent or inactive unit — was unprotected.
# Drive the decision instead, extracted the way test/drill.sh extracts its own.
# Its own scratch dir: $WORK is rm -rf'd at :3206, well before this block.
FJS_DIR="$(mktemp -d)"
FJSW="$FJS_DIR/fjs-note.sh"
{ printf '%s\n' 'log() { printf "rig-forgejo-runner: %s\\n" "$*"; }'
awk '/^forgejo_runner_liveness_note\(\) \{/,/^\}/' "$FJS"
} > "$FJSW"
check "forgejo-runner: liveness note extracted (guards the awk)" 0 "forgejo_runner_liveness_note() {" \
cat "$FJSW"
note_for() { bash -c '. "$1"; forgejo_runner_liveness_note "$2"' _ "$FJSW" "$1"; }
check "liveness note: an ACTIVE unit is told what active does not prove" 0 "not proof" note_for active
check "liveness note: …and is given the remedy" 0 "systemctl restart forgejo-runner" note_for active
note_is_empty() { [ -z "$(note_for "$1")" ]; }
check "liveness note: an INACTIVE unit gets nothing" 0 "" note_is_empty inactive
check "liveness note: an ABSENT unit (empty state) gets nothing" 0 "" note_is_empty ""
rm -rf "$FJS_DIR"
# The header contract at :25 — no token, no network call — survives this.
check "forgejo-runner: status still makes no network call" 1 "" \
grep -nE '^[^#]*(curl|wget) ' "$FJS"
rm -rf "$FRW"
# --- the checksum gate, DRIVEN not grepped (review !110) --------------------
@ -3689,7 +3460,7 @@ chmod +x "$CBSTUB/install"
cibox_run() { # cibox_run [VAR=val ...] — the REAL template install.sh, stubbed
rm -f "$CBW/installed"
env PATH="$CBSTUB:$PATH" CIBOX_BIN="$CBW/bin-under-test" \
env PATH="$CBSTUB:$PATH" \
CB_REDIRECT=https://code.forgejo.org/forgejo/runner/releases/tag/v9.9.9 \
CB_PAYLOAD="$CBW/payload" "$@" bash "$CIBOX"
}
@ -3829,44 +3600,6 @@ check "forgejo-runner: an explicit --labels on a rerun warns it was not applied"
grep -o -- "--labels was not applied" "$FR"
check "forgejo-runner: that warning is gated on --labels being EXPLICIT" 0 "LABELS_EXPLICIT" \
grep -o "LABELS_EXPLICIT" "$FR"
# #144 option B: plain converge warns only for known *retired* defaults — not
# for every RECORDED != current default (that would noise custom --labels
# forever, including drill's Leg 3). Drive the recogniser against fixtures
# (test/drill.sh extraction pattern) so a dead matcher cannot greppen green.
check "forgejo-runner: plain converge warns on a known retired default label set" 0 "registered with an older rig default label set" \
grep -o "registered with an older rig default label set" "$FR"
check "forgejo-runner: that warn says re-register only if you want the new label" 0 "re-register only if you want the new label" \
grep -o "re-register only if you want the new label" "$FR"
PRE_144_LABELS='ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04,docker:docker://node:22-bookworm'
CURRENT_LABELS="$(sed -n "s/^DEFAULT_LABELS='\\(.*\\)'$/\\1/p" "$FR")"
RETIRED_FNS="$(mktemp)"
awk '/^labels_are_a_retired_default\(\) \{/,/^\}/' "$FR" > "$RETIRED_FNS"
check "extraction guards the awk: labels_are_a_retired_default() landed" 0 "labels_are_a_retired_default() {" \
grep -F 'labels_are_a_retired_default() {' "$RETIRED_FNS"
# shellcheck source=/dev/null
. "$RETIRED_FNS"
check "retired-default: the pre-#144 default is recognised" 0 "" \
labels_are_a_retired_default "$PRE_144_LABELS"
check "retired-default: the CURRENT default is not drift" 1 "" \
labels_are_a_retired_default "$CURRENT_LABELS"
check "retired-default: an operator's own --labels map is never drift" 1 "" \
labels_are_a_retired_default 'drill:docker://node:22-bookworm'
check "retired-default: a near-miss of a past default is not a match" 1 "" \
labels_are_a_retired_default "${PRE_144_LABELS} "
rm -f "$RETIRED_FNS"
# Default map: slim ubuntu-latest (act), opt-in full, docker — pin the three
# so a silent drop of the full rider or a flip back to full-as-default fails.
check "forgejo-runner: DEFAULT_LABELS maps ubuntu-latest to act-22.04 (slim)" 0 "ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04" \
grep -o "ubuntu-latest:docker://ghcr.io/catthehacker/ubuntu:act-22.04" "$FR"
check "forgejo-runner: DEFAULT_LABELS offers ubuntu-latest-full as opt-in" 0 "ubuntu-latest-full:docker://ghcr.io/catthehacker/ubuntu:full-22.04" \
grep -o "ubuntu-latest-full:docker://ghcr.io/catthehacker/ubuntu:full-22.04" "$FR"
check "forgejo-runner: DEFAULT_LABELS comment records the slim/full size measurement" 0 "54.5 GB" \
grep -o "54.5 GB" "$FR"
# ci.yml must install shellcheck when the image lacks it (#144 option B).
check "ci.yml: shellcheck step installs the tool when missing" 0 "command -v shellcheck" \
grep -o "command -v shellcheck" "$ROOT/.github/workflows/ci.yml"
check "ci.yml: shellcheck install uses sudo (GitHub path; no-op on act as root)" 0 "sudo apt-get install -y shellcheck" \
grep -o "sudo apt-get install -y shellcheck" "$ROOT/.github/workflows/ci.yml"
# The GitHub sibling is the precedent this restores — pin that it still scopes
# its own write, so the two cannot drift apart again.
gh_labels_write_is_scoped() {
@ -3931,40 +3664,6 @@ check "release.yml: the caller was not absolutised (docs-sync would go red)" 0 "
check "labels.yml: the caller was not absolutised either" 0 "count=0" \
wf_count labels.yml '^[[:space:]]*uses:[[:space:]]*https://.*ceremony/\.github/workflows/'
# The reconciler reads the machine roster from labels.conf while contributors
# read the prose roster. A Forgejo migration that updates only one side makes
# the required-verdict set differ depending on who is reading it.
configured_panel() {
sed -n 's/^panel=//p' "$ROOT/.github/labels.conf"
}
documented_panel() {
# shellcheck disable=SC2016 # backticks are literal Markdown delimiters
sed -n '/^2\. \*\*The review panel\*\*/,/required verdicts/p' "$ROOT/CONTRIBUTING.md" \
| grep -oE '`[^`]+-reviewer-andresmgsl`' \
| tr -d '`' \
| paste -sd ' ' -
}
panel_is_forgejo_roster() {
[ "$(configured_panel)" = \
"cluade-reviewer-andresmgsl codex-reviewer-andresmgsl grok-reviewer-andresmgsl kimi-reviewer-andresmgsl" ]
}
panel_rosters_match() {
[ "$(documented_panel)" = "$(configured_panel)" ]
}
# Panel verdicts are all-required, so its roster cannot span disjoint account
# namespaces. Triage authorization is any-match, so the union keeps issue flow
# valid on both GitHub and Forgejo while both boards remain live.
triage_actors_cover_both_forges() {
[ "$(sed -n 's/^triage-actors=//p' "$ROOT/.github/labels.conf")" = \
"dan-claude-bot cluade-reviewer-andresmgsl" ]
}
check "labels: panel names exactly the four Forgejo reviewer accounts" 0 "" \
panel_is_forgejo_roster
check "labels: CONTRIBUTING panel matches labels.conf exactly" 0 "" \
panel_rosters_match
check "labels: triage actors cover GitHub and Forgejo exactly" 0 "" \
triage_actors_cover_both_forges
# One tag governs all eight references — the pin may be bumped, never split.
ceremony_tags() {
grep -rhoE 'heavy-duty/ceremony/[^@]+@[^[:space:]]+' "$ROOT/.github/workflows/" \

View file

@ -50,10 +50,10 @@ trap 'rm -rf "$WORK"' EXIT
# --- the functions under test, extracted -------------------------------------
FNS="$WORK/drill-fns.sh"
for fn in tree_of assert_installed_from classify_leg capture_state emit_record forgejo_run_verdict forgejo_token_verdict forgejo_max_task_id forgejo_leg_row; do
for fn in tree_of assert_installed_from classify_leg capture_state emit_record; do
awk "/^${fn}\(\) \{/,/^\}/" "$ROOT/drill/drill.sh" >> "$FNS"
done
for fn in tree_of assert_installed_from classify_leg capture_state emit_record forgejo_run_verdict forgejo_token_verdict forgejo_max_task_id forgejo_leg_row; do
for fn in tree_of assert_installed_from classify_leg capture_state emit_record; do
check "extraction guards the awk: ${fn}() landed" 0 "${fn}() {" grep -F "${fn}() {" "$FNS"
done
# shellcheck source=/dev/null
@ -201,177 +201,6 @@ check "an all-green record says every leg ran and passed" 0 "Every leg ran and e
# =============================================================================
# the shipped script itself
# =============================================================================
# =============================================================================
# forgejo_run_verdict — did OUR dispatched run land, and how (#129)
# =============================================================================
# The Forgejo half of the runner leg cannot reuse the GitHub reader. Measured
# against forgejo.heavyduty.builders (8.0.3+gitea-1.22.0) on 2026-07-30, a
# completed run in GET /repos/{o}/{r}/actions/tasks carries NO `conclusion`
# field at all — `status` holds the terminal outcome directly, where GitHub
# splits status:completed + conclusion:success. And `id` is a global task id
# (25) while the run's own URL ends in run_number (1), so the pre-dispatch
# guard has to compare `id`.
#
# The payload is also an ASSIGNED-task view: it reads total_count 0 for as
# long as a run sits queued (measured: 200s), so "no new id" is the ONLY
# signal that the runner never took the job. That is the verdict this leg
# exists to produce, which is why it gets its own function and its own tests.
FJ="$WORK/fj"; mkdir -p "$FJ"
printf '%s' '{"workflow_runs":[],"total_count":0}' > "$FJ/empty.json"
printf '%s' '{"workflow_runs":[{"id":25,"status":"success","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-success.json"
printf '%s' '{"workflow_runs":[{"id":25,"status":"failure","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-failure.json"
printf '%s' '{"workflow_runs":[{"id":25,"status":"cancelled","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/new-cancelled.json"
printf '%s' '{"workflow_runs":[{"id":24,"status":"success","run_number":1,"url":"https://f/o/r/actions/runs/1"}],"total_count":1}' > "$FJ/stale-only.json"
check "verdict: an empty task list is PENDING, never a pass" 0 "pending" \
forgejo_run_verdict "" "$FJ/empty.json"
check "verdict: a queued run the runner never took stays PENDING" 0 "pending" \
forgejo_run_verdict "24" "$FJ/stale-only.json"
check "verdict: OUR new run, status success, is SUCCESS" 0 "success" \
forgejo_run_verdict "24" "$FJ/new-success.json"
check "verdict: status carries the outcome — failure is FAILED, not success" 0 "failed" \
forgejo_run_verdict "24" "$FJ/new-failure.json"
check "verdict: a cancelled run is FAILED, not silently passed" 0 "failed" \
forgejo_run_verdict "24" "$FJ/new-cancelled.json"
check "verdict: the first run ever (no pre-id) still resolves" 0 "success" \
forgejo_run_verdict "" "$FJ/new-success.json"
# A task appears in this payload the moment it is ASSIGNED, which can be before
# it finishes — so a non-terminal status must read as pending, not as a failure.
# Calling a still-running job "failed" would make the leg flaky in exactly the
# window the leg is watching.
printf '%s' '{"workflow_runs":[{"id":25,"status":"running","run_number":1}],"total_count":1}' > "$FJ/new-running.json"
check "verdict: an assigned-but-running task is PENDING, not FAILED" 0 "pending" \
forgejo_run_verdict "24" "$FJ/new-running.json"
# grok/kimi on !130: the reader must not stop at the FIRST run. actions/tasks
# accumulates — the moment a repo is drilled twice, our run shares the payload
# with older ones, and nothing documents the sort order. Reading entry[0] makes
# a green job read as a timeout, which is a FALSE FAILURE on the very gate this
# leg exists to provide.
printf '%s' '{"workflow_runs":[{"id":24,"status":"success"},{"id":25,"status":"success"}],"total_count":2}' > "$FJ/oldest-first.json"
printf '%s' '{"workflow_runs":[{"id":25,"status":"success"},{"id":24,"status":"success"}],"total_count":2}' > "$FJ/newest-first.json"
printf '%s' '{"workflow_runs":[{"id":25,"status":"running"},{"id":26,"status":"success"}],"total_count":2}' > "$FJ/ours-not-first.json"
printf '%s' '{"workflow_runs":[{"id":23,"status":"success"},{"id":24,"status":"failure"}],"total_count":2}' > "$FJ/all-stale.json"
# Pretty-printed: the instance may or may not compact its JSON, and a parser
# that silently depends on one-line objects is a latent failure (kimi, !130).
# Written HERE, like every other fixture: a suite that copies from a scratch
# path passes only on the box that built it (grok/kimi, !130 round 2).
printf '%s\n' '{
"workflow_runs": [
{"id": 24, "status": "success"},
{"id": 25, "name": "drill", "status": "success"}
],
"total_count": 2
}' > "$FJ/pretty.json"
check "verdict: ours is LAST in the payload — order must not decide" 0 "success" \
forgejo_run_verdict "24" "$FJ/oldest-first.json"
check "verdict: ours is FIRST in the payload — same answer" 0 "success" \
forgejo_run_verdict "24" "$FJ/newest-first.json"
check "verdict: a stale RUNNING entry ahead of ours does not mask it" 0 "success" \
forgejo_run_verdict "24" "$FJ/ours-not-first.json"
check "verdict: every entry at or below pre is stale — PENDING" 0 "pending" \
forgejo_run_verdict "24" "$FJ/all-stale.json"
check "verdict: a pretty-printed payload parses too" 0 "success" \
forgejo_run_verdict "24" "$FJ/pretty.json"
# The PRE-DISPATCH snapshot has the same multi-entry hazard as the verdict, and
# getting it wrong is worse: a `head -n1` pre-id on an oldest-first payload
# names an OLD run as the baseline, so a later poll that finds the same body
# reports the PREVIOUS drill's run as ours — a false PASS on the take-a-job
# assertion, where the entry[0] bug only produced a false failure (grok, !130).
# Both sides must fold max over every id, which is why they share one function.
check "max id: oldest-first payload yields the NEWEST id, not the first" 0 "25" \
forgejo_max_task_id "$FJ/oldest-first.json"
check "max id: newest-first payload yields the same answer" 0 "25" \
forgejo_max_task_id "$FJ/newest-first.json"
check "max id: an empty payload has no id at all" 0 "" \
forgejo_max_task_id "$FJ/empty.json"
check "max id: a pretty-printed payload folds too" 0 "25" \
forgejo_max_task_id "$FJ/pretty.json"
# The false PASS, pinned end to end: snapshot the oldest-first body, dispatch,
# the runner never takes it so the body is unchanged — the verdict must stay
# pending. With head -n1 this returned success.
# The two halves composed exactly as the leg composes them.
verdict_after_no_new_run() { forgejo_run_verdict "$(forgejo_max_task_id "$1")" "$1"; }
check "no new run after dispatch: max-id baseline keeps it PENDING (false-PASS guard)" 0 "pending" \
verdict_after_no_new_run "$FJ/oldest-first.json"
check "…and the same composition on a pretty payload" 0 "pending" \
verdict_after_no_new_run "$FJ/pretty.json"
# =============================================================================
# forgejo_leg_row — the row is the WHOLE lifecycle, not just the job
# =============================================================================
# codex/grok/kimi on !130: keying the record row on the take-a-job outcome alone
# lets it read "PASS — registered, took a job, removed" when install failed, so
# long as SOMETHING answered runs-on: drill. That is not contrived — this leg
# removes locally and tells the operator to delete the stale runner by hand, so
# a leftover drill-labeled runner from the previous drill is the DESIGNED-FOR
# aftermath, and it answers the fixture exactly.
#
# drills/<v>.md is the release's durable evidence. A row claiming a lifecycle
# that did not happen is precisely what the gate exists to refuse, so PASS
# requires every assertion, not just the interesting one.
check "leg row: everything succeeded is the only PASS" 0 "PASS" \
forgejo_leg_row 1 1 success 1 1
check "leg row: install failed cannot PASS, even when a foreign runner took the job" 0 "FAIL" \
forgejo_leg_row 0 1 success 1 1
check "leg row: status failed cannot PASS either" 0 "FAIL" \
forgejo_leg_row 1 0 success 1 1
check "leg row: remove failed cannot PASS" 0 "FAIL" \
forgejo_leg_row 1 1 success 0 1
check "leg row: a runner still registered after remove cannot PASS" 0 "FAIL" \
forgejo_leg_row 1 1 success 1 0
check "leg row: no dispatch attempted, everything else clean, is PARTIAL" 0 "PARTIAL" \
forgejo_leg_row 1 1 none 1 1
check "leg row: a job that was never taken is a FAIL" 0 "FAIL" \
forgejo_leg_row 1 1 timeout 1 1
check "leg row: PARTIAL requires a clean lifecycle too" 0 "FAIL" \
forgejo_leg_row 0 1 none 1 1
# The end-to-end shape codex/grok/kimi asked for, composed the way the leg
# composes it: a tasks payload carrying a NEWER successful run (as a foreign
# drill-labeled runner would produce) must still not yield a PASS row when the
# drill's own install failed. This is the exact false-evidence case.
row_after_failed_install() {
forgejo_leg_row 0 1 "$(forgejo_run_verdict "$(forgejo_max_task_id "$1")" "$1")" 1 1
}
printf '%s' '{"workflow_runs":[{"id":24,"status":"success"},{"id":99,"status":"success"}],"total_count":2}' \
> "$FJ/foreign-runner-took-it.json"
check "install failed + a newer successful run in the payload is still FAIL, never PASS" 0 "FAIL" \
row_after_failed_install "$FJ/foreign-runner-took-it.json"
# …and the same payload with a clean lifecycle is the PASS, so the check above
# is discriminating rather than always-FAIL.
row_after_clean_install() {
forgejo_leg_row 1 1 "$(forgejo_run_verdict "" "$1")" 1 1
}
check "…while the same payload with a clean lifecycle does PASS" 0 "PASS" \
row_after_clean_install "$FJ/foreign-runner-took-it.json"
# =============================================================================
# forgejo_token_verdict — a configured leg that cannot mint must FAIL, not SKIP
# =============================================================================
# #129's own acceptance: "Token source present but the instance is unreachable
# -> the leg FAILS; it must not skip and must not pass". A mint that returns
# nothing because the instance is unreachable, the token is under-scoped or the
# repo name is wrong is a CONFIGURED leg failing — reporting "no token source"
# sends the operator to check an env var they already set, and writes SKIPPED
# where the record owes a FAIL. That is the UNREADABLE-vs-NONE shape
# drills/README.md names.
check "token: a resolved registration token is ok" 0 "ok" \
forgejo_token_verdict "reg-tok" ""
check "token: an explicit token wins even with no API token" 0 "ok" \
forgejo_token_verdict "reg-tok" ""
check "token: no token at all and no API token is a genuine SKIP" 0 "no-source" \
forgejo_token_verdict "" ""
check "token: API token offered but mint produced nothing is a FAILURE" 0 "mint-failed" \
forgejo_token_verdict "" "api-tok"
# The anti-false-positive guard, stated as its own case: an OLD completed run
# with the SAME id as pre_id must never be read as this dispatch's result.
check "verdict: a pre-existing success with the pre-id is NOT our run" 0 "pending" \
forgejo_run_verdict "25" "$FJ/new-success.json"
# Arg refusals fire before the root check (repo doctrine, bootstrap.sh:114),
# which is what makes them provable here without a throwaway machine.
check "drill.sh refuses to run without BOTH refs pinned (#103)" 2 "--box-ref" \
@ -388,14 +217,6 @@ check "an unknown flag dies loudly, exit 2" 2 "unknown option" \
bash "$ROOT/drill/drill.sh" --frobnicate
check "--help prints the header and exits 0" 0 "THROWAWAY" \
bash "$ROOT/drill/drill.sh" --help
check "--forgejo-instance is a known flag (the leg's opt-in)" 2 "--users <path> is required" \
bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-instance https://f.example.com --yes
check "--forgejo-runner-repo is a known flag" 2 "--users <path> is required" \
bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-runner-repo o/r --yes
check "--help names the forgejo runner leg's flags" 0 "--forgejo-instance" \
bash "$ROOT/drill/drill.sh" --help
check "--forgejo-ref is a known flag (Forgejo's dispatch needs a ref)" 2 "--users <path> is required" \
bash "$ROOT/drill/drill.sh" --rig-ref r --box-ref b --forgejo-ref dev --yes
echo "---"
echo "$PASS passed, $FAIL failed"