2026-07-11 08:25:48 +00:00
|
|
|
# rig
|
2026-07-10 20:37:08 +00:00
|
|
|
|
2026-07-10 20:56:05 +00:00
|
|
|
A CLI that turns a **pristine Debian server into a hardened, tailnet-joined
|
|
|
|
|
node** — one curl, one command. A second command installs a version-pinned
|
|
|
|
|
Coolify on a control-plane box.
|
|
|
|
|
|
|
|
|
|
Philosophy (shared with [claudebox](https://github.com/heavy-duty/claudebox)):
|
2026-07-11 08:25:48 +00:00
|
|
|
**public tool, private state**. rig carries plumbing logic only — no
|
2026-07-10 20:56:05 +00:00
|
|
|
hostnames, no bindings, no secrets, nothing about *your* infrastructure. It
|
|
|
|
|
takes arguments, does its work, and stores no credential, ever.
|
|
|
|
|
|
|
|
|
|
## Install
|
|
|
|
|
|
|
|
|
|
```sh
|
2026-07-11 08:25:48 +00:00
|
|
|
curl -fsSL https://raw.githubusercontent.com/heavy-duty/rig/main/install.sh | bash
|
2026-07-10 20:56:05 +00:00
|
|
|
```
|
|
|
|
|
|
2026-07-11 08:25:48 +00:00
|
|
|
Installs the tree to `~/.local/share/rig` and links `rig` onto your
|
2026-07-10 20:56:05 +00:00
|
|
|
PATH (`/usr/local/bin` when root). Re-run any time to upgrade.
|
|
|
|
|
|
|
|
|
|
## Commands
|
|
|
|
|
|
2026-07-17 15:51:58 +00:00
|
|
|
### `rig bootstrap <control-plane|workload|runner|staging>`
|
2026-07-10 20:56:05 +00:00
|
|
|
|
|
|
|
|
Run as root on the fresh box (over SSH). Convergent — safe to re-run; a
|
|
|
|
|
second run changes nothing.
|
|
|
|
|
|
|
|
|
|
```sh
|
2026-07-11 08:25:48 +00:00
|
|
|
rig bootstrap control-plane --hostname my-coolify-box
|
|
|
|
|
rig bootstrap workload --hostname my-prod-box
|
2026-07-11 18:25:47 +00:00
|
|
|
rig bootstrap runner --hostname my-ci-box
|
2026-07-17 15:51:58 +00:00
|
|
|
rig bootstrap staging --hostname my-vm-host
|
2026-07-10 20:56:05 +00:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
- `--hostname <name>` — tailnet hostname (default: the role name)
|
bootstrap: infer the tailnet tag from the pre-auth key, verify the granted tag
rig used to pass --ts-tag to `tailscale up --advertise-tags`, stating the
tailnet tag a second time with no way to know whether its request and the
key's own tags agreed. It asserted the tag it REQUESTED, never the tag control
GRANTED — the sshd first-wins bug in a different hat, and the same scar (both
M900s joined tag:server, retagged by hand, unnoticed).
Collapse the two sources of truth onto one: the key.
- `tailscale up` drops --advertise-tags; the key's tags apply.
- After join, poll `tailscale status --json` for `.Self.Tags` (netmap ground
truth, not `debug prefs`) until tags appear or BackendState=Running, on BOTH
the fresh-join and already-joined paths.
- UNTAGGED -> hard refusal: `tailscale logout` to back the user-owned node out,
then die naming the fix (mint a tagged key).
- Role policy moves onto the effective tag: a runner must not have tag:server
among the tags the key actually granted. Strictly stronger than before.
- --ts-tag is removed, and dies exit 2 with a message pointing at the key
(consuming its value), not an "unknown flag".
- New array-aware reader json_string_array in lib/runner-config.sh (jq-free,
never fails under set -e), with its own unit tests; bootstrap sources the lib.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:27:09 +00:00
|
|
|
|
|
|
|
|
There is **no `--ts-tag` flag**. A pre-auth key is minted *with* its tags, so
|
|
|
|
|
the key is the single source of truth for the tailnet tag — rig no longer states
|
|
|
|
|
a second one it might disagree with. It **verifies** the tag control actually
|
|
|
|
|
granted after join instead (see below). Passing `--ts-tag` now exits 2 with a
|
|
|
|
|
message pointing you at the key.
|
2026-07-10 20:56:05 +00:00
|
|
|
|
|
|
|
|
What it does: installs `curl ca-certificates unattended-upgrades` (and
|
|
|
|
|
enables periodic unattended upgrades); writes an sshd hardening drop-in
|
2026-07-12 15:26:34 +00:00
|
|
|
(`PermitRootLogin prohibit-password`, `PasswordAuthentication no`) and
|
|
|
|
|
**verifies it took effect** via `sshd -T`; sets the system hostname; installs
|
bootstrap: infer the tailnet tag from the pre-auth key, verify the granted tag
rig used to pass --ts-tag to `tailscale up --advertise-tags`, stating the
tailnet tag a second time with no way to know whether its request and the
key's own tags agreed. It asserted the tag it REQUESTED, never the tag control
GRANTED — the sshd first-wins bug in a different hat, and the same scar (both
M900s joined tag:server, retagged by hand, unnoticed).
Collapse the two sources of truth onto one: the key.
- `tailscale up` drops --advertise-tags; the key's tags apply.
- After join, poll `tailscale status --json` for `.Self.Tags` (netmap ground
truth, not `debug prefs`) until tags appear or BackendState=Running, on BOTH
the fresh-join and already-joined paths.
- UNTAGGED -> hard refusal: `tailscale logout` to back the user-owned node out,
then die naming the fix (mint a tagged key).
- Role policy moves onto the effective tag: a runner must not have tag:server
among the tags the key actually granted. Strictly stronger than before.
- --ts-tag is removed, and dies exit 2 with a message pointing at the key
(consuming its value), not an "unknown flag".
- New array-aware reader json_string_array in lib/runner-config.sh (jq-free,
never fails under set -e), with its own unit tests; bootstrap sources the lib.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:27:09 +00:00
|
|
|
tailscale and joins your tailnet — then **verifies the tag the key granted**
|
|
|
|
|
(see *The tag comes from the key* below).
|
2026-07-10 20:56:05 +00:00
|
|
|
|
2026-07-12 15:47:38 +00:00
|
|
|
**`--hostname` converges both names.** On a box that has already joined,
|
|
|
|
|
`bootstrap` skips `tailscale up` (so a re-run needs no pre-auth key) — but it
|
|
|
|
|
still reconciles the **tailnet** hostname via `tailscale set --hostname`. Without
|
|
|
|
|
that, a box which joined under the wrong name — say `--hostname` was omitted, so
|
|
|
|
|
it defaulted to the *role* — stayed misnamed forever, and re-running rig, the
|
|
|
|
|
documented repair, could not fix it. A machine you deliberately renamed in the
|
|
|
|
|
admin console keeps that name; rig will not fight it.
|
|
|
|
|
|
2026-07-12 15:26:34 +00:00
|
|
|
> **Why the drop-in is `00-rig.conf` and not `99-`.** `sshd_config` is
|
|
|
|
|
> **first-wins** — *"for each keyword, the first obtained value will be used"*
|
|
|
|
|
> (`sshd_config(5)`) — and `Include` expands its glob in lexical order. Cloud
|
|
|
|
|
> images ship `/etc/ssh/sshd_config.d/50-cloud-init.conf` carrying
|
|
|
|
|
> `PasswordAuthentication yes`, so a `99-` drop-in is read **second** and every
|
|
|
|
|
> keyword in it is silently discarded. This is the opposite of the
|
|
|
|
|
> last-wins convention most config systems use, and it shipped green here for
|
|
|
|
|
> a month: rig asserted the *file existed* rather than what `sshd` actually
|
|
|
|
|
> resolved, and the Incus rehearsal container has no cloud-init drop-in to
|
|
|
|
|
> lose to. Every Hetzner box rig had bootstrapped was still serving
|
|
|
|
|
> `passwordauthentication yes`. `bootstrap` now sweeps a stale `99-rig.conf`
|
|
|
|
|
> on re-run, and refuses to claim success unless `sshd -T` agrees.
|
|
|
|
|
|
2026-07-10 20:56:05 +00:00
|
|
|
**The pre-auth key:** provide it via the `TS_AUTHKEY` env var or type it at
|
bootstrap: infer the tailnet tag from the pre-auth key, verify the granted tag
rig used to pass --ts-tag to `tailscale up --advertise-tags`, stating the
tailnet tag a second time with no way to know whether its request and the
key's own tags agreed. It asserted the tag it REQUESTED, never the tag control
GRANTED — the sshd first-wins bug in a different hat, and the same scar (both
M900s joined tag:server, retagged by hand, unnoticed).
Collapse the two sources of truth onto one: the key.
- `tailscale up` drops --advertise-tags; the key's tags apply.
- After join, poll `tailscale status --json` for `.Self.Tags` (netmap ground
truth, not `debug prefs`) until tags appear or BackendState=Running, on BOTH
the fresh-join and already-joined paths.
- UNTAGGED -> hard refusal: `tailscale logout` to back the user-owned node out,
then die naming the fix (mint a tagged key).
- Role policy moves onto the effective tag: a runner must not have tag:server
among the tags the key actually granted. Strictly stronger than before.
- --ts-tag is removed, and dies exit 2 with a message pointing at the key
(consuming its value), not an "unknown flag".
- New array-aware reader json_string_array in lib/runner-config.sh (jq-free,
never fails under set -e), with its own unit tests; bootstrap sources the lib.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:27:09 +00:00
|
|
|
the interactive prompt. Use a **single-use, tagged, short-expiry** key — the
|
|
|
|
|
**tagged** part is now load-bearing, not advice (see below). It lives in process
|
|
|
|
|
memory only — rig never writes a credential to disk.
|
|
|
|
|
|
|
|
|
|
**The tag comes from the key, and rig verifies the one control granted.** rig
|
|
|
|
|
used to pass `--ts-tag` to `tailscale up --advertise-tags`, stating the tag a
|
|
|
|
|
*second* time — with no way to know whether its request and the key's own tags
|
|
|
|
|
agreed. It asserted the tag it **requested**, never the tag control **granted**;
|
|
|
|
|
this is the same shape as the sshd first-wins bug above, and it left the same
|
|
|
|
|
scar (both M900s joined carrying `tag:server` and had to be retagged by hand,
|
|
|
|
|
because nothing in rig ever read the effective tag back). So rig stops
|
|
|
|
|
overriding the key: `tailscale up` carries no `--advertise-tags`, the key's tags
|
|
|
|
|
apply, and after join rig polls `tailscale status --json` for `.Self.Tags` — the
|
|
|
|
|
netmap's ground truth, not `tailscale debug prefs`, which prints what was
|
|
|
|
|
*requested* — and asserts on that, on **first join and on every re-run** (which
|
|
|
|
|
catches a box bootstrapped before this change, or retagged behind rig's back).
|
|
|
|
|
|
|
|
|
|
> **An untagged key is a hard refusal.** Drop `--advertise-tags` and you also
|
|
|
|
|
> drop the accidental net that used to tag an untagged key's node anyway. An
|
|
|
|
|
> untagged node joins owned by the *key creator's user identity* — it inherits
|
|
|
|
|
> that human's ACL grants, expires with the key, and vanishes if the account is
|
|
|
|
|
> deleted. That is a fleet-shaped mistake, not a warning: rig runs `tailscale
|
|
|
|
|
> logout` to back the half-joined node out and dies telling you to mint a tagged
|
|
|
|
|
> key. A wrong tag **cannot** be fixed in place either — `tailscale set` has no
|
|
|
|
|
> tag flag, re-tagging needs a fresh key via `up --force-reauth` — so rig detects
|
|
|
|
|
> and refuses, and never claims a convergence it cannot perform.
|
2026-07-10 20:56:05 +00:00
|
|
|
|
2026-07-11 18:25:47 +00:00
|
|
|
`control-plane` and `workload` are identical today except the default
|
|
|
|
|
hostname; they exist because the boxes diverge over time, and because each
|
bootstrap: infer the tailnet tag from the pre-auth key, verify the granted tag
rig used to pass --ts-tag to `tailscale up --advertise-tags`, stating the
tailnet tag a second time with no way to know whether its request and the
key's own tags agreed. It asserted the tag it REQUESTED, never the tag control
GRANTED — the sshd first-wins bug in a different hat, and the same scar (both
M900s joined tag:server, retagged by hand, unnoticed).
Collapse the two sources of truth onto one: the key.
- `tailscale up` drops --advertise-tags; the key's tags apply.
- After join, poll `tailscale status --json` for `.Self.Tags` (netmap ground
truth, not `debug prefs`) until tags appear or BackendState=Running, on BOTH
the fresh-join and already-joined paths.
- UNTAGGED -> hard refusal: `tailscale logout` to back the user-owned node out,
then die naming the fix (mint a tagged key).
- Role policy moves onto the effective tag: a runner must not have tag:server
among the tags the key actually granted. Strictly stronger than before.
- --ts-tag is removed, and dies exit 2 with a message pointing at the key
(consuming its value), not an "unknown flag".
- New array-aware reader json_string_array in lib/runner-config.sh (jq-free,
never fails under set -e), with its own unit tests; bootstrap sources the lib.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 15:27:09 +00:00
|
|
|
follow-up command applies to exactly one role. `runner` is the box a CI agent
|
|
|
|
|
will live on, and it differs behaviorally: it **refuses `tag:server`**. That
|
|
|
|
|
refusal moved onto the *effective* tag and is strictly stronger for it — it is
|
|
|
|
|
no longer "don't advertise `tag:server`" but "the key you actually used must not
|
|
|
|
|
grant `tag:server` to repo-controlled code." A runner executes that code, and
|
|
|
|
|
`tag:server`'s grants (SSH between your servers, say) must never extend to it;
|
|
|
|
|
the check turns the worst misconfiguration from a documentation warning into a
|
|
|
|
|
hard, post-join error.
|
2026-07-10 20:56:05 +00:00
|
|
|
|
2026-07-17 15:51:58 +00:00
|
|
|
`staging` is the box that *hosts* staging boxes — Incus VMs minted by the
|
|
|
|
|
[`box`](https://github.com/heavy-duty/box) CLI, each converged from inside with
|
|
|
|
|
`rig bootstrap workload` and registered in the control plane as its own server.
|
|
|
|
|
Mint its key with `tag:local`: the host and its guests sit on opposite sides of
|
|
|
|
|
a trust boundary, and the *host* is never managed by the control plane — so the
|
|
|
|
|
role **refuses an effective `tag:server`**, same mechanism as `runner`. rig
|
|
|
|
|
deliberately installs no Incus and no box here — box's own `setup-host` is the
|
|
|
|
|
single owner of the Incus daemon's configuration, and two tools converging one
|
|
|
|
|
daemon is drift by construction. The closing log points you at it: install box,
|
|
|
|
|
run `box setup-host`, then `box new --template staging`. If `/dev/kvm` is
|
|
|
|
|
absent, rig warns (a host that exists to run VMs should have it) but does not
|
|
|
|
|
fail — the role is rehearsed in containers, which legitimately lack it.
|
|
|
|
|
|
2026-07-11 08:25:48 +00:00
|
|
|
### `rig coolify install --version <pin>`
|
2026-07-10 20:56:05 +00:00
|
|
|
|
|
|
|
|
Control-plane box only. Installs Coolify at exactly the pinned version with
|
|
|
|
|
`AUTOUPDATE=false` — your deploy tooling is verified against an API surface;
|
|
|
|
|
the platform must never move underneath it on its own. Upgrading is an
|
|
|
|
|
explicit re-run with a new pin. The pin is required; there is no default.
|
|
|
|
|
|
feat(coolify): install the control-plane dump as a systemd timer
The Coolify control-plane database holds the GitHub App private key, every
registered server's SSH key, and every environment value for every environment
it manages. Backing it up was a manual runbook step, and the dump script lived
in cast — the off-box tool, whose src never references it. It runs on the box,
as root, under a scheduler: that is rig's job description.
It matters beyond tidiness. The dump is forensics, not a restore path — a lost
control plane is rebuilt fresh and reconciled from the manifest. So there will
be a next control-plane box, and as a runbook step it was born un-backed-up,
depending on someone remembering mid-incident. Now it is backed up from birth.
rig installs the machinery and templates /etc/coolify-dump.env empty at 0600,
never reading it back — no credential passes through rig. The script's own
guards make an unfilled file fail the unit loudly rather than ship plaintext.
systemd timer over cron: EnvironmentFile is the right idiom for 0600 secrets,
failures surface in systemctl status instead of being mailed into the void, and
Persistent=true catches a run missed while the box was down.
Two hazards the cast script missed, carried into the unit:
- aws-cli >= 2.23 enables default upload checksums that S3-compatible backends
reject; Debian 13 ships 2.23.6, so the unit defaults both checksum knobs to
when_required.
- A failed pg_dump piped into age still yields a valid, tiny, encrypted file
that uploads cleanly every night and looks exactly like a working backup. The
script now refuses to upload an empty artifact.
Closes #8
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-12 19:14:07 +00:00
|
|
|
### `rig coolify backup install`
|
|
|
|
|
|
|
|
|
|
Control-plane box only. Installs a **nightly age-encrypted dump of Coolify's own
|
|
|
|
|
database** as a systemd timer.
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
rig coolify backup install
|
|
|
|
|
rig coolify backup install --schedule '*-*-* 02:30:00 UTC' --pg-container coolify-db
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
- `--schedule <OnCalendar>` — systemd calendar expression (default: `*-*-* 04:00:00 UTC`)
|
|
|
|
|
- `--pg-container` / `--pg-user` / `--pg-db` — Coolify's postgres (defaults: `coolify-db`,
|
|
|
|
|
`coolify`, `coolify`)
|
|
|
|
|
|
|
|
|
|
That database holds the GitHub App private key, every registered server's SSH key,
|
|
|
|
|
and every environment value for every environment the control plane manages. It is
|
|
|
|
|
`pg_dump`ed straight into `age` — encrypted **client-side, on the box** — and only
|
|
|
|
|
then shipped to S3. The bucket is never trusted with plaintext.
|
|
|
|
|
|
|
|
|
|
It is **forensics, not a restore path.** A lost control plane is rebuilt fresh and
|
|
|
|
|
reconciled from your manifest, never restored from this artifact. Which is exactly
|
|
|
|
|
why the plumbing belongs in rig: there *will* be a next control-plane box, and it
|
|
|
|
|
should be backed up from birth rather than depending on someone remembering a
|
|
|
|
|
runbook step mid-incident.
|
|
|
|
|
|
|
|
|
|
**rig installs the machinery; you supply the bindings.** rig writes
|
|
|
|
|
`/etc/coolify-dump.env` **empty**, `0600`, and never reads it back — no credential
|
|
|
|
|
ever passes through rig. You fill in the age recipient (a *public* key), the S3
|
|
|
|
|
bucket + endpoint, and the S3 credentials. Until you do, the unit **fails loudly on
|
|
|
|
|
every run**: a silent backup is worse than a missing one.
|
|
|
|
|
|
|
|
|
|
rig cannot verify that the upload works — that needs your credentials. So prove it
|
|
|
|
|
by hand once, rather than letting the timer discover it at 04:00:
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
systemctl start coolify-dump.service
|
|
|
|
|
journalctl -u coolify-dump.service -n 20 --no-pager
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
A backup you have never read back is not yet a backup.
|
|
|
|
|
|
2026-07-11 18:44:43 +00:00
|
|
|
### `rig runner install --repo <owner/repo>`
|
2026-07-11 17:45:38 +00:00
|
|
|
|
2026-07-11 18:25:47 +00:00
|
|
|
Runner box only, run after `rig bootstrap runner` (the same two-step rhythm
|
|
|
|
|
as `bootstrap control-plane` → `coolify install`):
|
2026-07-11 17:45:38 +00:00
|
|
|
|
|
|
|
|
```sh
|
2026-07-11 18:25:47 +00:00
|
|
|
rig bootstrap runner --hostname my-ci-box
|
2026-07-11 18:44:43 +00:00
|
|
|
rig runner install --repo acme/widgets
|
2026-07-11 17:45:38 +00:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Installs GitHub's official `actions/runner` as a systemd service under an
|
|
|
|
|
unprivileged user (default `github-runner`, created if absent, never root, no
|
|
|
|
|
supplementary groups). The runner is an agent, not a server: it long-polls
|
|
|
|
|
GitHub outbound and receives jobs down that already-established connection,
|
|
|
|
|
so it needs **zero inbound ports** and works fine behind a deny-all
|
|
|
|
|
firewall — it can even trigger deploys on hosts only it can reach, like a
|
|
|
|
|
tailnet-only control plane.
|
|
|
|
|
|
|
|
|
|
No Docker, deliberately: the Docker socket is a root API and `docker` group
|
|
|
|
|
membership is root-equivalent, which is a gratuitous path to root on a box
|
|
|
|
|
whose whole point is a narrow blast radius. Add Docker only once a job
|
|
|
|
|
genuinely needs it, and rethink the isolation model then.
|
|
|
|
|
|
2026-07-11 18:44:43 +00:00
|
|
|
- `--version <pin>` — actions/runner release to install (default: the
|
|
|
|
|
latest release, resolved at install time; e.g. `--version 2.335.1` —
|
|
|
|
|
the latest as of this writing). Pin it when you need a deterministic,
|
|
|
|
|
auditable install.
|
2026-07-11 17:45:38 +00:00
|
|
|
- `--name <name>` — runner name (default: this host's hostname)
|
2026-07-11 17:51:08 +00:00
|
|
|
- `--labels <csv>` — runner labels, replacing the `ci-runner` default — keep
|
|
|
|
|
any label your workflows' `runs-on` needs (GitHub adds `self-hosted` itself)
|
2026-07-11 17:45:38 +00:00
|
|
|
- `--user <name>` — the unprivileged service user (default: `github-runner`)
|
|
|
|
|
|
|
|
|
|
**The registration token:** provide it via the `RUNNER_TOKEN` env var or type
|
|
|
|
|
it at the interactive prompt. It's short-lived, consumed at registration, and
|
|
|
|
|
never written to disk by rig.
|
|
|
|
|
|
2026-07-11 18:44:43 +00:00
|
|
|
Why latest-by-default here when `coolify install` demands a pin: the two
|
|
|
|
|
tools age differently. Coolify never self-updates (`AUTOUPDATE=false`), so
|
|
|
|
|
its version is a contract your deploy tooling is verified against — stating
|
|
|
|
|
it is the point. The runner **self-updates regardless**: GitHub refuses jobs
|
|
|
|
|
from stale runners, so freezing it would just make it silently stop taking
|
|
|
|
|
work. The install-time version is a starting point either way; `--version`
|
|
|
|
|
exists for when you want that starting point deterministic and auditable.
|
2026-07-11 17:45:38 +00:00
|
|
|
|
fix(runner): install refuses a box registered to another repo
`rig runner install --repo <B>` on a box already registered to repo A
treated the mere existence of .runner as "already registered", skipped
configure, restarted the service still pointed at A, and reported success.
--repo was accepted, validated, and then ignored — leaving B with zero
runners and its `runs-on` jobs queued against one that will never come.
This is the natural next command after a partial `repoint`, and the failure
is worse than a no-op: moving a runner between repos is a trust-boundary
act, so quietly putting it back on the old one defeats the point of the move.
Gate install on the repo .runner actually names. Convergence — the property
worth keeping — is untouched: re-running against the repo the box is already
on still skips registration, never prompts for a token, and exits 0.
Skipping when the repo *differs* was never convergence, only a silently
ignored argument, so it now fails and names both repos, pointing at
`runner repoint` (move) or `runner remove` (start over). An unreadable
.runner is refused too — it is no licence to assume a match.
The .runner reader that `status` and `repoint` each carried is lifted into
commands/lib/runner-config.sh, which now also holds the guard. Its json_field
no longer dies bare under `set -o pipefail` when a key is missing, which is
what `status`'s own ${REPO_URL:-unknown} fallback always assumed.
Tests: the guard is exercised against a fixture .runner (refuses another repo
naming both, points at repoint, no-ops on the same repo, passes an
unregistered box, refuses an unreadable one) plus an ordering assertion that
it precedes svc.sh start — reaching it through the CLI would need root and a
really-registered runner, which the dependency-free harness cannot fabricate.
All three mutants (guard deleted, guard comparing nothing, guard moved below
the service start) go red.
Closes #13
2026-07-13 14:57:28 +00:00
|
|
|
Convergent **toward `--repo`** — re-running against the repo this box is
|
|
|
|
|
already on re-uses the binary, skips registration, and never asks for a token.
|
|
|
|
|
Pointed at a *different* repo it **refuses**, and names both: skipping there
|
|
|
|
|
would not be convergence, it would be ignoring the argument — restarting the
|
|
|
|
|
runner on the **old** repo while reporting success, leaving the repo you asked
|
|
|
|
|
for with no runner and its `runs-on` jobs queued forever. Moving a runner
|
|
|
|
|
between repos is a trust-boundary act; that verb is
|
|
|
|
|
[`rig runner repoint`](#rig-runner-repoint---repo-ownerrepo).
|
2026-07-11 17:45:38 +00:00
|
|
|
|
2026-07-13 13:25:27 +00:00
|
|
|
### `rig runner status`
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
rig runner status
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
What this box's runner is registered to — repo, runner name, labels,
|
|
|
|
|
install dir, systemd unit and its state. Reads the runner's own on-disk
|
|
|
|
|
config; no token, no network call. Exits 1 when no runner is installed.
|
|
|
|
|
|
|
|
|
|
The answer to "wait, which repo is this box wired to?" should not require
|
|
|
|
|
knowing that the config lives in a dotfile under an unprivileged user's home.
|
|
|
|
|
|
|
|
|
|
### `rig runner remove`
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
rig runner remove
|
|
|
|
|
rig runner remove --local # no token; leaves a stale entry to delete by hand
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Stops and uninstalls the systemd service, then deregisters the runner from
|
|
|
|
|
GitHub. The binary and its user stay put, so a later `rig runner install`
|
|
|
|
|
re-registers without downloading anything.
|
|
|
|
|
|
|
|
|
|
**The token here is a *removal* token, not a registration token** — a
|
|
|
|
|
different endpoint, and mixing them up is the easy mistake:
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
gh api -X POST repos/<owner/repo>/actions/runners/remove-token
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Supply it via `RUNNER_REMOVE_TOKEN` or the prompt; it never touches disk.
|
|
|
|
|
|
|
|
|
|
`--local` is the escape hatch for when the registration is already gone
|
|
|
|
|
server-side (or you can't mint a token): the box is cleaned, but a stale
|
|
|
|
|
offline runner stays listed in the repo, for you to delete from
|
|
|
|
|
Settings → Actions → Runners.
|
|
|
|
|
|
|
|
|
|
The service always comes down *first*, in both paths. GitHub's own removal
|
|
|
|
|
refuses to run while the service is installed ("Uninstall service first"),
|
|
|
|
|
and `--local` skips that check entirely — which would otherwise leave a
|
|
|
|
|
running service pointed at config that no longer exists.
|
|
|
|
|
|
|
|
|
|
Convergent — a box with no runner installed exits 0.
|
|
|
|
|
|
|
|
|
|
### `rig runner repoint --repo <owner/repo>`
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
rig runner repoint --repo acme/widgets
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Moves an installed runner from one repository to another: deregister,
|
|
|
|
|
re-register, reusing the binary already on the box. It keeps the runner's
|
|
|
|
|
existing name unless you pass `--name`.
|
|
|
|
|
|
fix(runner): install refuses a box registered to another repo
`rig runner install --repo <B>` on a box already registered to repo A
treated the mere existence of .runner as "already registered", skipped
configure, restarted the service still pointed at A, and reported success.
--repo was accepted, validated, and then ignored — leaving B with zero
runners and its `runs-on` jobs queued against one that will never come.
This is the natural next command after a partial `repoint`, and the failure
is worse than a no-op: moving a runner between repos is a trust-boundary
act, so quietly putting it back on the old one defeats the point of the move.
Gate install on the repo .runner actually names. Convergence — the property
worth keeping — is untouched: re-running against the repo the box is already
on still skips registration, never prompts for a token, and exits 0.
Skipping when the repo *differs* was never convergence, only a silently
ignored argument, so it now fails and names both repos, pointing at
`runner repoint` (move) or `runner remove` (start over). An unreadable
.runner is refused too — it is no licence to assume a match.
The .runner reader that `status` and `repoint` each carried is lifted into
commands/lib/runner-config.sh, which now also holds the guard. Its json_field
no longer dies bare under `set -o pipefail` when a key is missing, which is
what `status`'s own ${REPO_URL:-unknown} fallback always assumed.
Tests: the guard is exercised against a fixture .runner (refuses another repo
naming both, points at repoint, no-ops on the same repo, passes an
unregistered box, refuses an unreadable one) plus an ordering assertion that
it precedes svc.sh start — reaching it through the CLI would need root and a
really-registered runner, which the dependency-free harness cannot fabricate.
All three mutants (guard deleted, guard comparing nothing, guard moved below
the service start) go red.
Closes #13
2026-07-13 14:57:28 +00:00
|
|
|
This is the verb that was missing. `runner install` can create a runner but
|
|
|
|
|
never move one — pointed at a repo the box is not on, it fails and sends you
|
|
|
|
|
here — and re-pointing a box otherwise meant hand-rolled `config.sh`/`svc.sh`
|
|
|
|
|
incantations against an install path only rig knew.
|
2026-07-13 13:25:27 +00:00
|
|
|
|
|
|
|
|
Two short-lived tokens, each minted from **its own** repo — `RUNNER_REMOVE_TOKEN`
|
|
|
|
|
for the one it's leaving, `RUNNER_TOKEN` for the one it's joining. Both are
|
|
|
|
|
collected **before** anything is torn down: a token you turn out not to have
|
|
|
|
|
should fail while the runner is still registered and working, not halfway
|
|
|
|
|
through the move. If re-registration fails anyway, rig says so plainly and
|
|
|
|
|
prints the exact `runner install` line that finishes the job.
|
|
|
|
|
|
|
|
|
|
> **Labels do not survive a move on their own.** GitHub holds them; the runner
|
|
|
|
|
> does not persist them locally. rig now records what it registered with, so
|
|
|
|
|
> `repoint` and `status` can read it back — but a runner installed before rig
|
|
|
|
|
> did that has nothing to read, and `repoint` falls back to the `ci-runner`
|
|
|
|
|
> default and warns loudly before it touches anything. Labels are what
|
|
|
|
|
> `runs-on` matches, so a silent change there is a workflow that simply stops
|
|
|
|
|
> finding its runner. Pass `--labels` if yours differ.
|
|
|
|
|
|
|
|
|
|
Convergent — repointing to the repo it is already on changes nothing, exits 0,
|
|
|
|
|
and never asks for a token.
|
|
|
|
|
|
2026-07-11 08:25:48 +00:00
|
|
|
## What rig deliberately does NOT do
|
2026-07-10 20:56:05 +00:00
|
|
|
|
|
|
|
|
- **Provider firewalls** — Docker publishes ports past host firewalls, so
|
|
|
|
|
the real boundary is your cloud provider's firewall, configured outside
|
|
|
|
|
this tool.
|
|
|
|
|
- **Fetch your config** — boxes never receive repo credentials. Everything
|
2026-07-11 08:25:48 +00:00
|
|
|
rig needs arrives as arguments or an interactive prompt.
|
2026-07-10 20:56:05 +00:00
|
|
|
- **Manage deployments** — deploy manifests/executors are separate concerns.
|
2026-07-11 08:25:48 +00:00
|
|
|
(Planned: the `apply`/`diff` executor half joins rig as commands that
|
2026-07-10 20:56:05 +00:00
|
|
|
run on operator machines, never on boxes.)
|
|
|
|
|
|
|
|
|
|
## Testing
|
|
|
|
|
|
|
|
|
|
`bash test/cli.sh` (dependency-free assertions) + shellcheck run in CI. The
|
|
|
|
|
end-to-end rehearsal is a throwaway VM/container: pristine Debian → install →
|
|
|
|
|
`bootstrap workload` with a real single-use key → assert the sshd drop-in,
|
|
|
|
|
tailnet join, and a no-op second run → destroy, remove the node from the
|
|
|
|
|
tailnet.
|