diff --git a/CHANGELOG.md b/CHANGELOG.md index cc80cc6..6923726 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,39 @@ on the way to cutting its first release, and this file starts there. ## Unreleased +### Added + +- **`rig platform` — what is this machine, calculated at run time, stored + nowhere** (#64) — rig read no hardware at all; the single exception was + `uname -m` in `runner-install.sh`, used to pick a tarball and then + discarded. So "is this the 32GB one, or the M900?" was answered by logging + in and running `free -h`, `nproc`, `df -h` and `uname -r` by hand, four + commands deep, on a machine you were already unsure about. `rig platform` + prints hostname, OS, kernel, CPU, memory, disk and virtualization, then a + provenance block (which rig, when, and the role marker's traits). It + **computes rather than stores**: specs change without rig doing anything — + RAM added, root disk resized, the unattended-upgrades bootstrap itself + enables patching the kernel — so a stored spec is stale the moment the + machine changes, and refreshing one per run would collide with bootstrap's + "a second run changes nothing" contract. Nothing is written, so nothing can + go stale. The corollary is deliberate: reading only `/proc`, `uname`, + `/etc/os-release`, `df` and `systemd-detect-virt` means it needs no root, + makes no network call, and **runs on a pristine Debian box rig has never + bootstrapped** — useful for deciding what to converge a machine into, not + only for auditing it afterwards. That also makes it the rare rig command + the harness can RUN for real instead of grepping: the tests assert the + actual answer describes the actual test machine. Provenance is read, never + written, and degrades per-file — `/etc/rig/manifest` is #61 and does not + exist yet, so that line reads `not bootstrapped` on every machine today and + nothing else depends on it. Named `platform` and not `status` on purpose: + `users status` and `runner status` cross-check recorded against live state + and print `DRIFT`, and a command that records nothing cannot drift — which + also leaves `rig status` free for the machine-wide roll-up. Known + limitation, stated rather than guessed at: `CPU`/`MEMORY` are read from + `/proc` with no cgroup awareness, and whether an `lxc` guest sees its own + limits or the host's totals depends on whether `lxcfs` is in play — it is + unverified, so those two lines are unreliable there. + ### Changed - **BREAKING: `--class human|server` is now `--root-door closed|open`** (#77) — diff --git a/README.md b/README.md index 806c0ab..9743bda 100644 --- a/README.md +++ b/README.md @@ -719,6 +719,83 @@ default, dumps it, restores into a second container whose superuser differs, and asserts the rows and an ordered checksum survived — the same proof, done against throwaway containers on every push. +### `rig platform` + +```sh +rig platform +``` + +What is this machine — computed at run time, **stored nowhere**: + +``` +PLATFORM +HOSTNAME hetzner-cp-1 +OS Debian GNU/Linux 13 (trixie) +KERNEL 6.12.95+deb13-amd64 (x86_64) +CPU AMD Ryzen 7 3700X 8-Core Processor (16 cores) +MEMORY 31Gi total, 24Gi available +DISK 456Gi total, 201Gi free on / +VIRT kvm + +PROVENANCE +RIG 0.4.0, bootstrapped 2026-07-19T14:24:51Z +ROLE dev (class=human host=yes join=authkey) +``` + +"Is this the 32GB one, or the M900?" was previously a question you answered by +logging in and running `free -h`, `nproc`, `df -h` and `uname -r` by hand — +four commands deep, on a machine you were already unsure about. + +**Why this computes instead of storing.** It would be easy to write the specs +into a file at bootstrap. That is the wrong shape: specs change without rig +doing anything — someone adds RAM, resizes the root disk, or the +unattended-upgrades that bootstrap itself enables patches the kernel. A stored +spec is stale the moment the machine changes, and refreshing it on every run +would collide with bootstrap's contract that a second run changes nothing. +Computing at run time removes the problem instead of managing it: the answer +is correct by construction because there is nothing to go stale. + +The corollary is deliberate: **`rig platform` works on a machine rig has never +converged.** It reads only `/proc`, `uname`, `/etc/os-release`, `df` and +`systemd-detect-virt`, so it runs on bare Debian before bootstrap — useful for +deciding *what to converge this into*, not just for auditing afterwards. It +needs no root, makes no network call, and writes nothing, ever. + +The `PROVENANCE` block is the complementary half — which rig, and when, which +is *decided* rather than observed, so it is stored. It is **read, never +written**: `RIG` comes from `/etc/rig/manifest` and `ROLE` from +`/etc/rig/role`. Neither file is required — a machine missing one reads `not +bootstrapped` for that line, which is itself the useful answer. The manifest +is #61 and is not implemented yet, so today that line reads `not bootstrapped` +on every machine; nothing else in the command depends on it. + +**Known limitation — `CPU` and `MEMORY` inside a container-style guest are +unverified.** `CPU` and `MEMORY` are read straight from `/proc/cpuinfo` and +`/proc/meminfo`, with no cgroup awareness. Inside a box-minted guest (`VIRT` +says `lxc`) it is **not currently established** whether those files report the +instance's configured limits or the host's totals: neither file is namespaced +by the kernel, but `lxcfs` — when the guest has it — overmounts both with +limit-aware versions, so the answer depends on the guest's setup rather than +on anything rig controls. Until someone confirms it against a real guest, +treat those two lines as unreliable on `lxc` machines and check the instance +config if the number matters. Everything else (OS, kernel, disk, virt, +provenance) is the guest's own either way. + +Deliberately not guessed at: cgroup-aware limit detection would be the fix if +the numbers do turn out to be the host's, but writing it against a *reasoned* +answer rather than an *observed* one risks correcting a bug that isn't there +and papering over one that is. + +Deliberately **not** here: NIC names, MAC addresses, PCI inventory, mount +tables, sensors — this is a cheatsheet, not `inxi`, and the bar is "what would +I want to know before I SSH in". Nor any health judgement ("disk nearly +full"): that needs thresholds this command has no business owning. It is +called `platform` and not `status` on purpose — `rig users status` and `rig +runner status` cross-check recorded state against live state and print +`DRIFT`, and this command records nothing, so it cannot drift and must not +borrow a promise it structurally cannot make. That leaves `rig status` free +for the machine-wide roll-up it will eventually want to be. + ### `rig runner install --repo ` Runner box only, run after `rig bootstrap runner-server` (the same two-step rhythm diff --git a/bin/rig b/bin/rig index c352918..a077496 100755 --- a/bin/rig +++ b/bin/rig @@ -54,6 +54,15 @@ commands: writes a gzipped SQL artifact (--no-owner --no-acl, so it restores onto a different instance); `restore` loads one back, connecting as the container's own superuser, behind a confirm gate. Run as root. + platform + What is this machine: hostname, OS, kernel, CPU, memory, disk and + virtualization, then rig's own provenance (which rig, when, and the + role marker bootstrap wrote). Computed at run time from /proc, uname, + /etc/os-release, df and systemd-detect-virt and stored NOWHERE — a + spec goes stale the moment someone adds RAM, so there is nothing to + go stale here. Writes nothing, needs no root, makes no network call, + and therefore also runs on a pristine Debian box rig has never + bootstrapped, where the provenance block reads 'not bootstrapped'. runner install --repo [options] GitHub Actions runner as a systemd service under an unprivileged user — outbound-only, no Docker. Prompts for the short-lived @@ -383,6 +392,10 @@ case "$cmd" in ;; esac ;; + platform) + shift + exec "$ROOT/commands/platform.sh" "$@" + ;; runner) shift sub="${1:-}" diff --git a/commands/platform.sh b/commands/platform.sh new file mode 100755 index 0000000..12ccdaa --- /dev/null +++ b/commands/platform.sh @@ -0,0 +1,173 @@ +#!/usr/bin/env bash +# rig platform — what is this machine? Calculated at run time, stored nowhere. +# +# Read-only in the strongest sense rig has: it reads /proc, uname, +# /etc/os-release, df and systemd-detect-virt, and writes NOTHING, ever. That +# is the design, not an implementation detail — specs change without rig doing +# anything (RAM added, root disk resized, unattended-upgrades patching the +# kernel), so a stored spec is stale the moment the machine changes, and +# refreshing one on every run would collide with bootstrap's convergence +# contract ("safe to re-run; a second run changes nothing"). +# +# The corollary is worth having deliberately: this runs on a machine rig has +# never converged, and needs no root. It answers "what should I converge this +# into?", not only "what did I converge this into?". +set -euo pipefail + +HERE="$(cd "$(dirname "$(readlink -f "${BASH_SOURCE[0]}")")" && pwd)" +# shellcheck source=SCRIPTDIR/lib/users-config.sh +. "$HERE/lib/users-config.sh" # read_role_marker — one reader of /etc/rig/role + +die() { printf 'rig-platform: ERROR: %s\n' "$1" >&2; exit "${2:-1}"; } + +usage() { + cat <<'EOF' +usage: rig platform + +Describes the machine you are on: hostname, OS, kernel, CPU, memory, disk +and virtualization, then rig's own provenance (which rig, when, and the role +marker bootstrap wrote). + +Computed at run time from /proc, uname, /etc/os-release, df and +systemd-detect-virt. Writes nothing, needs no root, makes no network call — +so it also works on a pristine Debian box rig has never bootstrapped, where +the provenance block reads 'not bootstrapped'. +EOF +} + +# --- args ------------------------------------------------------------------ +while [ $# -gt 0 ]; do + case "$1" in + -h|--help) usage; exit 0 ;; + *) die "unknown flag: $1" 2 ;; + esac +done + +# One aligned column for every line, so the output diffs cleanly across a +# fleet and reads as one table rather than a log. +field() { printf '%-10s %s\n' "$1" "$2"; } + +# --- hostname --------------------------------------------------------------- +# uname -n is the coreutils fallback: `hostname` lives in its own package and a +# minimal image may not carry it, and this command's whole point is running +# before anything has been installed. +HOSTNAME_V="$(hostname 2>/dev/null || uname -n)" + +# --- OS --------------------------------------------------------------------- +# THE os-release TRAP: /etc/os-release defines VERSION, NAME and ID, so +# sourcing it in the MAIN shell silently clobbers same-named script variables. +# Every site in this tree sources it in a SUBSHELL instead (bootstrap.sh:305, +# bootstrap-tenant.sh:126, runner-install.sh:88, db.sh:52, +# coolify-backup-install.sh:88), and test/cli.sh greps commands/ to keep it +# that way. Follow the form verbatim. +if [ -r /etc/os-release ]; then + OS="$(. /etc/os-release && printf '%s %s' "${NAME:-}" "${VERSION:-${VERSION_ID:-}}")" +else + OS="" +fi + +# --- kernel ----------------------------------------------------------------- +KERNEL="$(uname -r) ($(uname -m))" + +# --- CPU -------------------------------------------------------------------- +# 'model name' is x86's spelling; arm64 /proc/cpuinfo has no such field, so an +# unnamed CPU still reports its core count rather than nothing at all. +CPU_MODEL="$(awk -F': ' '/^model name/ {print $2; exit}' /proc/cpuinfo 2>/dev/null || true)" +CORES="$(nproc 2>/dev/null || true)" + +# --- memory ----------------------------------------------------------------- +# /proc/meminfo is in kB. MemAvailable is the kernel's own estimate of what a +# new workload could claim (MemFree undercounts badly, reclaimable cache being +# most of a busy box's RAM); it predates every kernel rig targets, but degrade +# rather than print a wrong number if it is missing. +mem_kb() { awk -v k="$1" '$1 == k":" {print $2; exit}' /proc/meminfo 2>/dev/null || true; } +MEM_TOTAL_KB="$(mem_kb MemTotal)" +MEM_AVAIL_KB="$(mem_kb MemAvailable)" +human_kb() { # kB -> IEC, matching df's units below + [ -n "${1:-}" ] || { printf 'unknown'; return 0; } + numfmt --to=iec-i "$(( $1 * 1024 ))" 2>/dev/null || printf '%s kB' "$1" +} + +# --- disk ------------------------------------------------------------------- +# -P is the one-line-per-filesystem guarantee (a long device name otherwise +# wraps and breaks field positions); -B1 gives bytes, so numfmt renders the +# same IEC units as memory above instead of df's own bare 'G'. +DISK_TOTAL="" DISK_FREE="" +if DF="$(df -PB1 / 2>/dev/null)"; then + DISK_TOTAL="$(printf '%s\n' "$DF" | awk 'NR==2 {print $2}')" + DISK_FREE="$(printf '%s\n' "$DF" | awk 'NR==2 {print $4}')" +fi +human_b() { [ -n "${1:-}" ] && numfmt --to=iec-i "$1" 2>/dev/null || printf 'unknown'; } + +# --- virtualization --------------------------------------------------------- +# THE set -e TRAP: systemd-detect-virt exits NON-ZERO on bare metal while +# printing 'none'. That is a normal, correct answer — without the `|| true` a +# bare-metal machine would turn this whole command into a failed run. The +# substitution also swallows the binary being absent entirely (a non-systemd +# box), which lands as 'unknown'. +VIRT="$(systemd-detect-virt 2>/dev/null || true)" + +printf '%s\n' "PLATFORM" +field HOSTNAME "$HOSTNAME_V" +field OS "${OS:-unknown}" +field KERNEL "$KERNEL" +field CPU "${CPU_MODEL:-unknown}${CORES:+ ($CORES cores)}" +field MEMORY "$(human_kb "$MEM_TOTAL_KB") total, $(human_kb "$MEM_AVAIL_KB") available" +field DISK "$(human_b "$DISK_TOTAL") total, $(human_b "$DISK_FREE") free on /" +field VIRT "${VIRT:-unknown}" +echo + +# --- provenance: READ, never written ---------------------------------------- +# The complementary half of the answer — which rig, and when — is decided +# rather than observed, so unlike everything above it IS stored. rig writes it +# during bootstrap; this command only ever reads it. +# +# /etc/rig/manifest is #61 and is NOT implemented yet, so on every machine in +# existence today this block reads 'not bootstrapped'. That degradation is the +# point: the two features are independent and neither blocks the other. The +# parse is the flat key=value shape the manifest is specified to use — the +# same jq-free shape /etc/rig/users and /etc/rig/role already use, parseable +# with `read` on a box that has no YAML parser. +# +# RIG_MANIFEST / RIG_ROLE_MARKER override the paths so the harness can drive +# both the present and the absent case against fixtures, non-root, without a +# real marker on the machine running the tests (repo precedent: the +# RIG_ROLE_MARKER gate in bin/rig, install.sh and users-close-root.sh). +MANIFEST="${RIG_MANIFEST:-/etc/rig/manifest}" +MARKER="${RIG_ROLE_MARKER:-/etc/rig/role}" + +manifest_field() { # $1 = key — empty when absent, unreadable or unset + local k v + [ -r "$MANIFEST" ] || return 0 + while IFS='=' read -r k v; do + [ "$k" = "$1" ] || continue + printf '%s\n' "$v" + return 0 + done < "$MANIFEST" + return 0 +} + +printf '%s\n' "PROVENANCE" +if [ -r "$MANIFEST" ]; then + RIG_VER="$(manifest_field version)" + RIG_WHEN="$(manifest_field bootstrapped)" + field RIG "${RIG_VER:-unknown}${RIG_WHEN:+, bootstrapped $RIG_WHEN}" +else + field RIG "not bootstrapped (no $MANIFEST)" +fi + +# The role marker is bootstrap's own line — 'role=dev class=human host=yes +# join=authkey' — printed as the role plus its traits. +MARKER_LINE="$(read_role_marker "$MARKER")" +if [ -n "$MARKER_LINE" ]; then + ROLE_NAME="" ROLE_TRAITS="" + for kv in $MARKER_LINE; do + case "$kv" in + role=*) ROLE_NAME="${kv#role=}" ;; + *) ROLE_TRAITS="${ROLE_TRAITS:+$ROLE_TRAITS }$kv" ;; + esac + done + field ROLE "${ROLE_NAME:-unknown}${ROLE_TRAITS:+ ($ROLE_TRAITS)}" +else + field ROLE "not bootstrapped (no $MARKER)" +fi diff --git a/test/cli.sh b/test/cli.sh index ec6c0f3..a05ecfb 100644 --- a/test/cli.sh +++ b/test/cli.sh @@ -977,6 +977,56 @@ else echo "skip: runner status/remove/repoint non-root refusals (running as root)" fi +# --------------------------------------------------------------------------- +# rig platform (#64). Unusually testable for this repo: it needs no root, no +# network and no fixtures, and it WRITES NOTHING — so unlike every other +# command here the harness can RUN it for real on the machine running the +# tests and assert on the actual answer, instead of proving arg-parse +# refusals and grepping the rest. +# --------------------------------------------------------------------------- +check "platform: --help exits 0" 0 "usage:" "$ROOT/commands/platform.sh" --help +check "platform: unknown flag exits 2" 2 "unknown flag" "$ROOT/commands/platform.sh" --nope +check "platform: dispatches through bin/rig" 0 "PLATFORM" "$ROOT/bin/rig" platform + +# The real run: exit 0 and every field present, as the running user. +check "platform: runs as this user, exit 0" 0 "PLATFORM" "$ROOT/bin/rig" platform +for f in HOSTNAME OS KERNEL CPU MEMORY DISK VIRT; do + check "platform: reports $f" 0 "$f" "$ROOT/bin/rig" platform +done +# Not just the labels — the VALUES have to describe THIS machine. uname -r and +# the hostname are the two the harness can independently compute and compare, +# which is what separates "it printed a table" from "it read the machine". +check "platform: KERNEL is this kernel" 0 "$(uname -r)" "$ROOT/bin/rig" platform +check "platform: HOSTNAME is this host" 0 "$(uname -n)" "$ROOT/bin/rig" platform +# MemAvailable/df rendered, not left as the 'unknown' fallback: a numfmt or +# /proc parse that silently broke would still print the labels above. +check "platform: MEMORY carries real numbers" 0 "total," "$ROOT/bin/rig" platform + +# Provenance degrades on a machine rig never converged — #61's manifest does +# not exist yet, so 'not bootstrapped' is the state of the world today and the +# command must ship complete without it. Both paths driven against fixtures. +PLATWORK="$(mktemp -d)" +printf 'version=9.9.9\nbootstrapped=2026-07-19T14:24:51Z\n' > "$PLATWORK/manifest" +check "platform: no manifest reads 'not bootstrapped'" 0 "RIG not bootstrapped" \ + env RIG_MANIFEST="$PLATWORK/absent" RIG_ROLE_MARKER="$PLATWORK/absent" "$ROOT/bin/rig" platform +check "platform: no role marker reads 'not bootstrapped'" 0 "ROLE not bootstrapped" \ + env RIG_MANIFEST="$PLATWORK/absent" RIG_ROLE_MARKER="$PLATWORK/absent" "$ROOT/bin/rig" platform +# A manifest that DOES exist is read, never written — the forward-compatible +# half, so #61 landing needs no change here. +check "platform: reads the manifest version" 0 "9.9.9" \ + env RIG_MANIFEST="$PLATWORK/manifest" RIG_ROLE_MARKER="$PLATWORK/absent" "$ROOT/bin/rig" platform +check "platform: reads the manifest timestamp" 0 "bootstrapped 2026-07-19T14:24:51Z" \ + env RIG_MANIFEST="$PLATWORK/manifest" RIG_ROLE_MARKER="$PLATWORK/absent" "$ROOT/bin/rig" platform +printf 'role=dev class=human host=yes join=authkey\n' > "$PLATWORK/role" +check "platform: renders the role marker's traits" 0 "dev (class=human host=yes join=authkey)" \ + env RIG_MANIFEST="$PLATWORK/absent" RIG_ROLE_MARKER="$PLATWORK/role" "$ROOT/bin/rig" platform +# The defining property: it writes NOTHING. Not the manifest it just reported +# missing, not the marker, not anything else in the fixture directory — the +# whole design rests on this, so assert it rather than trust it. +env RIG_MANIFEST="$PLATWORK/absent" RIG_ROLE_MARKER="$PLATWORK/absent" "$ROOT/bin/rig" platform >/dev/null 2>&1 +check "platform: writes nothing (no manifest created)" 1 "" test -e "$PLATWORK/absent" +rm -rf "$PLATWORK" + check "bare users shows usage, exit 2" 2 "usage:" "$ROOT/bin/rig" users check "users: bad subcommand exits 2" 2 "usage:" "$ROOT/bin/rig" users frobnicate