# cast Point it at a repo and a state directory; it makes a **Coolify** instance match what the repo declares. One-way, idempotent, never deletes. Philosophy (shared with [rig](https://github.com/heavy-duty/rig) and [claudebox](https://github.com/heavy-duty/claudebox)): **public tool, private state.** cast holds no hostnames, no bindings, no secrets, nothing about *your* infrastructure. It reads what you point it at and stores nothing, ever. `rig` builds the boxes. `cast` fills them. ## Install ```sh curl -fsSL https://raw.githubusercontent.com/heavy-duty/cast/main/install.sh | bash ``` That installs the **latest release**, and a cast release is a **prebuilt asset**: the installer resolves the newest tag by following GitHub's `releases/latest` redirect (no API, no token) and downloads that release's `cast-.tgz` — `bin/`, compiled `dist/`, production `node_modules/`, `package.json`, built once in CI — so **no `npm ci`, no `tsc`, no devDependencies ever run on your machine**. Three channels from the same script; `CAST_REF` picks ([#96](https://github.com/heavy-duty/cast/issues/96)): ```sh curl -fsSL .../install.sh | bash # the latest release (prebuilt) curl -fsSL .../install.sh | CAST_REF=0.1.0 bash # pinned to a release curl -fsSL .../install.sh | CAST_REF=main bash # the development tree, built from source ``` A set ref tries its release asset first, then falls back to source — `refs/tags` before `refs/heads`, so a tag outranks a branch of the same name — and only the source path needs `npm`. > **Transitional, until 0.1.0 is cut** (right after cast#96 lands): cast has > no GitHub release yet, so the default channel has nothing to resolve — it > **fails loudly** naming `CAST_REF=main` as the way to install today, and > never silently falls back to main. Once 0.1.0 exists, the plain > `curl | bash` above is the normal path. Needs `node` >= 22.12 and [`age`](https://github.com/FiloSottile/age) (secrets are decrypted by shelling out to it); `npm` only if you install from source. Re-run any time to upgrade. Unlike rig — which is pure bash so it can run on a bare box — cast runs on **your** machine: it is an API client, and a server should never install it. Installs are **versioned**, the same layout box and rig use: each install lands at `~/.local/share/cast/versions/` (the version is the tree's `package.json` version), a `current` symlink names the default, and the `cast` on your PATH points through it. Versions install side by side: ```sh cast versions # list what is installed, marking (current) and (running) cast use # switch the default — atomic, then asserted cast uninstall [|--all] # remove one non-current version, or everything ``` Re-running the installer with an already-installed version is a converging no-op (`CAST_REINSTALL=1` reinstalls that version's tree); a new version installs beside the old one and becomes the default — `cast use ` switches back. A pre-versioning flat install is migrated in place on the next run. The channel only decides **which** tree arrives and whether it is built here; every channel lands it the same way, in `versions/` — a prebuilt `0.1.0` asset and a `CAST_REF=main` source build sit side by side like any two versions. The installer symlinks `cast` into `~/.local/bin` (or `/usr/local/bin` as root) and, if that directory is not already on your `PATH`, appends it to your shell profile — `.zshrc`, `.bashrc`/`.bash_profile`, or `config.fish`, whichever your `$SHELL` reads — marked `# added by cast-install` and written only once. The shell you ran the installer from does not inherit it (a `curl | bash` pipeline is a subshell), so open a new shell or `source` the profile it names. Set `CAST_NO_MODIFY_PATH=1` to be left alone and wire `PATH` yourself. ## The two inputs cast joins a **manifest** (what to deploy) with **state** (where, and with what values). Neither knows about the other, which is the whole point: a manifest can live in a product repo without leaking your infrastructure, and your infrastructure can be re-pointed at a new Coolify without touching a product. **1. The product repo's `.infra/`** — committed, instance-blind: ``` .infra/ manifest.yaml # applications, databases, services, per environment env/..env.template # var NAMES + non-secret values; ${SECRET} placeholders ``` **2. A state directory** — private, yours: ``` environments.yaml # bindings: the team each env's token must belong to, # which server it deploys onto, the S3 destination, # GitHub App name, guards — and, per project, # the destination it deploys onto + its smoke target # …plus `projects:`, the registry: which projects # exist, and in which environments secrets/..env.age # age-encrypted values for the ${…} placeholders .coolify.env # COOLIFY_BASE_URL + COOLIFY_ACCESS_TOKEN (never commit) .coolify/.env # …the same, for a NAMED instance (see below) ``` Pass it with `--state `, or set `CAST_STATE`. Defaults to the cwd. ## Commands ```sh cast apply / --env [--path ] [--hostname-overlay ] cast apply --env --all # no repo: EVERY registered project cast diff / --env [--full] cast diff --env --all [--full] # no repo: EVERY registered project cast capture / --env [--generated ] [--override ] cast capture / --env --generated-only [--from =] cast inventory / --env cast inventory --env [--emit-draft [--recipient age1…] [--no-secrets]] cast destroy / --env [--instance ] [--with-project] cast server add --ip --key --env [--user root] [--port 22] cast github-app create / --env [--name ] [--port 8765] cast github-app register / --env --app-id --installation-id \ --client-id --client-secret-stdin --private-key cast smoke / --env [--project ] [--environment ] cast team [--env ] ``` - **`apply`** — idempotent create-or-update of every manifest resource, then redeploy what changed. One-way: it never deletes a resource that Coolify has and the manifest doesn't. It creates the **project** and its **environment** when they are absent — the two things a resource create has to name — and then removes the empty `production` that Coolify hands every new project, which is the single delete cast performs and never touches a project built by hand ([docs/semantics.md](docs/semantics.md)). Clones the repo's default branch unless `--path` points at a local checkout (refused with `--env prod` — prod always reads the default branch). - **`diff`** — reports drift, manifest → Coolify. Structural by default; `--full` also compares env vars. Exits non-zero when dirty, so CI can gate on it. - **`--all`** — on `apply`/`diff`, act on **every project the registry lists for this environment** instead of one named repo. See *The whole environment at once* below. - **`inventory`** — what is actually *on* a box. **With no repo it sweeps the instance** (every project, every environment, every resource — no manifest involved); with a repo it reconciles, showing resources and env var **keys** (never values) sorted into on-both / manifest-only / box-only. Needs no store, no age key, and no recipient — it runs *before* adoption, which is the point of it. A document, read by a person; nothing here is consumed by `apply`. See *Adopting a hand-built instance*. - **`inventory --emit-draft `** — the sweep, written down as a **draft of cast's own inputs**: a manifest per project, env templates, an `environments.yaml` with the registry, an age store, and `UNCAPTURED.md`. A **proposal**, never desired state — `apply` does not read it. It is how a project that has *no* manifest gets its first one, and how you take a point-in-time blueprint of a box. See *Drafting a box that was never declared*. - **`capture`** — the adoption path: reads a hand-built instance's live env and writes the environment's age store from it. See *Adopting a hand-built instance* below. With **`--generated-only`** it is instead **pass 2 of a bootstrap**: run *after* `apply`, it fills the store's provider-generated names (a Coolify-made `DATABASE_URL`) with the values Coolify generated. See *The bootstrap is two-pass* below. - **`destroy`** — the **only** verb that deletes what a manifest declared, and the reason `apply` never has to. **Manifest-scoped**: it removes the resources this manifest declares in this project and this environment, in reverse dependency order (applications → services → databases), and **reports everything else it finds without touching it**. It refuses `--all`, refuses a read-only instance, refuses an absent project, and refuses any environment whose `environments.yaml` binding does not carry `destroy_allowed: true`. The last gate is typing the environment's name at a plan that says, for every database, whether it is backed up and when the last backup landed. See *Tearing an environment down* below. - **`server add`** — uploads a server's private key and registers it with Coolify. - **`github-app create`** — creates the GitHub App Coolify clones private repos with, by running GitHub's App Manifest flow, then registers it. Two browser clicks, zero transcription. See *The GitHub App* below. - **`github-app register`** — adopts an App you already hold: one created by hand, or a disaster-recovery restore from a stored private key. `create` ends by running exactly this. - **`smoke`** — contract test against the project's `smoke_target`: proves Coolify's bulk env endpoint still *upserts* rather than replacing. Run it after every Coolify upgrade — `apply`'s never-delete guarantee rests on that behavior, and the published OpenAPI does not describe it accurately. It **writes** (two canary env vars onto that one application, then deletes them), so the repo is required: the target is resolved *inside the project and environment it was declared under*, with `--project` / `--environment` if the box names either differently, and it refuses rather than guessing when no application of that name is there. A bare app name is unique nowhere else — one instance carrying prod and staging is enough for the first `core` on it to be prod's. - **`team`** — prints the team the configured token acts as. With `--env`, also checks it against that environment's `team:` binding and exits non-zero on a mismatch — the dry run for "would `apply` refuse?", answered without touching anything. Every command that reaches a live Coolify takes an `--env`, because every one of them first asserts the token's team (below). `--hostname-overlay` swaps domains for a pre-flight run against temporary hostnames; re-applying **without** it is the cutover. ## Cloning: cast authenticates, and never prompts `apply`, `diff` and `capture` clone the product repo (unless `--path` points at a local checkout — refused for prod, which always reads the default branch). For a private repo that needs credentials, and cast resolves them itself: 1. **`gh`**, borrowed as a credential helper for that one invocation — it does not touch your global git config. 2. **`GITHUB_TOKEN` / `GH_TOKEN`** from the environment (the CI path). 3. Whatever git's own credential helper does, if you have one. Being logged into `gh` is enough. You do **not** need `gh auth setup-git` — that separate act is what wires git's helper, and not running it is exactly how you end up at git's interactive username/password prompt, which GitHub no longer accepts. cast sets `GIT_TERMINAL_PROMPT=0` on every path, so it can never hang there or hide a credentials failure behind an error about *the repository*. With no credentials at all it says so, and names the fix. The token is never put in the clone URL or in `http.extraheader` — both leak it into `ps`, and the latter persists it into the clone's git config. ## The GitHub App: `cast github-app` That is how *cast* clones. **Coolify** clones with a GitHub App, and the App used to be the one piece of a Coolify instance cast could not reproduce: created by hand in a browser, its four identifiers copied out of the UI by eye, its private key downloaded to `~/Downloads`, its details fed to a shell script as a six-variable env pile. Nothing about that survived in state. Rebuild the instance and you redid the hoops from memory. ```sh cast github-app create heavy-duty/incubator --env prod --name hdb-coolify-prod ``` There is **no REST endpoint that creates a GitHub App** — no `POST /apps`, no GraphQL mutation, no `gh app` subcommand, and no PAT scope that unlocks one. The only programmatic path is GitHub's [App Manifest flow](https://docs.github.com/en/apps/sharing-github-apps/registering-a-github-app-from-a-manifest): a browser form POST whose authentication is your existing GitHub session, followed by an unauthenticated code exchange. It is how Coolify's own *Create GitHub App* button works, and it is why this command serves you a page instead of calling an API. What `create` does: 1. If `gh` is on `PATH` and authenticated, checks you are an **admin** of the org — so you learn you cannot create Apps there *before* the browser dance, not after. `gh` is never required; an absent one skips the check silently. 2. Resolves the App's Coolify-facing name from **`github_apps./` in `environments.yaml`**, which is what every later `cast apply` resolves this repo's App by. `--name` seeds that entry when it is absent and is **refused** when it disagrees with one that exists. 3. Serves a one-shot page on `127.0.0.1` that submits an App manifest — `contents: read` + `metadata: read`, webhook inactive, private. 4. You click *Create GitHub App*; GitHub redirects back to the loopback server, which checks the CSRF `state` and shuts down. 5. Exchanges the code. **This response is the only moment GitHub ever hands over the private key, the client secret and the webhook secret together.** 6. **Writes all three to disk immediately**, before waiting on anything — see [Where the credentials land](#where-the-credentials-land). Everything after this point can fail for ordinary reasons (a slow install screen, a dropped network, `Ctrl-C`), and none of those may cost you a key GitHub will not reissue. 7. Prints (and tries to open) the install URL; you pick the repository. 8. Recovers the installation id by minting an RS256 JWT with the App's own key — never from the `installation_id` GitHub appends to a redirect, which GitHub documents as a spoofable hint — then fills it into the record from step 6. 9. Uploads the key to Coolify and creates the App record — unless a Source of that name already exists, in which case it verifies that one rather than registering a second (Coolify does not enforce unique Source names). 10. **Asks Coolify which repositories the App can actually see, and fails if `/` is not among them.** This is the step that matters most: without it a misconfigured App fails silently and surfaces hours later, in a different command, as an unresolvable source at `cast apply` time. If the install never lands, `create` stops at step 8 and tells you the exact `register` command that finishes the job against the files from step 6. Nothing is lost and nothing has to be recreated — in particular, do **not** re-run `create`, which would mint a second App. For that same reason `create` refuses up front, before the browser flow, when `.pem` already exists. `register` is the same command from step 9 onwards, for an App you already hold — one made by hand, or a disaster-recovery restore from a stored PEM: ```sh pbpaste | cast github-app register heavy-duty/incubator --env prod \ --app-id 12345 --installation-id 99887766 --client-id Iv23li… \ --client-secret-stdin --private-key ~/Downloads/app.private-key.pem ``` The client secret is read from **stdin only** — argv is visible in `ps` and kept in shell history. `--webhook-secret` is optional: a webhook-**inactive** App is the right shape for a tailnet-only Coolify where deliveries can never arrive and deploys are CI-triggered, and cast generates a value rather than making you invent one. ### Where the credentials land Into the state directory you point cast at — cast itself stores nothing: ``` /github-apps/ ├── .gitignore # `*` — written by cast ├── .pem # 0600, the private key └── .json # 0600, app id, installation id, client id + secret, webhook secret ``` Both are written the instant GitHub yields them, which is *before* `create` waits for you to install the App. Until the install lands, `.json` carries `"installation_id": null` — that is the one field GitHub will answer again as often as it is asked, and it is filled in on success. Re-running against an existing file is idempotent on identical content and a **refusal** otherwise; `--force` is the deliberate escape hatch for a stale half-run. All three secrets, because GitHub shows them once and `register` needs the client secret to be re-runnable at all — a disaster-recovery restore that is missing it is not a restore. They are written **plaintext at 0600**, not into `secrets/`: that store holds per-repo-per-env *application* env vars, whose whole purpose is to be decrypted and injected into the running container, which is the last place an App private key belongs — and its age identity may not exist on the machine doing the bootstrap at all. Encrypting the one credential that makes recovery possible behind a key that might not be there is how DR fails at the moment it is needed. So the guard is structural rather than cryptographic: the `.gitignore` means `git add -A` in your state repo cannot commit these by accident. Committing them stays possible and has to be deliberate — encrypt them yourself and commit the ciphertext, or keep the directory out of the repo and back it up somewhere that is not a git remote. ### Until it has worked once `create`'s design rests on GitHub accepting a `redirect_url` on `http://127.0.0.1:`. The manifest docs are silent on the scheme (loopback HTTP is documented for *OAuth* redirect URIs), and the precedent is strong — Probot's setup flow does exactly this — but it is unvalidated, because validating it needs a logged-in GitHub session. **The manual path below stays supported until `create` has succeeded against a real GitHub once.** If it fails, create the App by hand in the browser and use `github-app register`, which does not depend on the assumption at all.
The manual path 1. Org → Settings → Developer settings → GitHub Apps → **New GitHub App**. Permissions: **Contents: Read-only**, **Metadata: Read-only**. Uncheck *Active* under Webhook. Uncheck *Any account* (keep it private). 2. Note the **App ID** and **Client ID**; generate a **client secret**; generate and download a **private key**. 3. **Install App** → pick the repository. The installation id is the last path segment of the URL you land on (`…/settings/installations/`). 4. Feed all of it to `cast github-app register` (above), which validates the name against state and verifies the repo is reachable.
## Many Coolifys `--instance ` reads `/.coolify/.env` instead of `/.coolify.env`. Every verb that reaches Coolify takes it. ```sh cast diff heavy-duty/incubator --env prod --full --instance legacy ``` An environment can bind one, so `--env` selects the right control plane with no flag at all: ```yaml environments: prod: server: prod-box team: { id: 1, name: heavy-duty } instance: prod-cp # → /.coolify/prod-cp.env ``` An explicit `--instance` still wins, so a one-off read against a legacy box needs no edit to that file either. **With no flag and no binding, nothing changes** — `.coolify.env` is read exactly as before. Two properties, both deliberate: - **An unknown `--instance` refuses**, and names the instances that do exist. Falling back to the default is how a diff meant for a legacy box gets run against production. - **An instance may declare `COOLIFY_READ_ONLY=true`**, and then `apply`, `smoke` and `server add` refuse it — *before their first call*, and even though the token itself would permit the writes. That turns "I pointed the wrong token at the wrong box" from a live incident into an exit code. Every command that reaches a Coolify now says which one, next to the team assert. It is the most consequential input to any run, and the least visible. ## Adopting a hand-built instance cast is otherwise scoped to the steady state: manifest → Coolify, forever. Adoption is the one way in, and it has two verbs and a fixed order: **`inventory` → you read it → a manifest PR → `capture` → `apply`** **Look before you adopt.** A box nobody declared does not use your vocabulary: its project is called whatever someone typed, its environment is Coolify's default (`production`, not `prod`), and its resources are named by whoever clicked *New Resource* that afternoon. `inventory` shows you both sides at once, so those differences arrive together, as a document — instead of one at a time, as refusals from a verb that is already halfway through a migration. **First, sweep it — you cannot aim at coordinates you do not have yet:** ```sh cast inventory --env prod --instance legacy ``` ``` sweep — instance legacy (https://coolify.example.com) Incubator production (empty) staging 2 applications, 2 databases, 1 service application Incubator Stack v2 application Incubator Landing database Incubator Database v2 … La Familia Site production 1 application application lafamilia-web ``` Note what that costs you to *not* have: Coolify auto-creates a `production` environment in every project, so the obvious guess is empty and the live system is somewhere else entirely — under a name someone typed, in a project you may not have known was there. An environment with **zero** resources is far more often the wrong coordinate than an empty one, and `inventory` says so rather than quietly reporting that the manifest has five things the box lacks. **Then reconcile**, against a target you now know exists: ```sh cast inventory heavy-duty/incubator --env prod --instance legacy \ --project Incubator --environment staging \ --resource core="Incubator Stack v2" ``` It never reads a value, needs no store and no key, and its output is **not** desired state. What you do with it is decide, resource by resource and key by key, what the manifest should *gain* and what is cruft that must not travel — and land that as a manifest PR. Only then: ```sh CAST_CAPTURE_ADMIN_EMAIL=me@example.com \ cast capture heavy-duty/incubator --env prod --instance legacy \ --override ADMIN_EMAIL ``` It reads the required secret **names** from the manifest's own env templates (the `${…}` refs — the manifest already declares exactly this set), reads the live values off the instance, and classifies every name: | | | | --- | --- | | **captured** | found live, value taken | | **generated** | the manifest's `generated_secrets` declares it provider-made → written as the literal `pending-coolify-generated`, never the live value | | **overridden** | supplied by you, for a value that must *not* be carried over | | **missing** | required by a template, absent live → **refuses** | Then it prints a plan of **names and provenance — never values** — and waits for you to type the environment's name. Before any of that, it checks that the resources the manifest names **exist**. An absent resource reads back exactly like one with no env vars set: every name it declares reports *missing*, and `--override` would then have you hand-carry values that are sitting right there under a different name — writing a perfectly valid store while the actual finding (the manifest and the box disagree about what this thing is called) is never discovered. So a resource that isn't there refuses, and names what is. `inventory` is how you reconcile it. ### Three names that are not yours A hand-built box names things without asking you, at three levels, and cast takes each as a coordinate to *read* with — never as a reason to rename anything of yours: | flag | when | | --- | --- | | `--project ` | the project isn't named after the repo (`Incubator`, not `incubator`) | | `--environment ` | the environment isn't named after `--env` (Coolify's default is `production`, not `prod`) | | `--resource =` | a resource isn't named after the manifest's (`core` is `Incubator Stack v2` over there). Repeatable | None of them is ever a manifest field: they are arguments to a single run, because a manifest that recorded a legacy box's names would carry a dead machine's vocabulary forever. `--project` and `--environment` are how a verb that must *find* a target says where to look — `diff`, `capture`, `inventory`, `apply`, and `smoke`, which resolves its `smoke_target` in exactly that project and that environment, and refuses when it is not there (#29). `--resource` is **read-side only** (`diff`, `capture`, `inventory`): `apply` refuses it outright, because it creates resources under the manifest's own names — an alias there could only mean *adopt the existing one instead*, which is a different operation and would otherwise silently create a duplicate beside the resource you were pointing at. `--env` stays **ours**: it selects the manifest block, the `environments.yaml` binding, the age key, the store path. `--environment` is *theirs*, on the wire, and nothing else. Collapsing the two lets a box that is being deleted next week name the environment of the box that replaces it — `apply` creates the environment from that value, so it would be inherited permanently. The mapping is not mechanical, and that is the whole design. A `DATABASE_URL` copied off the source box points at the *source box's* Postgres: confidently wrong, entirely plausible, and the target's real URL does not exist until Coolify creates the resource. So the manifest declares those names, and cast placeholds them: ```yaml environments: prod: generated_secrets: [DATABASE_URL_PROD, REDIS_URL_PROD, UMAMI_DATABASE_URL] ``` It is a manifest property rather than a flag you have to remember, because the manifest is what knows `DATABASE_URL` comes from a database it declares. (A `generated_secrets` entry no template refers to is a schema error — a guard standing over nothing is worse than no guard, because it reads like one. `--generated ` covers a manifest that hasn't declared them yet.) > **A database's own URL is better *derived* than stored.** If the value is the > URL of a database this manifest declares, write `DATABASE_URL=${resource:postgres.url}` > in the template instead of storing it. cast reads it back from the database it > created — never into the store, never through a terminal — so there is no > placeholder, no two-pass dance, and no stored copy to drift or overwrite. A > rotated password is simply *followed* on the next apply. `generated_secrets:` > and the pass below stay for the residual class — a provider-generated value > that genuinely is not derivable. See [semantics.md](docs/semantics.md) → > *Derived resource URLs*. > **A declared domain is better *derived* than transcribed.** If the value is a > base URL an app is told to call itself (or a sibling) at, write > `LANDING_BASE_URL=${domain:landing}` — or `${domain:core.admin}` for a compose > app's per-service domain — instead of copying the hostname into the template by > hand. It resolves to the manifest's own `domains` / `service_domains` (verbatim, > scheme and all), at plan time, so there is nothing to drift. Unlike a resource > URL a domain is **public**, not a secret: it is not stored, not captured, and it > prints in a diff like any literal. Applications only. See > [semantics.md](docs/semantics.md) → *Derived domains*. > **An application can declare HTTP basic auth, and its password is a `${REF}` > like any other secret.** Write > `basic_auth: { enabled: true, username: ops, password: ${ADMIN_PW_PROD} }` on > the application, and `apply` sets it — closing the "a rebuilt resource comes > back UNPROTECTED" hole for applications (services have no API for it at all, on > 4.1.2 or v4.2). The schema **refuses a literal password**: a manifest is a > committed file, so a literal there is a password in git forever. What cast > cannot do is *verify* it — the password reads back only to a token with > sensitive-data reads, so `diff` compares the toggle and the username (a UI flip > is still caught) and says on every run that the password was not compared, > rather than implying it matches. See [semantics.md](docs/semantics.md) → *HTTP > Basic Auth on an application*. An **`--override`**'s value is read from `$CAST_CAPTURE_`, never from the command line: argv is visible in `ps` to every process on the box. It exists for values that must not survive the copy — staging and prod sharing a Mailgun domain means a staging box carrying the real `ADMIN_EMAIL` can mail real users. The store is encrypted to the environment's `age_recipient` (add it to `environments.yaml` — it's the public half, safe to commit). Plaintext goes to `age` on stdin: it is never a temp file, never on stdout, never in your shell history. An existing store is not overwritten without `--force`. ## The bootstrap is two-pass: `capture --generated-only` Those placeheld names are the reason bootstrapping an environment **cannot be one pass**. The value does not exist until Coolify makes it: ```sh cast capture heavy-duty/incubator --env prod # 1. the store learns every # name; generated ones are # placeheld — nothing has # created them yet cast apply heavy-duty/incubator --env prod # 2. Coolify creates the # database, and generates # the real URL cast capture heavy-duty/incubator --env prod \ # 3. the store learns THAT --generated-only --from DATABASE_URL=incubator-db # value ``` Until step 3 runs, the store says `pending-coolify-generated` while the live value is real — the exact state in which the next routine `apply` overwrites a working secret. It is also on the DR path: *rebuild the control plane from state* means apply-from-nothing, so every generated secret in every store is a placeholder again. This used to be a hand `age` re-encrypt against production, with the prod key in a process substitution. **Greenfield needs zero passes** (#104): a manifest whose templates hold no `${…}` refs at all — databases only, or apps whose env is pure literals — applies from nothing. No store, no age key: `diff` and `apply` say out loud that the store is absent and was not needed, and proceed. The store appears the first time `capture` writes it, or the first time a template gains a placeholder — from then on, an absent store refuses exactly as before. `--generated-only` **inverts** capture's rule and changes nothing else: it fills the `generated_secrets` names and leaves every other name in the store **exactly as it is, byte for byte** — never re-read from the box, so a secret you rotated by hand last month survives it. Same typed confirmation, same names-never-values plan. The store must already exist: pass 2 *fills* names, it does not create them. It reads each value from the **database that owns it** (`internal_db_url`) — never from the consuming app's env, where a generated URL never appears (the app's env holds what the template resolved to, which at this point is *the placeholder itself*). And it resolves that database **inside your project and environment only**, never from the instance-wide `GET /databases` list, which holds every database on the box — other projects', and umami's own bundled Postgres. **It will not guess which database a name comes from.** Nothing carries that edge: `generated_secrets:` is a flat list of names, and the template knows only `DATABASE_URL=${DATABASE_URL}`. So cast infers it only when it *cannot* be wrong (one name, one database) and otherwise refuses, printing the flag you need. A pick made from the *name* (`REDIS_URL` → the redis one) is wrong **silently**, and what it writes is a well-formed URL to somebody else's database. Four refusals, each a thing that used to be a step in a runbook: | | | | --- | --- | | **UNMAPPED** | more than one database could be meant → say which, with `--from` | | **OCCUPIED** | the name already holds a *real* value → filling it silently rotates a live credential. `--force` to mean it | | **ABSENT** | the name is not in the store at all → pass 1 has not run, or you are pointed at the wrong store | | **PENDING** | a placeholder in a name nothing here fills → the store would still be a lie | Afterwards it **asserts the postcondition**, against the ciphertext now on disk: zero `pending-coolify-generated` remain, and the name count is unchanged. A store that lost a name re-encrypts perfectly and reads back perfectly — you would find out at the next `apply`, in an environment whose plaintext nobody has any more. `apply` deliberately does not do this for you after a create. It would close the window entirely, but it would make the verb that mutates Coolify also mutate the encrypted store — and hence the git repo — which is a much bigger blast radius for a verb people run on a schedule. **[docs/semantics.md](docs/semantics.md)** is the contract behind those commands: what `apply` guarantees (never deletes, never recreates a database, fails loudly rather than recreating on un-updatable drift), the `dockercompose` build pack, the hostname-overlay shapes, and the places Coolify 4.1.2 does not cooperate — each citation verified against `coollabsio/coolify` v4.1.2 and the vendored OpenAPI in `reference/`. Read it before changing `apply`. ## Drafting a box that was never declared `inventory` can see a whole instance. `--emit-draft` makes it **write down what it sees**, in the shape of cast's own inputs: ```sh cast inventory --env prod --instance box-b --emit-draft ./draft --recipient age1… ``` ``` draft/ environments.yaml # bindings as far as they can be read — with the projects: registry incubator/.infra/manifest.yaml # one per project incubator/.infra/env/*.env.template la-familia/.infra/manifest.yaml # …including the client sites nobody ever declared secrets/..env.age # encrypted to a recipient you name UNCAPTURED.md # ← the important file ``` Two uses: **bootstrapping a project that has no manifest** (the third-party sites on the box being drained were never declared, and never will be unless something writes the first draft — hand-transcribing them from a UI is exactly the work cast exists to eliminate), and **a point-in-time blueprint** you could rebuild an instance from. ### A draft is a PROPOSAL **It is never desired state, and `apply` never reads it.** It is emitted, reviewed by a human, and lands in a repo as a PR — the same shape as `terraform import` → HCL: **sweep → emit draft → you read it → manifest PR → `capture` → `apply`** That boundary is the only reason the verb is allowed to exist, and it is enforced, not merely documented: - It **never emits into a repo that already has a manifest.** For a declared project the manifest *is* the truth, and one regenerated from a live box would let that box's accumulated cruft overwrite the reviewed spec — in the one direction nobody reviews. Adoption is one-way. A non-empty target directory is refused, and so is writing a manifest over an existing one. - `--emit-draft` is **sweep-mode only**. With a repo, `inventory` is reconciling against a manifest that already exists, which is exactly the case where a draft must not be written. Refused. ### Two things would make a draft actively dangerous **1. Copied provider-generated values.** A `DATABASE_URL` read off the source points at the *source box's* Postgres. Emit it, rebuild elsewhere, and the new box comes up **working** — reading and writing the old box's database. You find out the day the old box is deleted. Same for `REDIS_URL`, and for Coolify's own magic vars (`SERVICE_FQDN_*`, `SERVICE_URL_*`, `SERVICE_PASSWORD_*`), which are generated per-instance and mean nothing anywhere else. So the draft applies **`capture`'s discipline**: a provider-generated name is **placeheld** with the same `pending-coolify-generated` literal, its live value is not written into any artifact, and it is listed for disposition. The emitted manifest declares it under `generated_secrets:`, so a later `capture` placeholds it again with no flag to remember. A draft that is confidently wrong in four entries out of seventeen is worse than one that is obviously incomplete. The rule is **by name** — two families: Coolify's `SERVICE_*` magic vars, and any name carrying a *datastore* word (`DATABASE`, `DB`, `POSTGRES`, `REDIS`, …) and a *connection* word (`URL`, `HOST`, `PASSWORD`, …) as segments. It errs **wide** on purpose, because the two errors are not symmetric: over-matching a real secret placeholds it loudly and you put it back, while under-matching a generated one copies it silently and rebuilds a box that quietly uses a dead machine's database. Every value cast read is printed with its disposition — names and provenance, never values — and a var that points at the source box under a name cast does not recognize **will** have been copied. Read the table. Every other live var becomes a `${REF}`, with its value in the **age store** — never a literal in a committed file. cast cannot know which of a box's vars are secret (nobody wrote it down; that is why this verb exists), and a live API key written as a literal is a key in a git repo. Move the plainly-not-secret ones back to literals yourself, in review. The store is encrypted to a recipient you **name** — `--recipient age1…`, or the environment's `age_recipient` binding. With neither, cast **refuses**: a draft whose secrets were silently skipped looks complete and holds not one value. `--no-secrets` says so deliberately. **2. Silent losses.** cast cannot express everything a Coolify holds: destinations (which Docker network a resource sits on — no API at all in 4.1.2), service hostnames (they live per-container on `service.applications[].fqdn`), Basic Auth on a *service* and custom Traefik labels anywhere (an application's Basic Auth is expressible — see below — but its **password** is not readable, so a draft reports it instead of emitting a block a rebuild could not honour), the *Include Source Commit in Build* toggle, whole database kinds (a MySQL is invisible to cast's manifest), backup schedules, and anything else configured in the UI with no manifest field. A blueprint that omits these **without saying so** is worse than no blueprint, because in a disaster you would trust it and rebuild a *different box*. So **`UNCAPTURED.md` is a first-class output**, listing per resource every live setting cast saw and could not express — and it is **written on every run**, even when it has little to say. ### What a blueprint still cannot restore Worth stating plainly, because "rebuild from the repo" is routinely over-claimed: | | | | --- | --- | | control plane | `rig coolify install` ✅ | | structure | draft → manifest PR → `apply` ✅ | | secret **values** | the age store + your key ✅ | | **data** | Coolify's DB backups → S3 ✅ (a separate path) | | **the GitHub App private key** | ❌ re-create by hand | | **S3 access keys** | ❌ re-mint by hand | The last two are **not in the repo** — correctly; it holds no live credentials — and cannot be regenerated from it. A DR runbook has to say so. The same table is emitted into every `UNCAPTURED.md`, because that is the file someone will be reading at the worst possible moment. One project per Coolify environment: a project with resources in **two** populated environments is a tie cast will not break (picking would emit a blueprint of half a box), so it refuses and `--environment ` says which. It is a tiebreak, not a filter — a project with only one populated environment is drafted from it either way, which is what keeps the client sites (each alone in Coolify's default `production`) in a blueprint that claims to describe the box. ## Secrets, and attended applies An environment's age identity is resolved in exactly two ways: 1. `$CAST_AGE_KEY_FILE_` — injected for this invocation. `` is the environment name uppercased, with every character a shell cannot carry in a variable name mapped to `_`: env `drill-b` reads `CAST_AGE_KEY_FILE_DRILL_B`. 2. `~/.config/cast/age-.key` — a standing key on this machine, under the environment's exact name. That is the whole mechanism behind attended vs unattended applies: **an environment whose key you never leave on disk can only be applied by someone who injects it.** Keep a standing key for staging if you like; keep prod's in a password manager and pass it per apply, straight from the manager with a process substitution: ```sh CAST_AGE_KEY_FILE_PROD=<(pm read cast-prod-key) cast apply heavy-duty/incubator --env prod … ``` cast reads the identity itself and hands it to age on stdin, so this works even though `<(…)` yields a path only cast's own process can resolve — and the key never becomes a file, never appears in argv, and never enters the environment. The state directory holds ciphertext. It must never hold the identity that opens it. ## Teams: the one assert cast makes before it touches anything Coolify API tokens are **team-scoped**, and a token pointed at another team's resources **does not error**. The API resolves what the token cannot see to `null` — and to a tool like cast, `null` is indistinguishable from *"this resource does not exist yet"*, which is an invitation to create it. An `apply` run with a wrong-team token would not fail; it would silently provision a **duplicate set of resources into the wrong team**, on whatever server that team owns. Silent, mutating, discovered late. So every environment declares the team its token must belong to, and cast refuses to do anything at all until it has checked: ```yaml environments: prod: server: prod-box team: { id: 1, name: heavy-duty } ``` Give `id`, `name`, or both — both are compared when both are given. `id` is the true identity (names can be renamed); `name` is what makes the file readable. Run `cast team` to print the values for the token you currently have configured. The check is **fail-closed**: an environment with no `team:` is one whose token cannot be verified, so it is a schema error, not a warning. It runs before the first *read*, not merely before the first write — an unasserted `diff` against the wrong team would report "everything is absent", which is precisely the lie that an `apply` would then act on. Nothing below the team scopes a token. A Coolify environment has no team of its own (it hangs off a project) and no API path scopes by one: **Coolify environments are an organizational construct, not an auth boundary.** The team is the only boundary there is, so it is the one cast asserts. ## The registry: which projects exist `environments.yaml` says where things deploy *to*, and how a project you have already named is placed once it is there. Until the `projects:` block, nothing in it said **which projects exist at all** — "every project" was a thing the operator remembered: ```yaml projects: heavy-duty/incubator: environments: [prod, staging] acme/client-site: environments: [prod] ``` Keyed by the **full `/` slug**, and the key *is* the repo — there is no `repo:` field inside, because a second place to write the same string is a second place for it to be wrong. Unlike `github_apps`, a bare `` key is **refused** rather than resolved: this block is new, so it has no state files in the wild to keep working, and a bare `` is unique only *within* an org — which is exactly why it is not a key. `environments:` lists **our** environment names (the values `--env` takes), never Coolify's. The block is optional; a state file written before it loads unchanged. **It has to be true, so cast checks that it is — at parse time, for every verb.** Two ways it could quietly stop being true, both refused: - an environment name that no `environments:` block defines (a typo). The project is real and its environment imaginary, so a fleet run visits nothing for it, reports nothing, and exits clean. - an `environments..projects.` binding — a destination, a smoke target — in an environment the registry does not register that project for. The two blocks then describe two different fleets: state real enough for a direct `cast apply` to use, invisible to every fleet run. (Checked only when `projects:` is present.) Both refusals defend one failure: **a silently skipped project reads exactly like a clean one.** Silence is the one report that must never be ambiguous. What it unlocks, neither of which was possible without a list to iterate: - **fleet operations** — `cast diff --all` / `apply --all` over every project in an environment (below). - **rebuild-from-state** — "restore this Coolify from the state repo" cannot even be *attempted* without knowing what was on it. The registry is the difference between a documented recovery and an archaeology exercise. ## The whole environment at once: `--all` Every other cast verb is single-project, so *"do this to the whole instance"* was a shell loop the operator wrote from memory — **and the project they forgot is the one that drifted.** `--all` iterates the registry instead: ```sh cast diff --env prod --all # every project registered for prod cast apply --env prod --all ``` It runs the **same per-project path** the single-repo form runs (there is one implementation of what a project run *is*), reports each project under its own heading, and then prints an aggregate that leads with **coverage**: ``` fleet diff — prod registered: 3 (heavy-duty/incubator, acme/client-site, acme/landing) read: 2 of 3 clean: 1 heavy-duty/incubator drift: 1 acme/landing UNREACHABLE: 1 acme/client-site acme/client-site: refusing to diff: no project named "client-site" … ``` **A registered project cast cannot reach is an ERROR, not a skip** — the clone failing, no manifest block for this environment, a missing or undecryptable store, an absent Coolify project or environment, any HTTP error. A skipped project reads exactly like a clean one, which would make the one report you get most often (silence) the one you cannot trust. So the fleet **fails closed**: a clean fleet diff means *every* project was read, not that the ones cast happened to look at were fine. | | | |---|---| | `diff --all` exit 0 | every registered project was **read**, and every one is clean | | `diff --all` exit 1 | every one was read, and at least one has drift | | `diff --all` exit 2 | a project could not be read — **outranks drift**, because an unread project is not a diff result, it is the absence of one | | `apply --all` exit 0 | every registered project applied | | `apply --all` non-zero | anything else | The two verbs take opposite dispositions on failure, and both are deliberate: - **`diff --all` runs every project to completion.** Stopping early would hide the drift in the projects it never reached — a partial read is exactly the report this flag exists to make impossible. - **`apply --all` stops at the first failure**, and says which projects it applied and which it did not touch. Continuing to *mutate* a fleet after an unexplained failure is not a thing cast gets to do. `apply` is idempotent, so re-running after a fix is a no-op over the ones that already applied. **An empty or absent registry refuses** (exit 2). `--all` over a state file with no `projects:` block does *not* print "0 projects, clean" and exit 0 — an empty fleet reading as a clean fleet is the whole failure this feature is against. The refusal names what it looked for and prints the YAML to write. **`--all` is mutually exclusive** with the repo positional and with every single-project coordinate — `--path`, `--project`, `--environment`, `--resource`, `--hostname-overlay`. Each of those names ONE project's checkout, ONE project's Coolify name, ONE box's resource names; fleet-wide they are meaningless at best and dangerous at worst (`--project X` applied to every project in the registry would point them all at the same Coolify project — a false report on `diff`, and on `apply` every manifest in the fleet written into one project). The refusal names the offending flag. ## Two projects, one box: destinations A **destination** is the Docker network a resource is created on. A server has a default one, and while a server hosts a single project that default is the right answer — which is why cast went so long without naming it. The moment a server hosts *two* projects, it stops being: they share one network, and "isolated" becomes a thing you believe rather than a thing that is true. So the destination is declared per **project**, inside the environment — an environment-scoped key could not express it, because `server:` is precisely the thing two projects share: ```yaml environments: prod: server: shared-box team: { id: 1, name: heavy-duty } projects: heavy-duty/incubator: destination_uuid: # the network THIS project's resources go on smoke_target: core # the app `cast smoke` writes its canary to acme/client-site: destination_uuid: ``` Keyed by repo, full `/` slug first, exactly like `github_apps` — a bare `` key still resolves, so existing state files keep working. Both fields are optional, and an environment whose server hosts one project needs neither. A **UUID and not a name**, unlike `server:` right above it. Coolify 4.1.2 has no destinations API whatsoever — no list, no read, nothing — so there is no name for cast to resolve. You read the UUID out of the Coolify UI, the same way you do for `s3_destination`. **What cast can and cannot promise here.** It sends `destination_uuid` on create, for applications, databases and services alike. It can never check it afterwards: Coolify takes a UUID on write and hands back an integer `destination_id` on read, and nothing maps between them. So `diff` does the one honest thing left — it groups the live resources by the id Coolify *does* report, and a project whose resources do not all share one network is **drift**: ``` split placement: these resources sit on 2 different destinations destination 1: application landing, database postgres destination 4: application core a project's resources must share one destination — that is what the isolation IS. apply never moves a live resource between networks: resolve manually (runbook act). ``` …and when you declare a destination, every `diff` says, out loud, that it did not verify it. That is deliberate. A setting that reads back as *absent* rather than *wrong* is the failure this whole file keeps trying not to be. When you declare **nothing**, every `diff` says that too: ``` placement: server's default destination (none declared) — cast sends no destination_uuid, so Coolify picks; a server with more than one destination refuses the create outright. ``` Declaring nothing is not the absence of a placement decision. It is one, and it used to be the only one cast made silently — the inference sat in a source comment ("the server's only destination, which is what Coolify picks anyway"), which is exactly where an assumption is invisible until it is wrong. One sharp edge worth knowing: on a server with exactly **one** destination, Coolify ignores the `destination_uuid` you send and never validates it — a typo there is invisible until a second destination exists. On a server with more than one, a create that omits it is a hard `400`, which is why cast could not deploy onto a shared box at all until it could send this. The asymmetry hides itself: the day a server gains its second destination, every project on it that declared no destination stops being able to create. cast cannot warn you before that create — Coolify 4.1.2 serves no destinations API, so a server's destination *count* is not knowable until a create has already been attempted, and the 400 therefore lands **after** apply has made the project and the environment. What cast does instead is answer it: ``` cannot create application core: prod-box has multiple destinations, so a create must say which one to use. Coolify said: POST /applications/private-github-app → 400: {"message":"Server has multiple destinations and you do not set destination_uuid."} Read the destination UUID from the Coolify UI (4.1.2 exposes no API for it) and declare it as: environments.prod.projects.heavy-duty/incubator.destination_uuid Placement is create-time — a resource cannot be moved between networks later, so a wrong or missing destination is repaired by delete + recreate, never by a later apply. Re-run this apply once the UUID is declared: anything it already created (the project, its environment) is adopted, not made twice — apply reads before it writes. ``` Details, with citations: [reference/README.md](reference/README.md). ## Guarding an environment An environment may refuse variables by name pattern: ```yaml environments: prod: server: prod-box team: { id: 1, name: heavy-duty } forbidden_var_patterns: ["^ALLOW_"] ``` `apply` then refuses if any such var is **present** on any resource, regardless of value. `ALLOW_SEED=false` still fails: a var that exists can be flipped on later in the Coolify UI without touching a manifest, so "off" has to mean absent. This guard lives in your private state deliberately — not in the product's manifest. A product-side change must not be able to lower its own guard. ## Tearing an environment down: `cast destroy` ```sh cast destroy heavy-duty/incubator --env staging [--with-project] ``` `apply` fails closed on an immutable field (`build_pack`, `type`, `version`, placement) with *"resolve manually"* — which used to mean a hand deletion in the Coolify UI, against an instance whose token can see every project on it. That is how you delete the wrong project. `destroy` is that act, scoped and gated. **What it deletes:** the resources **the manifest declares**, in this project and this environment, in reverse dependency order — applications, then services, then databases. Nothing else. A resource it finds that the manifest does **not** declare is **reported and left standing**, and that report is also how you find out something was created outside cast. It is not an instance wipe and not an environment wipe: the boxes in this fleet are multi-project by design (one of them hosts two third-party client sites), and a delete you can point at a whole box is one wrong argument away from somebody else's production. **What it refuses:** | refusal | why | |---|---| | `--all` | the one verb that must never iterate a fleet. `apply --all` is safe to loop because it is idempotent and never deletes; a loop over a delete has no honest use. | | a read-only instance | the same `COOLIFY_READ_ONLY` assert `apply`/`smoke`/`server add` take. | | an absent project | an absent target reads back exactly like an empty one — and an empty one gives *this* verb a plan that deletes nothing, which renders as a perfectly clean teardown of an environment that is still standing. It names what *is* there instead. | | an environment without `destroy_allowed: true` | below. | | anything but the environment's name, typed | the same ceremony `capture` uses. | | `--with-project`, when anything cast did not declare is still in the project | Coolify refuses that delete too (`400 Project has resources`) — but it refuses it *after* your resources are gone. | There is no `--project`, no `--environment` and no `--resource`. Those coordinates exist to point cast at names **somebody else** chose in a UI, and that is exactly the box a delete must never be aimed at. **The interlock lives in state, not in argv:** ```yaml environments: staging: server: staging-box team: { id: 0, name: Root Team } destroy_allowed: true # absent = destroy refuses. Removed at cutover, forever. ``` A `--yes` flag is not a gate; it is a thing you type without reading, and by the second week it is in the shell history above the command it guards. This is a line a human edits, commits and merges — in the **private state repo**, for the same reason `forbidden_var_patterns` lives there: *a change on one side must not be able to lower its own guard.* It is `true` while an environment is empty and being battle-tested. **The cutover checklist deletes it the moment that environment carries real data**, and from then on destroying it costs a PR. **The plan says what the delete costs.** Coolify's delete takes the resource's volumes with it, so a database line carries its backup schedule and when the last backup actually landed — the difference between *recreate this* and *this is gone*. A backup configuration cast cannot read prints `backup schedule: UNKNOWN` and is treated as unrecoverable; it never rounds down to "none". ``` destroy plan — heavy-duty/incubator staging project: incubator (on Coolify) environment: staging scope: 3 resource(s) the manifest declares for staging, and nothing else DELETE, in reverse dependency order (applications → services → databases): application core a1 database cache d2 backup schedule: NONE — nothing has ever been scheduled for this database. its volume goes with it, and cast cannot bring it back. UNRECOVERABLE. database postgres d1 backup schedule: 0 2 * * * last backup: 2026-07-13T02:00:11Z (success) LEFT STANDING — on this box, and NOT declared by the manifest: service metabase type the environment name to DESTROY the resources above (staging): ``` `--with-project` additionally removes the **environment** and then the **project** — both only if they are empty, and only after Coolify's delete queue has actually drained (its DELETE returns *"deletion request queued"*, not *"deleted"*). It is the way back to zero from a half-applied first run. ## Scripts Operational helpers, all argument-driven (`scripts/`): restore a database backup into a target container. (`register-github-app.sh` is gone — it is `cast github-app register` now.) **They run where cast runs — off the box.** They drive the Coolify API, or reach a box over SSH; none of them expects to be executing *on* a server. Anything that belongs on a box, as root, under a scheduler is [rig](https://github.com/heavy-duty/rig)'s job, not cast's — including the nightly age-encrypted dump of the control-plane database, which is now `rig coolify backup install`. ## Development ```sh npm ci && npm run build && npm test npm run check # biome ```