Manifest-driven Coolify reconciler - rig builds the boxes, cast fills them
Find a file
claude-hdb b07563a815 fix: smoke resolves its target inside the project it was declared under (#29)
`smoke` found the application it WRITES to by name against GET /applications —
every app the token can see, across every project and every environment on the
instance — and took the first name match. So `smoke_target: core` did not name
an application; it named whichever `core` Coolify happened to list first. One
instance carrying prod and staging is enough for `cast smoke --env staging` to
POST its canary vars onto prod's `core`, and on the failure path leave them
there.

It now resolves the target through fetchLive(project, environment) — the same
lookup every read-side verb makes — and takes the coordinates that lookup needs:
--project and --environment, with diff/capture/inventory's semantics and
defaults. An application that is not in that project + environment is not an
empty result, it is the absence of anything to write to, so smoke refuses:
naming what it looked for, where the name came from, and what is actually there
(including when the name belongs to a service or a database, which would 404 on
the /envs endpoint smoke writes to).

The <org>/<repo> positional is now REQUIRED, and the deprecated state-file-scoped
`smoke_target` is dropped: it named an app from a key with no project to scope
to, so it could not be fixed, only carried. It is still declared in the schema —
refused with a migration message rather than a strict-mode "unrecognized key",
because loadBindings runs for every verb and an unmigrated state file must not
take `diff` and `apply` down with it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:08:35 +00:00
.github/workflows feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
bin feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
docs feat: a project registry — the list of what exists (#25) 2026-07-13 20:03:35 +00:00
reference feat: place a resource on a destination — and a state file that can say which (#21) 2026-07-13 19:31:02 +00:00
scripts chore: re-home the control-plane dump to rig 2026-07-12 19:18:46 +00:00
src fix: smoke resolves its target inside the project it was declared under (#29) 2026-07-13 20:08:35 +00:00
test fix: smoke resolves its target inside the project it was declared under (#29) 2026-07-13 20:08:35 +00:00
.gitignore feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
biome.jsonc feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
install.sh fix(install): wire BINDIR onto PATH in the shell profile, not a warning 2026-07-11 15:12:14 +00:00
package-lock.json feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
package.json feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
README.md fix: smoke resolves its target inside the project it was declared under (#29) 2026-07-13 20:08:35 +00:00
tsconfig.json feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00
vitest.config.ts feat: cast — the Coolify executor, extracted from the infra state repo 2026-07-11 12:25:44 +00:00

cast

Point it at a repo and a state directory; it makes a Coolify instance match what the repo declares. One-way, idempotent, never deletes.

Philosophy (shared with rig and claudebox): public tool, private state. cast holds no hostnames, no bindings, no secrets, nothing about your infrastructure. It reads what you point it at and stores nothing, ever.

rig builds the boxes. cast fills them.

Install

curl -fsSL https://raw.githubusercontent.com/heavy-duty/cast/main/install.sh | bash

Needs node >= 22.12 and age (secrets are decrypted by shelling out to it). Re-run any time to upgrade. Unlike rig — which is pure bash so it can run on a bare box — cast runs on your machine: it is an API client, and a server should never install it.

The installer symlinks cast into ~/.local/bin (or /usr/local/bin as root) and, if that directory is not already on your PATH, appends it to your shell profile — .zshrc, .bashrc/.bash_profile, or config.fish, whichever your $SHELL reads — marked # added by cast-install and written only once. The shell you ran the installer from does not inherit it (a curl | bash pipeline is a subshell), so open a new shell or source the profile it names. Set CAST_NO_MODIFY_PATH=1 to be left alone and wire PATH yourself.

The two inputs

cast joins a manifest (what to deploy) with state (where, and with what values). Neither knows about the other, which is the whole point: a manifest can live in a product repo without leaking your infrastructure, and your infrastructure can be re-pointed at a new Coolify without touching a product.

1. The product repo's .infra/ — committed, instance-blind:

.infra/
  manifest.yaml                  # applications, databases, services, per environment
  env/<app>.<env>.env.template   # var NAMES + non-secret values; ${SECRET} placeholders

2. A state directory — private, yours:

environments.yaml               # bindings: the team each env's token must belong to,
                                #   which server it deploys onto, the S3 destination,
                                #   GitHub App name, guards — and, per project,
                                #   the destination it deploys onto + its smoke target
                                #   …plus `projects:`, the registry: which projects
                                #   exist, and in which environments
secrets/<repo>.<env>.env.age    # age-encrypted values for the ${…} placeholders
.coolify.env                    # COOLIFY_BASE_URL + COOLIFY_ACCESS_TOKEN (never commit)
.coolify/<name>.env             # …the same, for a NAMED instance (see below)

Pass it with --state <dir>, or set CAST_STATE. Defaults to the cwd.

Commands

cast apply     <org>/<repo> --env <env> [--path <dir>] [--hostname-overlay <file>]
cast diff      <org>/<repo> --env <env> [--full]
cast capture   <org>/<repo> --env <env> [--generated <NAME>] [--override <NAME>]
cast inventory <org>/<repo> --env <env>
cast server add <name> --ip <ip> --key <file> --env <env> [--user root] [--port 22]
cast smoke     <org>/<repo> --env <env> [--project <name>] [--environment <name>]
cast team [--env <env>]
  • apply — idempotent create-or-update of every manifest resource, then redeploy what changed. One-way: it never deletes a resource that Coolify has and the manifest doesn't. Clones the repo's default branch unless --path points at a local checkout (refused with --env prod — prod always reads the default branch).
  • diff — reports drift, manifest → Coolify. Structural by default; --full also compares env vars. Exits non-zero when dirty, so CI can gate on it.
  • inventory — what is actually on a box. With no repo it sweeps the instance (every project, every environment, every resource — no manifest involved); with a repo it reconciles, showing resources and env var keys (never values) sorted into on-both / manifest-only / box-only. Needs no store, no age key, and no recipient — it runs before adoption, which is the point of it. A document, read by a person; nothing here is consumed by apply. See Adopting a hand-built instance.
  • capture — the adoption path: reads a hand-built instance's live env and writes the environment's age store from it. See Adopting a hand-built instance below.
  • server add — uploads a server's private key and registers it with Coolify.
  • smoke — contract test against the project's smoke_target: proves Coolify's bulk env endpoint still upserts rather than replacing. Run it after every Coolify upgrade — apply's never-delete guarantee rests on that behavior, and the published OpenAPI does not describe it accurately. It writes (two canary env vars onto that one application, then deletes them), so the repo is required: the target is resolved inside the project and environment it was declared under, with --project / --environment if the box names either differently, and it refuses rather than guessing when no application of that name is there. A bare app name is unique nowhere else — one instance carrying prod and staging is enough for the first core on it to be prod's.
  • team — prints the team the configured token acts as. With --env, also checks it against that environment's team: binding and exits non-zero on a mismatch — the dry run for "would apply refuse?", answered without touching anything.

Every command that reaches a live Coolify takes an --env, because every one of them first asserts the token's team (below).

--hostname-overlay swaps domains for a pre-flight run against temporary hostnames; re-applying without it is the cutover.

Cloning: cast authenticates, and never prompts

apply, diff and capture clone the product repo (unless --path points at a local checkout — refused for prod, which always reads the default branch). For a private repo that needs credentials, and cast resolves them itself:

  1. gh, borrowed as a credential helper for that one invocation — it does not touch your global git config.
  2. GITHUB_TOKEN / GH_TOKEN from the environment (the CI path).
  3. Whatever git's own credential helper does, if you have one.

Being logged into gh is enough. You do not need gh auth setup-git — that separate act is what wires git's helper, and not running it is exactly how you end up at git's interactive username/password prompt, which GitHub no longer accepts. cast sets GIT_TERMINAL_PROMPT=0 on every path, so it can never hang there or hide a credentials failure behind an error about the repository. With no credentials at all it says so, and names the fix.

The token is never put in the clone URL or in http.extraheader — both leak it into ps, and the latter persists it into the clone's git config.

Many Coolifys

--instance <name> reads <state>/.coolify/<name>.env instead of <state>/.coolify.env. Every verb that reaches Coolify takes it.

cast diff heavy-duty/incubator --env prod --full --instance legacy

An environment can bind one, so --env selects the right control plane with no flag at all:

environments:
  prod:
    server: prod-box
    team: { id: 1, name: heavy-duty }
    instance: prod-cp        # → <state>/.coolify/prod-cp.env

An explicit --instance still wins, so a one-off read against a legacy box needs no edit to that file either. With no flag and no binding, nothing changes.coolify.env is read exactly as before.

Two properties, both deliberate:

  • An unknown --instance refuses, and names the instances that do exist. Falling back to the default is how a diff meant for a legacy box gets run against production.
  • An instance may declare COOLIFY_READ_ONLY=true, and then apply, smoke and server add refuse it — before their first call, and even though the token itself would permit the writes. That turns "I pointed the wrong token at the wrong box" from a live incident into an exit code.

Every command that reaches a Coolify now says which one, next to the team assert. It is the most consequential input to any run, and the least visible.

Adopting a hand-built instance

cast is otherwise scoped to the steady state: manifest → Coolify, forever. Adoption is the one way in, and it has two verbs and a fixed order:

inventory → you read it → a manifest PR → captureapply

Look before you adopt. A box nobody declared does not use your vocabulary: its project is called whatever someone typed, its environment is Coolify's default (production, not prod), and its resources are named by whoever clicked New Resource that afternoon. inventory shows you both sides at once, so those differences arrive together, as a document — instead of one at a time, as refusals from a verb that is already halfway through a migration.

First, sweep it — you cannot aim at coordinates you do not have yet:

cast inventory --env prod --instance legacy
sweep — instance legacy (https://coolify.example.com)

  Incubator
    production     (empty)
    staging        2 applications, 2 databases, 1 service
      application  Incubator Stack v2
      application  Incubator Landing
      database     Incubator Database v2
      …
  La Familia Site
    production     1 application
      application  lafamilia-web

Note what that costs you to not have: Coolify auto-creates a production environment in every project, so the obvious guess is empty and the live system is somewhere else entirely — under a name someone typed, in a project you may not have known was there. An environment with zero resources is far more often the wrong coordinate than an empty one, and inventory says so rather than quietly reporting that the manifest has five things the box lacks.

Then reconcile, against a target you now know exists:

cast inventory heavy-duty/incubator --env prod --instance legacy \
  --project Incubator --environment staging \
  --resource core="Incubator Stack v2"

It never reads a value, needs no store and no key, and its output is not desired state. What you do with it is decide, resource by resource and key by key, what the manifest should gain and what is cruft that must not travel — and land that as a manifest PR. Only then:

CAST_CAPTURE_ADMIN_EMAIL=me@example.com \
  cast capture heavy-duty/incubator --env prod --instance legacy \
    --override ADMIN_EMAIL

It reads the required secret names from the manifest's own env templates (the ${…} refs — the manifest already declares exactly this set), reads the live values off the instance, and classifies every name:

captured found live, value taken
generated the manifest's generated_secrets declares it provider-made → written as the literal pending-coolify-generated, never the live value
overridden supplied by you, for a value that must not be carried over
missing required by a template, absent live → refuses

Then it prints a plan of names and provenance — never values — and waits for you to type the environment's name.

Before any of that, it checks that the resources the manifest names exist. An absent resource reads back exactly like one with no env vars set: every name it declares reports missing, and --override would then have you hand-carry values that are sitting right there under a different name — writing a perfectly valid store while the actual finding (the manifest and the box disagree about what this thing is called) is never discovered. So a resource that isn't there refuses, and names what is. inventory is how you reconcile it.

Three names that are not yours

A hand-built box names things without asking you, at three levels, and cast takes each as a coordinate to read with — never as a reason to rename anything of yours:

flag when
--project <name> the project isn't named after the repo (Incubator, not incubator)
--environment <name> the environment isn't named after --env (Coolify's default is production, not prod)
--resource <manifest>=<live> a resource isn't named after the manifest's (core is Incubator Stack v2 over there). Repeatable

None of them is ever a manifest field: they are arguments to a single run, because a manifest that recorded a legacy box's names would carry a dead machine's vocabulary forever.

--project and --environment are how a verb that must find a target says where to look — diff, capture, inventory, apply, and smoke, which resolves its smoke_target in exactly that project and that environment, and refuses when it is not there (#29). --resource is read-side only (diff, capture, inventory): apply refuses it outright, because it creates resources under the manifest's own names — an alias there could only mean adopt the existing one instead, which is a different operation and would otherwise silently create a duplicate beside the resource you were pointing at.

--env stays ours: it selects the manifest block, the environments.yaml binding, the age key, the store path. --environment is theirs, on the wire, and nothing else. Collapsing the two lets a box that is being deleted next week name the environment of the box that replaces it — apply creates the environment from that value, so it would be inherited permanently.

The mapping is not mechanical, and that is the whole design. A DATABASE_URL copied off the source box points at the source box's Postgres: confidently wrong, entirely plausible, and the target's real URL does not exist until Coolify creates the resource. So the manifest declares those names, and cast placeholds them:

environments:
  prod:
    generated_secrets: [DATABASE_URL_PROD, REDIS_URL_PROD, UMAMI_DATABASE_URL]

It is a manifest property rather than a flag you have to remember, because the manifest is what knows DATABASE_URL comes from a database it declares. (A generated_secrets entry no template refers to is a schema error — a guard standing over nothing is worse than no guard, because it reads like one. --generated <NAME> covers a manifest that hasn't declared them yet.)

An --override's value is read from $CAST_CAPTURE_<NAME>, never from the command line: argv is visible in ps to every process on the box. It exists for values that must not survive the copy — staging and prod sharing a Mailgun domain means a staging box carrying the real ADMIN_EMAIL can mail real users.

The store is encrypted to the environment's age_recipient (add it to environments.yaml — it's the public half, safe to commit). Plaintext goes to age on stdin: it is never a temp file, never on stdout, never in your shell history. An existing store is not overwritten without --force.

docs/semantics.md is the contract behind those commands: what apply guarantees (never deletes, never recreates a database, fails loudly rather than recreating on un-updatable drift), the dockercompose build pack, the hostname-overlay shapes, and the places Coolify 4.1.2 does not cooperate — each citation verified against coollabsio/coolify v4.1.2 and the vendored OpenAPI in reference/. Read it before changing apply.

Secrets, and attended applies

An environment's age identity is resolved in exactly two ways:

  1. $CAST_AGE_KEY_FILE_<ENV> — injected for this invocation
  2. ~/.config/cast/age-<env>.key — a standing key on this machine

That is the whole mechanism behind attended vs unattended applies: an environment whose key you never leave on disk can only be applied by someone who injects it. Keep a standing key for staging if you like; keep prod's in a password manager and pass it per apply.

The state directory holds ciphertext. It must never hold the identity that opens it.

Teams: the one assert cast makes before it touches anything

Coolify API tokens are team-scoped, and a token pointed at another team's resources does not error. The API resolves what the token cannot see to null — and to a tool like cast, null is indistinguishable from "this resource does not exist yet", which is an invitation to create it. An apply run with a wrong-team token would not fail; it would silently provision a duplicate set of resources into the wrong team, on whatever server that team owns. Silent, mutating, discovered late.

So every environment declares the team its token must belong to, and cast refuses to do anything at all until it has checked:

environments:
  prod:
    server: prod-box
    team: { id: 1, name: heavy-duty }

Give id, name, or both — both are compared when both are given. id is the true identity (names can be renamed); name is what makes the file readable. Run cast team to print the values for the token you currently have configured.

The check is fail-closed: an environment with no team: is one whose token cannot be verified, so it is a schema error, not a warning. It runs before the first read, not merely before the first write — an unasserted diff against the wrong team would report "everything is absent", which is precisely the lie that an apply would then act on.

Nothing below the team scopes a token. A Coolify environment has no team of its own (it hangs off a project) and no API path scopes by one: Coolify environments are an organizational construct, not an auth boundary. The team is the only boundary there is, so it is the one cast asserts.

The registry: which projects exist

environments.yaml says where things deploy to, and how a project you have already named is placed once it is there. Until the projects: block, nothing in it said which projects exist at all — "every project" was a thing the operator remembered:

projects:
  heavy-duty/incubator:
    environments: [prod, staging]
  acme/client-site:
    environments: [prod]

Keyed by the full <org>/<repo> slug, and the key is the repo — there is no repo: field inside, because a second place to write the same string is a second place for it to be wrong. Unlike github_apps, a bare <repo> key is refused rather than resolved: this block is new, so it has no state files in the wild to keep working, and a bare <repo> is unique only within an org — which is exactly why it is not a key. environments: lists our environment names (the values --env takes), never Coolify's.

The block is optional; a state file written before it loads unchanged.

It has to be true, so cast checks that it is — at parse time, for every verb. Two ways it could quietly stop being true, both refused:

  • an environment name that no environments: block defines (a typo). The project is real and its environment imaginary, so a fleet run visits nothing for it, reports nothing, and exits clean.
  • an environments.<env>.projects.<repo> binding — a destination, a smoke target — in an environment the registry does not register that project for. The two blocks then describe two different fleets: state real enough for a direct cast apply to use, invisible to every fleet run. (Checked only when projects: is present.)

Both refusals defend one failure: a silently skipped project reads exactly like a clean one. Silence is the one report that must never be ambiguous.

What it unlocks, neither of which was possible without a list to iterate:

  • fleet operationscast diff --all / apply --all over every project in an environment.
  • rebuild-from-state — "restore this Coolify from the state repo" cannot even be attempted without knowing what was on it. The registry is the difference between a documented recovery and an archaeology exercise.

Two projects, one box: destinations

A destination is the Docker network a resource is created on. A server has a default one, and while a server hosts a single project that default is the right answer — which is why cast went so long without naming it.

The moment a server hosts two projects, it stops being: they share one network, and "isolated" becomes a thing you believe rather than a thing that is true. So the destination is declared per project, inside the environment — an environment-scoped key could not express it, because server: is precisely the thing two projects share:

environments:
  prod:
    server: shared-box
    team: { id: 1, name: heavy-duty }
    projects:
      heavy-duty/incubator:
        destination_uuid: <uuid>     # the network THIS project's resources go on
        smoke_target: core           # the app `cast smoke` writes its canary to
      acme/client-site:
        destination_uuid: <other>

Keyed by repo, full <org>/<repo> slug first, exactly like github_apps — a bare <repo> key still resolves, so existing state files keep working. Both fields are optional, and an environment whose server hosts one project needs neither.

A UUID and not a name, unlike server: right above it. Coolify 4.1.2 has no destinations API whatsoever — no list, no read, nothing — so there is no name for cast to resolve. You read the UUID out of the Coolify UI, the same way you do for s3_destination.

What cast can and cannot promise here. It sends destination_uuid on create, for applications, databases and services alike. It can never check it afterwards: Coolify takes a UUID on write and hands back an integer destination_id on read, and nothing maps between them. So diff does the one honest thing left — it groups the live resources by the id Coolify does report, and a project whose resources do not all share one network is drift:

split placement: these resources sit on 2 different destinations
  destination 1: application landing, database postgres
  destination 4: application core
  a project's resources must share one destination — that is what the isolation IS.
  apply never moves a live resource between networks: resolve manually (runbook act).

…and when you declare a destination, every diff says, out loud, that it did not verify it. That is deliberate. A setting that reads back as absent rather than wrong is the failure this whole file keeps trying not to be.

One sharp edge worth knowing: on a server with exactly one destination, Coolify ignores the destination_uuid you send and never validates it — a typo there is invisible until a second destination exists. On a server with more than one, a create that omits it is a hard 400, which is why cast could not deploy onto a shared box at all until it could send this. Details, with citations: reference/README.md.

Guarding an environment

An environment may refuse variables by name pattern:

environments:
  prod:
    server: prod-box
    team: { id: 1, name: heavy-duty }
    forbidden_var_patterns: ["^ALLOW_"]

apply then refuses if any such var is present on any resource, regardless of value. ALLOW_SEED=false still fails: a var that exists can be flipped on later in the Coolify UI without touching a manifest, so "off" has to mean absent.

This guard lives in your private state deliberately — not in the product's manifest. A product-side change must not be able to lower its own guard.

Scripts

Operational helpers, all argument-driven (scripts/): register a GitHub App with Coolify, restore a database backup into a target container.

They run where cast runs — off the box. They drive the Coolify API, or reach a box over SSH; none of them expects to be executing on a server. Anything that belongs on a box, as root, under a scheduler is rig's job, not cast's — including the nightly age-encrypted dump of the control-plane database, which is now rig coolify backup install.

Development

npm ci && npm run build && npm test
npm run check          # biome