Commit graph

70 commits

Author SHA1 Message Date
claude-hdb
d927a7ad57 fix(diff): converge live projection on compose domains and is_static
Two distinct root causes made `cast diff`/`apply` re-diff and redeploy an
application on every run against a live Coolify 4.1.2 (cast#68).

1. `docker_compose_domains` parse bug. Coolify 4.1.2 returns this field as a
   JSON-encoded, service-KEYED object ({ "<svc>": { "domain": "<comma-joined>" } }),
   not the [{name,domain}] array cast expected. The array-only parser bailed to
   `undefined`, so cast diffed the desired map against nothing forever.
   `parseDockerComposeDomains` now decodes the real object shape into the
   internal {service: string[]} map while still tolerating the legacy array
   shape (the write-side round-trip). Malformed/scalar/empty still → undefined.

2. `is_static` unreadable. Coolify 4.1.2 returns `is_static: null` on the read
   path even for a genuinely-static app, so projecting `false` diffed false→true
   forever. `projectLiveFields` now omits is_static when the live value is
   null/absent; `fetchLive` flags the app `staticNotCompared`; `computeDiff`
   skips the comparison (once-per-run warn), degrading is_static to a
   create-time-only setting. A real live boolean is still projected and diffed.

Closes #68

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 16:23:27 +00:00
Daniel Marin
95c9befd2a
Merge pull request #67 from claude-hdb/feat/derive-domain-refs
feat(resolve): derive base-URL env vars from manifest domains via ${domain:...} (#66)
2026-07-15 13:43:49 +01:00
claude-hdb
6f0b0f6fc1 feat(resolve): derive base-URL env vars from manifest domains via ${domain:...} (#66)
A public base URL an app reads (LANDING_BASE_URL, ADMIN_WEB_BASE_URL) is a
fact the manifest already states in `domains`/`service_domains` — the same
fields cast parses to reconcile Coolify domains. Hand-transcribing it into an
env template is a second copy that drifts (incubator's prod LANDING_BASE_URL
silently kept a pre-apex host). So a template can now say it directly:

    LANDING_BASE_URL=${domain:landing}
    ADMIN_WEB_BASE_URL=${domain:core.admin}

- ${domain:<app>}            -> applications.<app>.domains[0]
- ${domain:<app>.<service>}  -> applications.<app>.service_domains.<service>[0]

Symmetric with ${resource:...} (#60) — parse -> sentinel -> validate -> fill —
but a domain is PURE MANIFEST DATA, known at plan time, so it resolves fully in
desiredFromManifest against a map built from the manifest: no live read, no
executor deferral, no unresolved-at-write path. Domains are PUBLIC, so they
resolve to secret:false (printed in diffs) and read as plain literals
downstream (no diff.ts change). Not secrets: excluded from templateRefs, never
captured. assertDomainRefs is the single validation gate (apply/diff/capture),
refusing an undeclared app/service, a wrong-shape ref, or an empty/blank domain
list before the sentinel can escape. Applications only (Coolify 4.1.2 can't set
service domains). REPORTING_TZ-style operator literals stay literal.

Closes #66.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 12:02:33 +00:00
Daniel Marin
35d147f6c4
Merge pull request #65 from claude-hdb/feat/derive-resource-url
feat(resolve): derive DATABASE_URL/REDIS_URL from the database cast created (#60)
2026-07-15 00:49:49 +01:00
Daniel Marin
9084713a5d
Merge pull request #64 from claude-hdb/fix/static-build-fields
fix(apply): express static-site build settings so a monorepo app is served, not run (#63)
2026-07-15 00:49:30 +01:00
claude-hdb
ea526d6598 feat(resolve): derive DATABASE_URL/REDIS_URL from the database cast created (#60)
Add a ${resource:<name>.url} env-template ref that resolves to the internal URL
of a database the same manifest declares, read back from the live resource's
internal_db_url — never stored in the age store, never decrypted, never printed.
This deletes the two-pass generated-secret bootstrap for a database's own URL
rather than automating it: no placeholder, no stored copy to drift or overwrite,
and a rotated password is simply followed on the next apply.

Resolution runs in one function (fillDerivedEnv) against two URL maps: at diff
time against databases already on the box (so a matching app shows no drift —
killing the "secret DATABASE_URL differs" noise that ran on every plan), and in
the executor at apply time against a database created earlier in the same run
(the from-nothing case; apply acts databases-before-applications, #45). The
unresolved sentinel is never written — the executor refuses, rather than write a
blank that boots the app pointed at nothing, and re-running once the database is
up resolves it as an ordinary update.

A ${resource:X.url} naming a database the manifest does not declare, or an
attribute other than .url, is a hard plan-time error refused by every verb that
opens a template (apply, diff, capture). generated_secrets and the two-pass
bootstrap remain for the residual class — a provider-generated value that
genuinely is not derivable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 23:46:26 +00:00
claude-hdb
cd1864aaad fix(apply): express static-site build settings so a monorepo app is served, not run (#63)
apply created applications but dropped install_command, build_command, and
is_static — settings the manifest had no field for — so a static site in an
npm-workspace monorepo (landing) was built and RUN from the repo-root
package.json, booting the core API server, which crash-looped on a missing
DATABASE_URL.

The build block gains install_command / build_command / start_command
(free-form strings) and static (-> Coolify is_static). apply writes and diffs
them; draft emits them (they left its NO_HOME list, and is_static was never in
it — the silent loss that caused the crash), and only emits static alongside a
publish_directory so a draft always loads.

Managing is_static is opt-in: declaring `static:` is required to serve a static
app, and NOT emitting is_static by default avoids the first apply PATCHing
static serving OFF on an un-migrated app (or fighting a pack:static coupling
forever). static:true with no publish_directory, and any of the four on a
dockercompose app, are parse-time refusals.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 23:46:14 +00:00
Daniel Marin
8f0abe2ab8
Merge pull request #61 from claude-hdb/feat/backup-schedule-diff
feat: diff and apply a database's backup schedule (#51)
2026-07-15 00:12:24 +01:00
claude-hdb
d8d5cf8397 feat: diff and apply a database's backup schedule (#51)
Backup schedules were write-only, filed under "known limitations" on the
claim that "live Coolify state doesn't expose it back". The parenthesis was
load-bearing and false: a schedule is not on the database's own GET, but it
was never meant to be — it has its own route, GET /databases/{uuid}/backups,
which cast had been POSTing to all along and had simply never read.

The cost was exact. A database created before its `backup:` block was
declared never got one (apply set the schedule only inside the create
branch); a schedule deleted in the UI was invisible; and the `--full` diff
that gates a production cutover passed with an unbacked-up production
database.

Shape settled from the source rather than the vendored spec, which documents
the body as "Content is very complex. Will be implemented later.":
DatabasesController@database_backup_details_uuid (v4.1.2) returns a raw
Eloquent collection — a JSON array of ScheduledDatabaseBackup rows, columns
per $fillable (uuid, enabled, frequency,
database_backup_retention_amount_locally). `frequency` round-trips verbatim:
the controller validates it and stores $request->only(...) unchanged, with no
mutator on the model. The "diffing it would flag spurious drift" fear was a
guess about a read nobody had performed.

- `backup` becomes a diffed field like any other (resolve.ts), replacing the
  side channel that carried it around the diff.
- The live side reads the route (coolify.ts, fetchLive), and apply sets the
  schedule on UPDATE as well as create — POST or PATCH, decided by a read.
- A disabled schedule is a row that backs nothing up: neither clean nor
  absent. cast diffs it and re-enables it.

Degrades honestly, since no live box was probed: an unreachable or
unrecognized response can only ever produce "declared, NOT compared — verify
in the Coolify UI", never invented drift and never a clean bill on an
unread database. On the write side the same failure raises rather than
guessing — POSTing blind would duplicate a schedule that may already exist.
2026-07-14 23:07:54 +00:00
Daniel Marin
0fecfbf5d7
Merge pull request #55 from claude-hdb/fix/generated-secret-guard
fix(apply): refuse to write the generated-secret placeholder over a live value (#47)
2026-07-15 00:04:24 +01:00
claude-hdb
bab33b1e6f fix(apply): refuse to write the generated-secret placeholder over a live value
The bootstrap is two-pass and only the first pass was ever safe to repeat.
The store holds `pending-coolify-generated` for a provider-generated secret;
the first apply sends it, Coolify creates the Postgres/Redis and replaces it
with the real URL. From that moment the store is known-wrong — and `diff` and
`apply` had never heard of the literal cast itself invented to say so.

`diff` printed `secret DATABASE_URL differs`, which is word for word what a
legitimate rotation prints, and `apply` stood ready to PATCH the placeholder
back over the live URL and redeploy every consumer onto it. Coolify's bulk env
endpoint is a plain upsert (create_bulk_envs, v4.1.2: an existing key is found
and its value overwritten), so nothing on the far side stopped it either.

- diffEnv gives the placeholder its own state, `placeholder-conflict`, when the
  store holds it and the live resource holds anything else. Live-also-
  placeholder, absent live, and the create path are unchanged.
- renderDiff says it in words no rotation prints, and counts it in the summary.
- applyPlan REFUSES on it, before any resource is touched — same fail-closed
  shape as the not-updatable refusal. The message names the key and the
  resource, never the live value, and points at the remedy (#48).

Keyed on the store's VALUE, not the manifest's `generated_secrets:` list: that
list names store refs (DATABASE_URL_PROD) while an env diff is keyed by env var
key (DATABASE_URL). Matching the list against these keys would have sailed past
the very case that motivated the issue.

Closes #47.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:56:27 +00:00
Daniel Marin
0764f0018e
Merge pull request #56 from claude-hdb/fix/reserved-env-names
fix: never write an env var whose name Coolify injects itself (SOURCE_COMMIT, COOLIFY_*)
2026-07-14 23:54:15 +01:00
Daniel Marin
90d95321eb
Merge pull request #59 from claude-hdb/feat/destroy
feat(destroy): cast destroy — a scoped, state-gated teardown verb (#43)
2026-07-14 23:53:59 +01:00
claude-hdb
6849d29f0e feat(destroy): a scoped teardown verb, gated in state (#43)
`apply` fails closed on an immutable field with "resolve manually" — which
meant a hand deletion in the Coolify UI, unscoped and unconfirmed, against an
instance whose token can see every project on it. That is how the wrong project
gets deleted.

`cast destroy <org>/<repo> --env <env> [--with-project]` is that act, scoped:

- MANIFEST-SCOPED. It deletes the resources the manifest declares in that
  project and that environment, in reverse dependency order (applications →
  services → databases). Anything else it finds is reported and LEFT STANDING —
  that report is how a resource created outside cast gets discovered, and the
  boxes in this fleet are multi-project by design.
- Not a flag on apply. `apply never deletes` is the invariant that makes it safe
  to run on a schedule; apply.ts and diff.ts are untouched.
- REFUSES rather than no-ops: --all (always), a read-only instance, an absent
  project (D-237 — an absent target must never read as a clean empty plan), a
  manifest that declares nothing this environment holds, and --with-project
  while anything undeclared is still in the project.
- The prod interlock lives in STATE, not argv: environments.<env>.destroy_allowed
  in environments.yaml, absent = refuse. A flag is a thing you type without
  reading; this is a line a human edits, commits and merges.
- The plan says what the delete COSTS: every database line carries its backup
  schedule and when the last backup landed. A backups route cast cannot read
  prints UNKNOWN and is treated as unrecoverable — it never rounds down to NONE.
- Last gate: the environment's name, typed (capture's ceremony).

Coolify's DELETE query params are sent explicitly (all four default to true):
delete_volumes, delete_connected_networks, delete_configurations — and
docker_cleanup=FALSE, because that one prunes the whole SERVER, and these boxes
host other people's production.
2026-07-14 22:52:17 +00:00
Daniel Marin
d5d984f631
Merge pull request #58 from claude-hdb/feat/capture-generated-only
feat(capture): --generated-only, the bootstrap's missing pass 2
2026-07-14 23:49:49 +01:00
Daniel Marin
0417177b00
Merge pull request #62 from claude-hdb/fix/stale-compose-fixture
test(resolve): fix stale compose_file fixture that reddens main after #49+#46 merge
2026-07-14 23:49:35 +01:00
claude-hdb
a060444520 test(resolve): absolute compose_file in the source-commit fixture
#54's fixture predates #49's leading-slash refinement (both cut from the
same base); merged together the fixture violates the new rule and reddens
main. One character — the product change and its own tests are unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:46:06 +00:00
Daniel Marin
97c1db35f8
Merge pull request #57 from claude-hdb/fix/domain-preflight
fix(apply): pre-flight domain uniqueness, and translate Coolify's 409 (#44)
2026-07-14 23:43:17 +01:00
Daniel Marin
ab0eb27449
Merge pull request #53 from claude-hdb/fix/create-order
fix(apply): create databases and services before the applications that need them (#45)
2026-07-14 23:42:46 +01:00
Daniel Marin
7c6092dd87
Merge pull request #54 from claude-hdb/feat/source-commit-notice
feat(resolve): warn that apply cannot enable "Include Source Commit in Build" (#46)
2026-07-14 23:42:29 +01:00
Daniel Marin
73a92254b4
Merge pull request #52 from claude-hdb/fix/compose-file-path
fix(manifest): refuse checkout paths that are not absolute (#49)
2026-07-14 23:41:59 +01:00
claude-hdb
5e10375837 feat(capture): --generated-only, the bootstrap's missing pass 2
A manifest that declares `generated_secrets:` bootstraps in two passes by
construction: pass 1 `capture` placeholds those names (their values do not
exist yet), `apply` creates the database and Coolify generates the real URL —
and nothing then taught the store that value. The operator did it by hand:
decrypt a fourteen-name store, edit two lines, re-encrypt to the environment's
age recipient, against production, holding the prod key.

`capture --generated-only` inverts capture's disposition rule and changes
nothing else — it fills the generated names and leaves every other name in the
store exactly as it is, byte for byte. Same verb, same ceremony, same
store-writing code path.

- reads the value from the resource that OWNS it (`internal_db_url` on the
  database), never from a consuming app's env, where a generated URL never
  appears — the app's env holds the placeholder itself at this point.
- resolves the database inside the project+environment via
  GET /projects/{uuid}/{env}, never the instance-wide GET /databases (which
  lists other projects' databases and umami's bundled Postgres — #29's bug in
  another hat). The scoping is structural, not a filter.
- refuses to guess which database a name comes from: nothing in the manifest,
  the templates or the box carries that edge, so it infers only when it cannot
  be wrong (one name, one database) and otherwise hands back `--from`.
- refuses to overwrite a generated name holding a real value without --force
  (a silent credential rotation), a name absent from the store, and a
  placeholder nothing fills.
- asserts the postcondition against the ciphertext on disk: zero
  pending-coolify-generated remain, and the name count is unchanged. That
  assertion was a line in a human runbook.

`apply` deliberately does NOT do this after a create — it would make the verb
that mutates Coolify also mutate the encrypted store, and hence the git repo.

Closes #48.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:32:07 +00:00
claude-hdb
2e201fb58a fix(apply): pre-flight domain uniqueness, and translate Coolify's 409 (#44)
Coolify enforces domain uniqueness across the whole instance; cast plans
inside one project + one environment. So apply could produce a plan that
was internally consistent, correct against everything cast can observe,
and still be refused — by a resource in a project cast never queries,
arriving as a raw 409 mid-apply, after the project and the environment
had already been created.

- Pre-flight the create plan: before the first write (project and
  environment are created lazily, by the first create), check the domains
  the plan is about to claim against GET /applications. A conflict is now
  a refusal that costs nothing, not a half-applied run. One GET, and only
  on a plan that creates an application with a domain — a first apply.
  Covers both live shapes: fqdn, and per-service docker_compose_domains.
- Translate the 409 when one gets through anyway (a conflict with a
  service fqdn or the instance fqdn is not visible in GET /applications,
  so the pre-flight is a strict subset of Coolify's check). Names the
  domain, the resource, its uuid — and whether it is outside the applied
  project, which is the part the operator cannot get from Coolify.
- Never send force_domain_override=true. Coolify suggests it in the error
  text; two resources on one domain is a routing coin-flip, and Coolify
  says so in the same response.

Closes #44.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:30:34 +00:00
claude-hdb
965541bbc1 fix: never write an env var whose name Coolify injects itself (#50)
Coolify injects SOURCE_COMMIT and the COOLIFY_* family into an application's
runtime environment itself, and SKIPS its own injection of a name the resource
already carries a var of (ApplicationDeploymentJob.php v4.1.2, line 2994 —
`->where('key', 'SOURCE_COMMIT')->isEmpty()`). A resource-level var of that
name therefore SUPPRESSES the platform's value. An empty one suppresses it
just as completely: presence, not value.

And it fails green — the deploy succeeds, health checks pass, and the only
symptom is /version reporting "unknown", the endpoint a production cutover is
gated on (D-266).

The rule is now a property of cast, not of one code path. A new src/reserved.ts
owns it, and every place cast touches an env var honors it:

- resolve — every manifest read (desiredFromManifest, requiredSecrets,
  manifestResources) refuses a template declaring a reserved name, before any
  write. So apply, diff, capture and inventory all refuse identically.
- draft — a reserved name read off a live box gets its own provenance,
  `suppressed`: out of the template, out of the age store, its live value read
  into no artifact, and named in UNCAPTURED.md with the consequence.
- diff — promoted out of the remove-candidate orphan list ("apply never removes
  these; read them by eye") and printed as a FINDING with its consequence. Not
  clean. apply still never deletes: cast reports, the human removes it.
- capture (classify) and cli (syncEnv) carry the same assertion at the file and
  at the wire — unreachable through the CLI today, and kept because the
  invariant is "cast never writes one", not "the CLI happens to check first".
- smoke writes an env var too; its probe names are asserted outside the space.

The rule lives in cast's code, NOT beside forbidden_var_patterns in private
state: that one is policy an environment may set for itself, this one is a fact
about Coolify, true on every box — nothing a manifest change could lower.

19 tests in test/reserved.test.ts, one per path.

Closes #50.
2026-07-14 22:29:23 +00:00
claude-hdb
93bca37890 feat(resolve): warn that apply cannot enable "Include Source Commit in Build"
Coolify 4.1.2 gates the SOURCE_COMMIT *build arg* behind a per-application
setting (ApplicationSetting.include_source_commit_in_build, default false)
that has no API surface: it appears in zero API controllers, and both the
create (l.914) and PATCH (l.2368) allowlists in ApplicationsController.php
reject unrecognized keys outright ("This field is not allowed."), so it
cannot be smuggled through — sending it would 422 the whole request. Its
only writer in v4.1.2 is the Livewire Advanced tab (Advanced.php:128), i.e.
a human in the UI.

So apply says it out loud, once per dockercompose application, via the same
desiredFromManifest mechanism and in the same voice as the existing umami
service-domains warning. A manual step the tool knows about and does not
mention is a manual step that gets forgotten — and this one fails green.

Note the toggle gates the BUILD-time arg only; Coolify's runtime injection
of SOURCE_COMMIT is unconditional (ApplicationDeploymentJob.php:2949), so
a service reading process.env.SOURCE_COMMIT per request never needed it.
The warning therefore does not repeat #46's original (incorrect) premise
that this toggle is why /version reported sha "unknown" — the real cause is
an app-level env var suppressing Coolify's own injection (#50).

Closes #46.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:24:10 +00:00
claude-hdb
21070636bc fix(apply): create databases and services before the applications that need them (#45)
apply walked report.changes in manifest order — applications, then databases,
then services (desiredFromManifest) — so a first apply created AND deployed
the compose app before the Postgres and Redis it talks to existed at all: a
full build and a guaranteed-red deployment on every first run.

Sort the walk by a fixed kind-order instead: databases -> services ->
applications. It cannot be a computed graph — nothing in a manifest declares
that `core` needs `postgres`, no resource names another — so the order is a
constant (KIND_ORDER), ranked as a Record<ResourceKind, number> so a fourth
kind fails the build rather than silently sorting ahead of databases.

Creates and updates alike: an application restarted against a database whose
own pending change has not landed is the same failure one apply later. The
report is copied, never sorted in place — renderDiff and the fleet summary
still read in manifest order, and clean/orphans/placement are untouched. Only
WHEN apply acts changes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:22:30 +00:00
claude-hdb
9a1b7003b7 fix(manifest): refuse checkout paths that are not absolute (#49)
Coolify 4.1.2 validates every checkout-relative path on create and 422s
anything without a leading slash: `docker_compose_location` against
ValidationPatterns::FILE_PATH_PATTERN, `base_directory`/`publish_directory`
against DIRECTORY_PATH_PATTERN. cast passed all three through verbatim and
validated nothing about their shape — and docs/semantics.md taught the exact
value Coolify rejects (`compose_file: docker-compose.yaml`).

The 422 is a property of the manifest, not of the instance, so it is knowable
before a single API call. Refine at parse time, on every verb: a refusal that
names the fix, not a silent normalization — the manifest is the artifact under
review, so a bad value is fixed in the file, once.

Fixes the docs example and the fixtures that carried the broken value.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 22:20:27 +00:00
Daniel Marin
7cc3a3b64e
Merge pull request #42 from claude-hdb/fix/dest-and-default-env
fix: the first apply against a fresh multi-destination box (#40, #41)
2026-07-14 18:30:59 +01:00
claude-hdb
1af5eeba0a fix: the first apply against a fresh multi-destination box (#40, #41)
Both of these were found by the same run — the genuinely-from-nothing apply that
#38 was also hiding in, against a box that shares its server with another project.
Neither is a bug in what apply DOES; both are bugs in what it leaves behind and
what it says.

#40 — cast removes the default environment it made Coolify create.

POST /projects hands a new project Coolify's OWN default environment, `production`.
#39 taught apply to create the environment its resources actually name, so a project
cast creates from nothing now ends up carrying two: ours, holding everything, and an
empty `production` that nothing will ever use. That is precisely the shape that makes
a box unreadable later, and we have the live example — on the box being migrated away
from, `production` is empty and everything runs in `staging`, and "the obvious guess
is the wrong one" is a note we had to write down for ourselves. Shipping more of those
is not neutrality.

This is the only delete cast performs, so it argues for itself against apply-never-
deletes: what that rule protects is things cast did not make, and this is a byproduct
of cast's own POST /projects seconds earlier, holding nothing and having never held
anything. Three conditions, jointly, or nothing is touched — cast created the project
in THIS run (never a project someone built by hand), the environment is EMPTY (asked
of Coolify via the details route, the only one that eager-loads resources — not
inferred from the first condition), and its name is NOT ours (an --environment
production keeps its production, since that is where everything is about to live).
Best-effort: a delete that fails is reported and never fails an apply that worked.

#41 — the multi-destination 400 says what to do, and the plan says what it assumed.

A create against a server with more than one destination that names none is rejected
with "Server has multiple destinations and you do not set destination_uuid." — a
message that names neither the remedy nor the file it goes in, arriving at the FIRST
create, after apply has already made the project and the environment.

cast cannot pre-flight it and that half is not fixable: 4.1.2 serves no destinations
API at all, and GET /servers/{uuid} does not carry them either, so a server's
destination COUNT is unknowable until a create has been attempted. The diagnosis is
what is fixable. The 400 is now answered with the failing resource, the server by the
name the operator wrote (not its UUID), the exact path the UUID goes in
(environments.<env>.projects.<org>/<repo>.destination_uuid), the create-time warning —
placement is repaired by delete + recreate, never by a later apply — and Coolify's own
words kept verbatim, so the next person's search still works.

And the assumption behind an undeclared destination is now on screen at the moment it
is made: `placement: server's default destination (none declared)`. This reverses a
judgment cast held explicitly ("a line on every diff that says nothing is how a report
stops being read" — the test it replaces). The line does not say nothing; it says which
network the next create lands on. It stays on a clean run that creates nothing, too,
because the trap is set for projects that are already built: the day their server gains
a second destination, every one of them that declared no destination stops being able
to create, and nothing will have warned them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 17:25:29 +00:00
Daniel Marin
19f8d6f14e
Merge pull request #39 from claude-hdb/fix/ensure-environment
fix: apply creates the environment its resources name (#38)
2026-07-14 17:44:29 +01:00
claude-hdb
6b363cb5be fix: apply creates the environment its resources name (#38)
POST /projects hands a new project Coolify's OWN default environment,
`production` — never ours. cast then created every resource with
`environment_name: <our --env>`, so the first apply against a project that
did not exist yet 404'd on its first resource ("Environment not found") and
left the project behind, created and empty.

Two comments in the source already asserted the behaviour as though it were
implemented (cli.ts:716, :808), and the README says it outright — the route
existed in the vendored 4.1.2 spec, cast just never called it. It went unseen
because every environment cast had touched until now was hand-built in a UI
and adopted, so it already existed under whatever name someone typed. The
genuinely-from-nothing apply is the one path nobody had run.

apply now reconciles project + environment once per run, before the first
create. Read-before-write: an environment that already exists is never written
to, so adoption is untouched and this cannot regress an apply that works today.
A 409 is read as "present" (the race between our read and our write).

Coolify's default environment is LEFT ALONE, per apply-never-deletes — deleting
it would be the first delete cast ever performs. An empty `production` beside
the environment everything lives in is reported, the same courtesy an orphan
gets, and removed by hand or not at all.

The regression test drives the real failure, not a call count: the fake Coolify
404s a create whose environment_name it does not carry, exactly as a live box
does — against the old executor it reproduces the reported error verbatim.
2026-07-14 16:43:12 +00:00
Daniel Marin
2b202580ea
Merge pull request #35 from claude-hdb/fix/age-key-stdin
fix: hand the age identity to age on stdin — fd paths resolve only in cast's process
2026-07-13 23:51:04 +01:00
claude-hdb
b9ded195da fix: hand the age identity to age on stdin — fd paths resolve only in cast's process
CAST_AGE_KEY_FILE_PROD=<(pm read …) — the documented way to inject a prod
key that never touches disk — expands to /proc/self/fd/N, a path meaningful
only inside the process holding the fd. cast passed that string to a
freshly-spawned age, which resolved it against its own fd table and failed
with ENOENT, for every password manager, on every shell.

node owns the fd, so cast now reads the identity itself and hands it to age
as `-i -` on stdin. The key still never becomes a file, never appears in
argv, and never enters the environment. Not `-i /dev/stdin`: node closes
the pipe before age re-opens it by path (ENXIO).

The regression test reproduces the shape exactly — a key path that only
this process can resolve — and fails against the old code with the same
age ENOENT hit live during the incubator prod migration.

Fixes #34

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 22:43:00 +00:00
Daniel Marin
7d65524054
Merge pull request #33 from claude-hdb/feat/inventory-emit-draft
inventory --emit-draft: a reviewable blueprint of a live instance (#27)
2026-07-13 21:46:10 +01:00
Daniel Marin
94e0d7e8ea
Merge pull request #32 from claude-hdb/feat/fleet-all
fleet operations: cast diff/apply --all over the project registry (#26)
2026-07-13 21:45:55 +01:00
Daniel Marin
637028dc28
Merge pull request #31 from claude-hdb/fix/smoke-project-scoped
smoke resolves its target inside the project it was declared under (#29)
2026-07-13 21:45:40 +01:00
Daniel Marin
dba865f0f0
Merge pull request #30 from claude-hdb/feat/project-registry
a project registry — the list of what exists (#25)
2026-07-13 21:45:23 +01:00
claude-hdb
e96bab5d79 feat: emit a draft of what a box holds — a proposal, never desired state (#27)
`cast inventory` could already see a whole instance (#22). It can now write
down what it sees, in the shape of cast's own inputs:

    cast inventory --env prod --instance box-b --emit-draft ./draft

    draft/
      environments.yaml                 # bindings as far as they can be read — with the projects: registry (#25)
      incubator/.infra/manifest.yaml    # one per project
      incubator/.infra/env/*.env.template
      la-familia/.infra/manifest.yaml   # …including the client sites nobody ever declared
      secrets/<project>.<env>.env.age   # encrypted to a recipient you name
      UNCAPTURED.md                     # ← the important file

Two uses: bootstrapping a project that has no manifest (the third-party sites
on the box being drained were never declared, and never will be unless
something writes the first draft), and a point-in-time blueprint.

A DRAFT IS A PROPOSAL. It is never desired state, and `apply` never reads it:

    sweep → emit draft → a human reads it → manifest PR → capture → apply

Same shape as `terraform import` → HCL, and the boundary is enforced, not
merely documented. It never emits into a repo that already has a manifest —
for a declared project the manifest IS the truth, and one regenerated from a
live box would let that box's accumulated cruft overwrite a reviewed spec, in
the one direction nobody reviews. Adoption is one-way. So: a non-empty target
refuses, a manifest at the path it would write refuses, and --emit-draft with
a repo positional refuses (that is the reconcile path, and it is exactly the
case where a draft must not be written).

Two things would make a draft actively dangerous, and both are the point:

1. COPIED PROVIDER-GENERATED VALUES. A DATABASE_URL read off the source points
   at the SOURCE box's Postgres; rebuild elsewhere and the new box comes up
   WORKING, reading and writing the old box's database, and you find out the
   day the old box is deleted. So the draft applies capture's discipline: a
   provider-generated name is placeheld with the same GENERATED_PLACEHOLDER
   literal, its live value is written into no artifact, and the emitted
   manifest declares it under generated_secrets: so a later capture placeholds
   it again with no flag to remember. The rule is by NAME — Coolify's SERVICE_*
   magic vars, and any name carrying a datastore word and a connection word —
   and it errs wide, because over-matching a real secret is loud and
   recoverable while under-matching a generated one is silent and is not.
   Every other var becomes a ${REF} with its value in the age store, never a
   literal in a committed file: cast cannot know which of a box's vars are
   secret, and a live key written as a literal is a key in a git repo.

2. SILENT LOSSES. UNCAPTURED.md is a first-class output, emitted on every run:
   per resource, every live setting cast saw and could not express —
   destinations (#21), service hostnames, Basic Auth/Traefik labels, backup
   schedules, database kinds cast does not model, env names a template cannot
   hold — plus what no API in 4.1.2 will tell it, and the table of what a
   blueprint still cannot restore (the GitHub App private key and the S3 keys:
   re-create by hand). A blueprint that omits these without saying so is worse
   than no blueprint, because in a disaster you would trust it and rebuild a
   different box.

Secrets are encrypted to a recipient you NAME (--recipient, or the
environment's age_recipient binding). With neither, cast refuses rather than
quietly emitting a draft that looks complete and holds not one value;
--no-secrets says so deliberately. A project with resources in two populated
environments is a tie cast will not break — it refuses, and --environment says
which, as a tiebreak rather than a filter (filtering by name would drop the
client sites, each alone in Coolify's default `production`, out of a blueprint
that claims to describe the box).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:32:25 +00:00
claude-hdb
8dde6dbc05 feat: --all — every project in an environment, and a report that says so (#26)
Every cast verb was single-project, so "do this to the whole instance" was a
shell loop the operator wrote from memory — and the project they forgot is the
one that drifted. `cast diff --env prod --all` and `cast apply --env prod --all`
iterate the registry (#25) instead.

The bulk of this is a refactor: the apply/diff block in cli.ts was one long
inline body, and it is now `runProject` — checkout → secrets → desired →
bindings → live → diff → optionally apply. Both the single-repo path and the
`--all` loop call it, so there is exactly ONE implementation of what a project
run is. A second, parallel fleet path is how the two would drift, and drift is
the subject of this tool. `openCoolify` and the team assert are hoisted out of
it: one --env means one instance and one team, so asserting once still lands
strictly before the FIRST project's first read — the read is already the lie.

Fails closed on the aggregate. A registered project cast cannot reach is an
ERROR, never a skip: the clone failing, no manifest block for this environment,
an absent or undecryptable store, an absent Coolify project/environment, any
HTTP error. A silently skipped project reads exactly like a clean one — #12/#18/
#22 at fleet scale — so the report leads with COVERAGE (registered / read /
clean / drifted / unreachable), and:

  diff --all   0  every registered project was READ, and every one is clean
               1  every one was read, and at least one has drift
               2  a project could not be read — outranking drift, because an
                  unreadable project is not a diff result but the absence of one
  apply --all  0  every registered project applied; non-zero otherwise

`diff --all` runs every project to completion (stopping hides the drift in the
projects it never reached); `apply --all` STOPS at the first failure and names
what it applied and what it did not touch (continuing to mutate a fleet after an
unexplained failure is not a thing cast gets to do).

Two refusals. An empty or absent registry refuses rather than printing
"0 projects, clean" — an empty fleet reading as a clean fleet is the whole
failure this is against; the message distinguishes an unmigrated state file from
a registry pointed elsewhere and prints the YAML to write. And `--all` is
mutually exclusive with the repo positional and with every single-project
coordinate (--path, --project, --environment, --resource, --hostname-overlay):
each names ONE project's checkout, ONE project's Coolify name, ONE box's
resource names, and `--project X` across a fleet would point every project at
the same Coolify project — a false report on diff, and on apply every manifest
in the fleet written into one project.

Also: `projectsIn`'s doc-comment guessed that `[]` made a fleet verb over an
unmigrated state file "a clean no-op rather than a crash". It is precisely
backwards, and now says so. And the --path/--env-prod refusal is hoisted to the
CLI's up-front flag validation (one rule, one string, two call sites in
resolve.ts) — it used to be caught only by accident of resolveCheckout running
before the bindings load.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:23:23 +00:00
claude-hdb
b07563a815 fix: smoke resolves its target inside the project it was declared under (#29)
`smoke` found the application it WRITES to by name against GET /applications —
every app the token can see, across every project and every environment on the
instance — and took the first name match. So `smoke_target: core` did not name
an application; it named whichever `core` Coolify happened to list first. One
instance carrying prod and staging is enough for `cast smoke --env staging` to
POST its canary vars onto prod's `core`, and on the failure path leave them
there.

It now resolves the target through fetchLive(project, environment) — the same
lookup every read-side verb makes — and takes the coordinates that lookup needs:
--project and --environment, with diff/capture/inventory's semantics and
defaults. An application that is not in that project + environment is not an
empty result, it is the absence of anything to write to, so smoke refuses:
naming what it looked for, where the name came from, and what is actually there
(including when the name belongs to a service or a database, which would 404 on
the /envs endpoint smoke writes to).

The <org>/<repo> positional is now REQUIRED, and the deprecated state-file-scoped
`smoke_target` is dropped: it named an app from a key with no project to scope
to, so it could not be fixed, only carried. It is still declared in the schema —
refused with a migration message rather than a strict-mode "unrecognized key",
because loadBindings runs for every verb and an unmigrated state file must not
take `diff` and `apply` down with it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:08:35 +00:00
claude-hdb
18660041f9 feat: a project registry — the list of what exists (#25)
environments.yaml could say where things deploy to, and how a project you
have already named is placed once it is there. It could not say which
projects exist. "Every project" was a thing the operator remembered — so
fleet operations (#26) had nothing to iterate, and rebuild-from-state (#27)
was an assumption, since you cannot restore what you cannot enumerate.

A new optional top-level block, keyed by the full <org>/<repo> slug:

  projects:
    heavy-duty/incubator:
      environments: [prod, staging]

The key IS the repo — no `repo:` field, because a second place to write the
same string is a second place for it to be wrong. No bare-<repo> fallback,
unlike github_apps and environments.<env>.projects: those carry one because
state files in the wild are keyed that way, and this block has none to
support. A bare <repo> is unique only within an org, which is why it is not
a key (#12, twice learned).

Validated in loadBindings, so every verb refuses a registry that lies:

- an environment no `environments:` block defines is an error — the project
  would be registered into an environment no command can visit
- every environments.<env>.projects.<slug> binding must be registered for
  that env, or the two blocks describe two different fleets: a destination
  or smoke_target real enough for a direct apply, invisible to every fleet
  run. Only enforced when `projects:` is present, so pre-registry state
  files keep loading unchanged.

Both defend one failure: a silently skipped project reads exactly like a
clean one. Errors render multi-line now — zod's own .message is the issue
array as JSON, which flattened the refusals into a line of \n escapes.

projectsIn(bindings, env) gives an environment's slugs, sorted; [] with no
registry. The --all flag that consumes it is #26's, not here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 20:03:35 +00:00
Daniel Marin
ee53494806
Merge pull request #28 from claude-hdb/feat/destination-placement
place a resource on a destination — and a state file that can say which (#21)
2026-07-13 20:45:20 +01:00
Daniel Marin
97188b128b
Merge pull request #24 from claude-hdb/feat/inventory-sweep
inventory sweeps the instance — and an empty environment shouts (#22)
2026-07-13 20:45:06 +01:00
claude-hdb
68bf3f1876 docs: point smoke's instance-wide lookup at #29
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 19:42:23 +00:00
claude-hdb
8deaeac07b feat: place a resource on a destination — and a state file that can say which (#21)
A destination is the Docker network a resource is created on. cast never sent
one, so everything landed on the server's default — invisible and harmless while
each server hosts one project, and neither the moment a server hosts two.

The state file had nowhere to say otherwise, either. A destination is scoped
project × environment, and `environments.<env>` is scoped by environment alone:
a `destination:` key there would mean "one network shared by every project in
this environment", which is the isolation it is meant to provide, inverted.

So:

- `environments.<env>.projects.<repo>` — per-project state, keyed by repo, full
  `<org>/<repo>` slug first with a bare-`<repo>` fallback, exactly like
  `github_apps`. It carries `destination_uuid` and `smoke_target`.
- `smoke_target` moves there. It was state-file-scoped: it named ONE project's
  app (`core`) from a key that could not tell two projects apart — or even prod's
  app from staging's. The old key is still read (with a warning), so an unmigrated
  state file keeps smoking, and `cast smoke` now takes an optional `<org>/<repo>`.
- `apply` sends `destination_uuid` on create, for applications, databases and
  services alike — Coolify runs identical destination logic in all three.

The API turns out to be worse than the issue assumed, in a way that changes what
"diff should compare the destination" can honestly mean. Verified against
coollabsio/coolify v4.1.2 (routes/api.php + the three Api controllers), and
written up in reference/README.md:

- There is NO destinations API. Zero routes. A destination cannot be listed, read
  or resolved by name — only a raw UUID from the UI identifies one, exactly as
  with `s3_destination`. Hence `destination_uuid:` and not `destination:`.
- The field is WRITE-ONLY. Coolify takes `destination_uuid` on write and returns
  `destination_id` (an integer PK) on read, with nothing mapping between them.
- On a server with >1 destination, a create that OMITS it is a hard 400. So cast
  could not deploy onto a shared box at all — it did not silently misplace there,
  it simply failed. On a single-destination server the uuid is ignored entirely
  and never validated, so a wrong one is invisible until a second one exists.

A declared UUID therefore cannot be verified against the resource it was sent
for — by cast or by anything else. Diffing it as a field would compare a UUID to
an int and report drift that never clears, so it is reported rather than compared,
and the limit is stated out loud: every diff that declares a destination says it
did not verify it. Silence would make an unverified setting read as a verified
one, which is the failure shape #12/#14/#17/#18 are all about.

What IS comparable is the live side to itself. `diff` groups live resources by the
`destination_id` Coolify does report, and a project whose resources do not all
share one network is drift — non-clean, both sides named, and never repaired
(apply moves nothing between networks). That catches the thing actually worth
catching, including on a box whose destinations were made by hand.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 19:31:02 +00:00
claude-hdb
619b94b7df fix(inventory): biome lint — a bare template literal and a string concat
CI caught what my local gate did not: `npm run check 2>&1 | tail -1 && commit`
takes the exit status of `tail`, not of biome, so the `&&` gated nothing. The
check had been failing locally too; the pipe was swallowing it.

- noUnusedTemplateLiteral: a backtick string with no interpolation
- useTemplate: string concatenation where a template literal belongs

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 19:22:07 +00:00
claude-hdb
10a161ee07 feat: inventory sweeps the instance — a discovery verb that needed you to have discovered
`inventory` (#19/#20) reconciled a manifest against one project and one
environment THAT YOU NAME. But the premise of the verb is that you are looking
at a box you did not build — so you do not know those coordinates yet. It was
a discovery tool that required you to have already discovered, and the operator
went straight back to hand-curling /projects to find out where anything lived.

  cast inventory --env prod --instance box-b     # no repo → sweep

Every project, every environment, every resource the token can see. No manifest,
no store, no age key, no recipient. With a repo it reconciles exactly as before.

Worse than the missing sweep was how the targeted path FAILED. Pointed at a
project's `production` environment — auto-created by Coolify, and empty — it
reported:

    on the box, NOT in the manifest
        (none)
    5 difference(s) between the manifest and this box.

Every word true; the overall impression ("the box has nothing, the manifest has
five things") exactly the D-237 lie cast refuses everywhere else. The resources
were alive and serving production the whole time, in an environment named
`staging` that nobody had ever swapped. An environment with ZERO resources is
far more often the wrong coordinate than an empty one, so it now says so, and
names the sweep.

The sweep asserts the team first, and that matters more here than anywhere:
Coolify scopes what a token can see to its team, so a wrong-team token would
sweep an instance and truthfully report that it is empty.

Environment enumeration takes two roads — GET /projects/{uuid}/environments,
falling back to the relation on GET /projects/{uuid}. The vendored OpenAPI has
been wrong before, and this is the one path where failing to enumerate is worse
than being slow.

The stub in the new suite is shaped like the box this came from: three projects
(two of them unrelated third-party client sites nobody knew were there), an
empty auto-created `production`, and the real system in `staging`.

npm run check + build clean; 173 tests passing (was 169).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 19:07:17 +00:00
Daniel Marin
f7670938ee
Merge pull request #23 from claude-hdb/feat/resource-aliases
--resource: the third name a hand-built box does not share with you
2026-07-13 20:02:27 +01:00
claude-hdb
40322bbf75 feat: --resource, the third name a hand-built box does not share with you
#20 shipped the refusal without shipping the resolution: capture correctly
refuses when a manifest resource does not exist on the source, and then
there was no way to say "it's over there, under another name."

Found immediately, on the box that motivated it. The manifest says `core`,
`landing`, `postgres`, `redis`, `umami`. The box says `Incubator Stack v2`,
`Incubator Landing`, `Incubator Database v2`, `Incubator Redis v2`,
`Incubator Umami`. Neither is wrong — one names things for a human reading a
UI, the other for a machine reading a diff — and neither gets to overwrite the
other.

  --resource core="Incubator Stack v2"     (repeatable)

Applied at the boundary: live resources are renamed to the manifest's
vocabulary once, immediately after the lookup, so computeDiff / classify /
reconcile all match by name exactly as before and none of them needs to know a
hand-built box was involved.

Read-side only, and `apply` refuses it up front — before a clone, a decrypt or
a single call. apply CREATES under the manifest's names, so an alias there
could only mean "adopt the existing one instead": a different operation nobody
has asked for, whose silent failure mode is a duplicate resource created beside
the one you were pointing at.

diff needed this as much as capture did. Without it, a --full diff against a
box whose resources are named differently reports every manifest resource as
"to create" and never mentions the live ones — the D-237 lie by another route,
a confident full-create plan against a box that has all of it under other names.
That diff is the staleness gate of a live migration.

Two smaller things, both about not laundering a naming gap into a pass:

- inventory, when NOTHING matched and yet the box has resources, now says so
  and prints the --resource lines to paste. "The box is empty" is exactly the
  wrong conclusion, and it was the easy one to draw.
- the absent-resource refusal prints the same, per absent resource.

An alias whose left side names no manifest resource is an error, not a no-op:
a typo would otherwise map nothing, leave the real resource looked-up under its
own name, and refuse with no hint that the flag had missed.

inventory keeps the box's own name beside ours in the report (`core ←
"Incubator Stack v2" on the box`) — a document that renamed the box's resources
to our vocabulary and never mentioned theirs would be unusable against the UI
it describes.

npm run check + build clean; 169 tests passing (was 164).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 18:50:06 +00:00
Daniel Marin
bc726db4c1
Merge pull request #20 from claude-hdb/feat/inventory-and-read-side-coordinates
read-side coordinates (#17, #18) + cast inventory (#19)
2026-07-13 19:22:30 +01:00