feat(deploy): Komodo fleet deployment — resources.toml + decision
CI / check (push) Successful in 36s

Adopt Komodo Periphery (over the NetBird mesh) as the booth fleet control plane,
superseding SSH-and-booth.sh. The booth runs the SAME compose files; Komodo Core
drives them remotely. booth.sh is demoted to a break-glass local fallback.

- komodo/resources.toml mirrors the working park-buzi Stack (built by hand in the
  Core UI, then exported to TOML — field names match the running v2.2). Stack-only:
  servers are created by the agent onboarding OUTBOUND (one-time onboarding key →
  Periphery self-registers, auto-rotating keys, booth opens no inbound port), so
  there is no [[server]] block. Per-booth secrets via [[...]] refs to Core's store.
- komodo/README.md + .env.komodo.example document the flow and the hard rules
  (no webhook; onboarding/outbound/mesh-only; per-booth unique secrets; never
  down -v the ledger volume).
- wiki/decisions/fleet-deployment-komodo.md records the decision + threat-model
  analysis (Periphery is a root agent → mesh-bound; EVENT_SIGNING_KEY-in-Core is a
  fraud-root blast radius until ATECC608 signs; Core is now Tier-0; GPL-3.0 is fine
  as external ops tooling). container-deployment reframed (booth.sh = fallback);
  index + log updated.

Verified end-to-end against a real booth (park-buzi): onboarded OK, Stack deployed,
all containers green, admin seeded.

Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
This commit is contained in:
2026-06-27 12:14:04 +02:00
parent 83298bc0c5
commit 9918f278b2
7 changed files with 321 additions and 0 deletions
+6
View File
@@ -12,6 +12,12 @@ How the parking system's runtime apps are packaged as containers, tagged, and pu
Settled 2026-06-22. Companion to [[vision-service-packaging]] (which scopes the vision service
into the monorepo) and the desktop [[desktop-shell-tauri]] (a separate, tag-only bundle).
> **The build/tag/registry pipeline below is current.** What changed (2026-06-27): the
> *deploy mechanism* is no longer "SSH in and run `booth.sh`". At fleet scale that's superseded
> by [[fleet-deployment-komodo]] (Komodo Periphery over a NetBird mesh, driving these same
> compose files). `scripts/booth.sh` is now a **break-glass local fallback**, not the primary
> deploy path.
## Two images (the desktop app is NOT containerized)
- **`parking-server`** — the Fastify API **plus the built React SPA**. One container serves both:
+151
View File
@@ -0,0 +1,151 @@
---
type: decision
tags: [parking, deployment, fleet, komodo, netbird, offline-first, threat-model]
sources: []
updated: 2026-06-27
status: settled
---
# Fleet deployment — Komodo Periphery over a NetBird mesh
How the parking appliance is deployed and managed **at fleet scale**, superseding the
single-box, SSH-and-`booth.sh` model. The image build/tag/registry pipeline
([[container-deployment]]) is unchanged — this decides only the *control plane* that drives
those same compose files onto many booths. Settled 2026-06-27.
## The problem `booth.sh` couldn't solve
[[container-deployment|`scripts/booth.sh`]] is a thin wrapper over `docker compose -f base -f
prod --env-file .env`. It works for **one** appliance you can get a shell on, but as the fleet
grows (the stated direction is **many/growing** sites) it gives us none of:
- **Remote, no-SSH operation** — an update means someone gets a root shell on the booth.
- **A fleet view** — which booth runs which `dev-<sha>`, which is healthy/offline.
- **A deploy audit trail** — who deployed what, when.
- **One-click rollback** to a previous immutable `dev-<sha>`.
These are exactly the gaps a deployment controller fills. We already run every prerequisite
(a **Komodo Core**, a **NetBird** zero-trust mesh, the **Gitea registry**), so the marginal
cost is low.
## Decision
Adopt **Komodo Periphery** on each appliance, driven by the existing **Komodo Core** over the
**NetBird** mesh. Keep the compose files and the [[container-deployment|image pipeline]]
verbatim — Komodo consumes them as a *Stack*; it does not replace them. `booth.sh` is demoted
to a **break-glass local fallback** for when the mesh/Core is unreachable.
```
Gitea push ─▶ build-images.yml ─▶ registry (parking-server:dev-<sha>, parking-vision:dev-<sha>)
│
Komodo Core (off-site) ──── NetBird mesh ─┼─▶ Periphery @ booth-A ─▶ docker compose up (pinned sha)
• fleet table / history / rollback ├─▶ Periphery @ booth-B
• per-booth secret injection └─▶ Periphery @ booth-C …
• NO deploy webhook (manual + pinned)
```
### The three load-bearing choices (settled with the user 2026-06-27)
1. **Fleet size: many/growing.** Komodo is treated as load-bearing infrastructure, not a
convenience. This is what tips the decision away from "SSH-over-NetBird + a playbook".
2. **Deploy trigger: always manual + pinned.** **No deploy webhook on a booth Stack.** A human
deploys a specific immutable `TAG=dev-<sha>` from Core. This preserves the determinism we
chose when pinning the booth tag (a moving `:dev` auto-redeploying a production booth is the
surprise we explicitly rejected). A *staging* booth MAY track `:dev`; a production booth
never does.
3. **Secrets: Komodo-managed (per-booth, unique).** Core's secret store injects `JWT_SECRET`
and `EVENT_SIGNING_KEY` into the Stack at deploy. This scales (no SSH-to-N-booths to rotate
a key) — but see the threat-model tension below; the keys MUST be **distinct per booth**.
## Why this is safe (against the project's two forces)
### Offline-first ([[offline-first]]) — Core is orchestration, never a runtime dependency
The booth must run **fully when the mesh is down**. Komodo's agent model satisfies this:
Periphery + the local containers keep operating if Core is unreachable; we lose *remote
management* until the mesh returns, **not operation**. There must be **no runtime path** from
booth operation to Core — Core only deploys. (Periphery's own liveness is irrelevant to entry/
exit; the Fastify server and SQLite ledger run independently of it.)
### Threat model — the adversary is the booth operator ([[threat-model]])
This is the sharp edge, and the reason this page is explicit rather than a footnote.
- **Periphery is a root-capable remote-exec agent on the appliance.** If the operator
compromises the box, the agent is a lever. Mitigations: bind Periphery **only to the NetBird
interface** (never `0.0.0.0`), enforce its **passkey + TLS**, and fold the agent into the
[[disk-os-hardening]] surface. It is part of the trusted computing base now.
- **`EVENT_SIGNING_KEY` is the anti-fraud root.** It signs the [[append-only-event-chain|
append-only ledger]] — the control between us and a booth operator forging entry/exit events.
Holding it in Core means **a Core compromise can forge any booth's ledger that shares a key**.
Two mitigations make central management acceptable:
- **Per-booth, unique keys.** Never reuse a signing key across sites, so a single leak taints
one booth, not the fleet.
- **The [[atecc608|ATECC608]] is the real long-term signer.** The
`EVENT_SIGNING_KEY` HMAC is the *interim* mechanism; once the secure element signs the
chain, the key in Core stops being the fraud root. Tracked in [[open-questions]].
- **Core becomes a Tier-0 asset.** It now holds login + ledger keys for the whole fleet, so it
must be hardened to the booths' bar: Komodo API bound to the NetBird mesh only, never a public
interface; access-controlled; backed up.
### Licensing — Komodo is GPL-3.0, and that's fine here
The hard MIT/Apache/BSD constraint ([[technology-stack]]) is about **shipped app dependencies**
(code we distribute/link). Komodo is **external ops tooling we self-host and don't distribute**,
so its GPL-3.0 does not taint the product — exactly like the [[vision-service|AGPL ANPR
exception]] reasoning (a separate process / external boundary, not a linked dependency). Noted
here so it isn't re-litigated.
## What lives where
| Concern | Where | Notes |
| --- | --- | --- |
| Image build + tags | Gitea CI ([[container-deployment]]) | unchanged: `:dev` moving + `:dev-<sha>` immutable |
| Compose files | the repo + on the booth | unchanged base + `docker-compose.prod.yml` |
| Stack / deploy definition | **Komodo Core** | git-synced from `komodo/` (infra-as-code) |
| Which sha is deployed | **Komodo Core**, manual | `TAG=dev-<sha>`, pinned, no webhook |
| `JWT_SECRET`, `EVENT_SIGNING_KEY` | **Komodo Core** secret store | **per-booth, unique** |
| `COOKIE_SECURE=0`, `TAG`, `REGISTRY` | Komodo Stack env | per-environment |
| Registry pull creds | **Komodo Core** | so Periphery can pull from Gitea |
| Local break-glass | `booth.sh` + a local `.env` | mesh-down fallback only |
## Setup outline
**On each appliance** (after [[appliance-provisioning]]):
1. Install **Komodo Periphery** (binary or container), bound **only** to the NetBird interface;
set its passkey/TLS.
2. Point its compose/stack dir at `/opt/parking_systems/` (the existing files).
3. Keep `booth.sh` + a minimal local `.env` (no real secrets) as break-glass.
**In Komodo Core:**
1. Add the booth as a **Server**, address = its **NetBird IP** (mesh, not LAN/WAN).
2. Define the **Stack** = base + `docker-compose.prod.yml`, env from Core's secret store, secrets
**per booth**.
3. **No deploy webhook** on the booth Stack — deploys are manual; set `TAG=dev-<sha>` explicitly.
4. Add Gitea registry creds so Periphery can pull.
5. Sync the Stack/Server definitions from the repo's `komodo/` directory (infra-as-code:
`komodo/resources.toml` + README) so the control plane is itself reviewable +
version-controlled.
## Open / not yet done
- **Per-booth secret generation + rotation flow** — how a new site's unique `EVENT_SIGNING_KEY`
is generated and registered in Core (vs. on-site `openssl rand`). Tie-in: [[open-questions]]
JWT-key item.
- **ATECC608 as the signer** supersedes `EVENT_SIGNING_KEY`-in-Core as the fraud root — until
then central secrets carry the blast-radius noted above.
- **Periphery hardening checklist** folded into [[disk-os-hardening]] (interface binding, passkey,
TLS, agent as TCB).
- **Staging vs production booth split** (a staging booth on `:dev` with a webhook; production
manual+pinned) — not yet modelled in `komodo/`.
- **Core backup / DR** — Core is now Tier-0; its loss = no fleet management (operation
unaffected, per offline-first). Backup story TBD.
## Supersedes / relates
- **Supersedes** the "SSH + `booth.sh` is the deploy mechanism" assumption in
[[container-deployment]] (that page's *build/tag/registry* content stands; its `booth.sh`-as-
primary-deploy framing is now the fallback). Cross-linked there.
- Companion: the `komodo/` infra-as-code sketch (in the repo, not the wiki),
[[appliance-provisioning]] (what runs *before* Periphery), [[disk-os-hardening]] (the
appliance's hardening surface).