Files
parking_solution/wiki/decisions/fleet-deployment-komodo.md
T

15 KiB

type, tags, sources, updated, status
type tags sources updated status
decision
parking
deployment
fleet
komodo
netbird
offline-first
threat-model
2026-09-03 settled

Fleet deployment — Komodo Periphery over a NetBird mesh

How the parking appliance is deployed and managed at fleet scale, superseding the single-box, SSH-and-booth.sh model. The image build/tag/registry pipeline (container-deployment) is unchanged — this decides only the control plane that drives those same compose files onto many booths. Settled 2026-06-27.

The problem booth.sh couldn't solve

[[container-deployment|scripts/booth.sh]] is a thin wrapper over docker compose -f base -f prod --env-file .env. It works for one appliance you can get a shell on, but as the fleet grows (the stated direction is many/growing sites) it gives us none of:

  • Remote, no-SSH operation — an update means someone gets a root shell on the booth.
  • A fleet view — which booth runs which dev-<sha>, which is healthy/offline.
  • A deploy audit trail — who deployed what, when.
  • One-click rollback to a previous immutable dev-<sha>.

These are exactly the gaps a deployment controller fills. We already run every prerequisite (a Komodo Core, a NetBird zero-trust mesh, the Gitea registry), so the marginal cost is low.

Decision

Adopt Komodo Periphery on each appliance, driven by the existing Komodo Core over the NetBird mesh. Keep the compose files and the container-deployment verbatim — Komodo consumes them as a Stack; it does not replace them. booth.sh is demoted to a break-glass local fallback for when the mesh/Core is unreachable.

Gitea push ─▶ build-images.yml ─▶ registry (parking-server:dev-<sha>, parking-vision:dev-<sha>)
                                          │
Komodo Core (off-site) ──── NetBird mesh ─┼─▶ Periphery @ booth-A  ─▶ docker compose up (pinned sha)
   • fleet table / history / rollback     ├─▶ Periphery @ booth-B
   • per-booth secret injection           └─▶ Periphery @ booth-C …
   • NO deploy webhook (manual + pinned)

The three load-bearing choices (settled with the user 2026-06-27)

  1. Fleet size: many/growing. Komodo is treated as load-bearing infrastructure, not a convenience. This is what tips the decision away from "SSH-over-NetBird + a playbook".
  2. Deploy trigger: always manual + pinned. No deploy webhook on a booth Stack. A human deploys a specific immutable TAG=<branch>-<sha> from Core. This preserves the determinism we chose when pinning the booth tag (a moving tag auto-redeploying a booth is the surprise we explicitly rejected). This holds for every booth — including the staging booth: it tracks the stage branch but still runs a pinned stage-<sha> (an earlier sketch floated a moving :dev + webhook for staging; rejected in favour of pinned-everywhere). See "Promotion tiers".
  3. Secrets: Komodo-managed (per-booth, unique). Core's secret store injects JWT_SECRET, EVENT_SIGNING_KEY and BACKUP_KEY into the Stack at deploy. This scales (no SSH-to-N-booths to rotate a key) — but see the threat-model tension below; the keys MUST be distinct per booth.

Promotion tiers — dev → stage → main (settled 2026-06-29)

Three git branches map to three image tags and three booth roles:

Branch Image tag Role Deploy
dev :dev / :dev-<sha> working branch no booth runs it
stage :stage / :stage-<sha> staging booth (park-buzi) — real-world test manual + pinned stage-<sha>
main :main / :main-<sha> production booths manual + pinned main-<sha>
  • Promotion is a merge, not a build trigger. When dev is confident-ready, merge dev → stage; CI (build-images.yml, which now triggers on dev/stage/main) builds :stage + :stage-<sha>; then bump TAG=stage-<sha> in komodo/resources.toml and deploy from Core. main only ever receives what survived staging.
  • stage was branched from dev (2026-06-29) so the first real-world test carries the full current app, not a main that predates this work. park-buzi's Stack branch + TAG both point at stage.
  • DB migrations run at boot (migrate-runtime.mjs), so a promotion auto-migrates the staging booth's SQLite ledger — staging is exactly where a bad migration is caught before production.
  • Per-booth secrets must pre-exist in Core for park-buzi: [[park_buzi_jwt_secret]], [[park_buzi_event_signing_key]], [[park_buzi_backup_key]] — distinct, never shared. Once the booth signs real entries under its EVENT_SIGNING_KEY, that key is load-bearing for its ledger forever (escrow it; see backup-recovery).

Why this is safe (against the project's two forces)

Offline-first (offline-first) — Core is orchestration, never a runtime dependency

The booth must run fully when the mesh is down. Komodo's agent model satisfies this: Periphery + the local containers keep operating if Core is unreachable; we lose remote management until the mesh returns, not operation. There must be no runtime path from booth operation to Core — Core only deploys. (Periphery's own liveness is irrelevant to entry/ exit; the Fastify server and SQLite ledger run independently of it.)

Threat model — the adversary is the booth operator (threat-model)

This is the sharp edge, and the reason this page is explicit rather than a footnote.

  • Periphery is a root-capable remote-exec agent on the appliance. If the operator compromises the box, the agent is a lever. Mitigations: bind Periphery only to the NetBird interface (never 0.0.0.0), enforce its passkey + TLS, and fold the agent into the disk-os-hardening surface. It is part of the trusted computing base now.
  • EVENT_SIGNING_KEY is the anti-fraud root. It signs the [[append-only-event-chain| append-only ledger]] — the control between us and a booth operator forging entry/exit events. Holding it in Core means a Core compromise can forge any booth's ledger that shares a key. Two mitigations make central management acceptable:
    • Per-booth, unique keys. Never reuse a signing key across sites, so a single leak taints one booth, not the fleet.
    • The atecc608 is the real long-term signer. The EVENT_SIGNING_KEY HMAC is the interim mechanism; once the secure element signs the chain, the key in Core stops being the fraud root. Tracked in open-questions.
  • Core becomes a Tier-0 asset. It now holds login + ledger keys for the whole fleet, so it must be hardened to the booths' bar: Komodo API bound to the NetBird mesh only, never a public interface; access-controlled; backed up.

Licensing — Komodo is GPL-3.0, and that's fine here

The hard MIT/Apache/BSD constraint (technology-stack) is about shipped app dependencies (code we distribute/link). Komodo is external ops tooling we self-host and don't distribute, so its GPL-3.0 does not taint the product — exactly like the [[vision-service|AGPL ANPR exception]] reasoning (a separate process / external boundary, not a linked dependency). Noted here so it isn't re-litigated.

What lives where

Concern Where Notes
Image build + tags Gitea CI (container-deployment) unchanged: :dev moving + :dev-<sha> immutable
Compose files the repo + on the booth unchanged base + docker-compose.prod.yml
Stack / deploy definition Komodo Core git-synced from komodo/ (infra-as-code)
Which sha is deployed Komodo Core, manual TAG=dev-<sha>, pinned, no webhook
JWT_SECRET, EVENT_SIGNING_KEY Komodo Core secret store per-booth, unique
COOKIE_SECURE=0, TAG, REGISTRY Komodo Stack env per-environment
Registry pull creds Komodo Core so Periphery can pull from Gitea
Local break-glass booth.sh + a local .env mesh-down fallback only

Setup outline

On each appliance (after appliance-provisioning):

  1. Install Komodo Periphery (binary or container), bound only to the NetBird interface; set its passkey/TLS.
  2. Point its compose/stack dir at /opt/parking_systems/ (the existing files).
  3. Keep booth.sh + a minimal local .env (no real secrets) as break-glass.

In Komodo Core:

  1. Add the booth as a Server, address = its NetBird IP (mesh, not LAN/WAN).
  2. Define the Stack = base + docker-compose.prod.yml, env from Core's secret store, secrets per booth.
  3. No deploy webhook on the booth Stack — deploys are manual; set TAG=dev-<sha> explicitly.
  4. Add Gitea registry creds so Periphery can pull.
  5. Sync the Stack/Server definitions from the repo's komodo/ directory (infra-as-code: komodo/resources.toml + README) so the control plane is itself reviewable + version-controlled.

A non-booth stack (2026-09-06)

The Stack model turned out to fit a service that is not a booth: the Car Wash review collector (vision-review-outbox) runs on the reviewer's GPU host as its own [[stack]] (wash-collector, server = "art-docker-station", file_paths = ["docker-compose.collector.yml"]). Same repo, branch and pinned TAG promotion, its own secret references, and — because a stack names its compose files — nothing booth-side lands on that host and nothing of it on a booth. The same stack carries the phase-B trainer as a compose profile (train, bodytype-classifier-training): a deploy never starts it; the owner runs it by hand on the host with docker compose … --profile train run --rm trainer …. So after a deploy of that stack docker ps shows one container — expected; and a deploy pulls nothing for the profile, the first run does (host Docker must be logged in to the registry). Its two env lines (TRAINER_OUT, the TRAINER_PUBLISH_TOKEN secret reference) stay commented in resources.toml until the first publish.

Open / not yet done

  • Per-booth secret generation + rotation flow — how a new site's unique EVENT_SIGNING_KEY is generated and registered in Core (vs. on-site openssl rand). Tie-in: open-questions JWT-key item.
  • ATECC608 as the signer supersedes EVENT_SIGNING_KEY-in-Core as the fraud root — until then central secrets carry the blast-radius noted above.
  • Periphery hardening checklist folded into disk-os-hardening (interface binding, passkey, TLS, agent as TCB).
  • Staging vs production booth split — ✅ MODELLED 2026-06-29 (see "Promotion tiers" below). park-buzi is the first staging booth, tracking the stage branch / :stage image, deployed manual + pinned (TAG=stage-<sha>, no webhook — we hold the no-moving-tag-on-a-booth line even on staging, not the webhook-on-staging option the earlier sketch floated). komodo/resources.toml
    • komodo/README.md updated.
  • Core backup / DR — Core is now Tier-0; its loss = no fleet management (operation unaffected, per offline-first). Backup story TBD.

Supersedes / relates

  • Supersedes the "SSH + booth.sh is the deploy mechanism" assumption in container-deployment (that page's build/tag/registry content stands; its booth.sh-as- primary-deploy framing is now the fallback). Cross-linked there.
  • Companion: the komodo/ infra-as-code sketch (in the repo, not the wiki), appliance-provisioning (what runs before Periphery), disk-os-hardening (the appliance's hardening surface).

park-lab — the lab bench joins the fleet (2026-07-07)

Second [[stack]] in komodo/resources.toml: park-lab (server = the lab box's Periphery connect_as), the first non-booth member and the proof of the tier model in practice:

Stack compose branch image tag secrets
park-lab dev moving dev (a lab may float) park_lab_*
park-buzi stage pinned stage-<sha> park_buzi_*

The three knobs are independent per stack — the ResourceSync's own branch only governs where the FILE is read from, each stack's branch picks its compose files, TAG picks the image. Per-box secrets even in the lab (blast radius). The lab box earned its keep immediately: it caught the USB close-cancel truncation, the printer/controller wizard gate, and the Periphery v2.2.0 root_directory default before any of them reached a real booth (appliance-provisioning §7a).

ResourceSync branch drift — the exact gotcha this page already warned about (2026-09-03)

This page's own §"park-lab" note (2026-07-07) already spelled it out: "the ResourceSync's own branch only governs where the FILE is read from" — independent of any [[stack]]'s own branch field. It bit anyway. resource-sync-park-systems in Komodo Core was pointed at dev, while park-buzi and park-2 are stage-tier Stacks (branch = "stage", pinned TAG=stage-<sha>, per the promotion-tiers model above). resources.toml had been byte-identical on dev and stage since park-buzi's Stack was first written, so this had zero observable effect for months — until a desktop-app debugging session (see desktop-shell-tauri) landed 9 real commits on dev (including a WS_ALLOWED_ORIGINS fix) that were never merged to stage, creating the first genuine divergence between the two branches.

Symptom: merged dev → stage, pushed, bumped TAG in resources.toml on stage, committed, pushed — then destroyed + recreated the park-2 Stack in Komodo Core and it STILL came back running the old image. Every sync was silently re-reading resources.toml from dev (which still had the stale TAG), overwriting the correct value just committed on stage. No error, no warning — the sync just quietly did what it was configured to do, from the wrong branch.

Fix: pointed resource-sync-park-systems at stage in Komodo Core's UI (Sync config → branch field), then re-synced + redeployed park-2 — confirmed via /api/version (previously 404, proving a stale image; correctly 401-auth-gated after the fix, proving the new image + route exist).

Standing lesson, now written twice: a [[stack]]'s promotion tier (which branch its own branch/TAG fields track) and the ResourceSync resource's own git branch are two independently configured settings in Komodo Core — nothing enforces they agree, and a mismatch is invisible until the two branches' resources.toml actually diverge. Check this FIRST whenever a redeploy doesn't pick up an expected resources.toml change, before assuming the change itself, the CI build, or the deploy step is broken.