apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash collector's SQLite + crops read-only off its volume; time split (validation = newest slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against the torch model; NO model file below the validation floor (exit 3, report still written); exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package. Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip. apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract: crop margin, input size, RGB 0-255, normalisation inside the graph) and RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained on, min-confidence, `detector_class` on the result; path set but no file = phase B off without an error; a broken file is a health detail. models/bodytype.version (tracked, empty) pins the published version the Dockerfile fetches at build (BuildKit secret; a pin that cannot be fetched fails the build). Verified: a trainer model gives identical probabilities inside the vision service; both images built and smoke-tested. Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI. Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site, secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox, vision-service-packaging, fleet-deployment-komodo, index, log. Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
14 KiB
type, tags, sources, updated, status
| type | tags | sources | updated | status | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| decision |
|
2026-09-03 | settled |
Fleet deployment — Komodo Periphery over a NetBird mesh
How the parking appliance is deployed and managed at fleet scale, superseding the
single-box, SSH-and-booth.sh model. The image build/tag/registry pipeline
(container-deployment) is unchanged — this decides only the control plane that drives
those same compose files onto many booths. Settled 2026-06-27.
The problem booth.sh couldn't solve
[[container-deployment|scripts/booth.sh]] is a thin wrapper over docker compose -f base -f prod --env-file .env. It works for one appliance you can get a shell on, but as the fleet
grows (the stated direction is many/growing sites) it gives us none of:
- Remote, no-SSH operation — an update means someone gets a root shell on the booth.
- A fleet view — which booth runs which
dev-<sha>, which is healthy/offline. - A deploy audit trail — who deployed what, when.
- One-click rollback to a previous immutable
dev-<sha>.
These are exactly the gaps a deployment controller fills. We already run every prerequisite (a Komodo Core, a NetBird zero-trust mesh, the Gitea registry), so the marginal cost is low.
Decision
Adopt Komodo Periphery on each appliance, driven by the existing Komodo Core over the
NetBird mesh. Keep the compose files and the container-deployment
verbatim — Komodo consumes them as a Stack; it does not replace them. booth.sh is demoted
to a break-glass local fallback for when the mesh/Core is unreachable.
Gitea push ─▶ build-images.yml ─▶ registry (parking-server:dev-<sha>, parking-vision:dev-<sha>)
│
Komodo Core (off-site) ──── NetBird mesh ─┼─▶ Periphery @ booth-A ─▶ docker compose up (pinned sha)
• fleet table / history / rollback ├─▶ Periphery @ booth-B
• per-booth secret injection └─▶ Periphery @ booth-C …
• NO deploy webhook (manual + pinned)
The three load-bearing choices (settled with the user 2026-06-27)
- Fleet size: many/growing. Komodo is treated as load-bearing infrastructure, not a convenience. This is what tips the decision away from "SSH-over-NetBird + a playbook".
- Deploy trigger: always manual + pinned. No deploy webhook on a booth Stack. A human
deploys a specific immutable
TAG=<branch>-<sha>from Core. This preserves the determinism we chose when pinning the booth tag (a moving tag auto-redeploying a booth is the surprise we explicitly rejected). This holds for every booth — including the staging booth: it tracks thestagebranch but still runs a pinnedstage-<sha>(an earlier sketch floated a moving:dev+ webhook for staging; rejected in favour of pinned-everywhere). See "Promotion tiers". - Secrets: Komodo-managed (per-booth, unique). Core's secret store injects
JWT_SECRET,EVENT_SIGNING_KEYandBACKUP_KEYinto the Stack at deploy. This scales (no SSH-to-N-booths to rotate a key) — but see the threat-model tension below; the keys MUST be distinct per booth.
Promotion tiers — dev → stage → main (settled 2026-06-29)
Three git branches map to three image tags and three booth roles:
| Branch | Image tag | Role | Deploy |
|---|---|---|---|
dev |
:dev / :dev-<sha> |
working branch | no booth runs it |
stage |
:stage / :stage-<sha> |
staging booth (park-buzi) — real-world test |
manual + pinned stage-<sha> |
main |
:main / :main-<sha> |
production booths | manual + pinned main-<sha> |
- Promotion is a merge, not a build trigger. When
devis confident-ready, mergedev → stage; CI (build-images.yml, which now triggers ondev/stage/main) builds:stage+:stage-<sha>; then bumpTAG=stage-<sha>inkomodo/resources.tomland deploy from Core.mainonly ever receives what survived staging. stagewas branched fromdev(2026-06-29) so the first real-world test carries the full current app, not amainthat predates this work.park-buzi's Stackbranch+TAGboth point atstage.- DB migrations run at boot (
migrate-runtime.mjs), so a promotion auto-migrates the staging booth's SQLite ledger — staging is exactly where a bad migration is caught before production. - Per-booth secrets must pre-exist in Core for
park-buzi:[[park_buzi_jwt_secret]],[[park_buzi_event_signing_key]],[[park_buzi_backup_key]]— distinct, never shared. Once the booth signs real entries under itsEVENT_SIGNING_KEY, that key is load-bearing for its ledger forever (escrow it; see backup-recovery).
Why this is safe (against the project's two forces)
Offline-first (offline-first) — Core is orchestration, never a runtime dependency
The booth must run fully when the mesh is down. Komodo's agent model satisfies this: Periphery + the local containers keep operating if Core is unreachable; we lose remote management until the mesh returns, not operation. There must be no runtime path from booth operation to Core — Core only deploys. (Periphery's own liveness is irrelevant to entry/ exit; the Fastify server and SQLite ledger run independently of it.)
Threat model — the adversary is the booth operator (threat-model)
This is the sharp edge, and the reason this page is explicit rather than a footnote.
- Periphery is a root-capable remote-exec agent on the appliance. If the operator
compromises the box, the agent is a lever. Mitigations: bind Periphery only to the NetBird
interface (never
0.0.0.0), enforce its passkey + TLS, and fold the agent into the disk-os-hardening surface. It is part of the trusted computing base now. EVENT_SIGNING_KEYis the anti-fraud root. It signs the [[append-only-event-chain| append-only ledger]] — the control between us and a booth operator forging entry/exit events. Holding it in Core means a Core compromise can forge any booth's ledger that shares a key. Two mitigations make central management acceptable:- Per-booth, unique keys. Never reuse a signing key across sites, so a single leak taints one booth, not the fleet.
- The atecc608 is the real long-term signer. The
EVENT_SIGNING_KEYHMAC is the interim mechanism; once the secure element signs the chain, the key in Core stops being the fraud root. Tracked in open-questions.
- Core becomes a Tier-0 asset. It now holds login + ledger keys for the whole fleet, so it must be hardened to the booths' bar: Komodo API bound to the NetBird mesh only, never a public interface; access-controlled; backed up.
Licensing — Komodo is GPL-3.0, and that's fine here
The hard MIT/Apache/BSD constraint (technology-stack) is about shipped app dependencies (code we distribute/link). Komodo is external ops tooling we self-host and don't distribute, so its GPL-3.0 does not taint the product — exactly like the [[vision-service|AGPL ANPR exception]] reasoning (a separate process / external boundary, not a linked dependency). Noted here so it isn't re-litigated.
What lives where
| Concern | Where | Notes |
|---|---|---|
| Image build + tags | Gitea CI (container-deployment) | unchanged: :dev moving + :dev-<sha> immutable |
| Compose files | the repo + on the booth | unchanged base + docker-compose.prod.yml |
| Stack / deploy definition | Komodo Core | git-synced from komodo/ (infra-as-code) |
| Which sha is deployed | Komodo Core, manual | TAG=dev-<sha>, pinned, no webhook |
JWT_SECRET, EVENT_SIGNING_KEY |
Komodo Core secret store | per-booth, unique |
COOKIE_SECURE=0, TAG, REGISTRY |
Komodo Stack env | per-environment |
| Registry pull creds | Komodo Core | so Periphery can pull from Gitea |
| Local break-glass | booth.sh + a local .env |
mesh-down fallback only |
Setup outline
On each appliance (after appliance-provisioning):
- Install Komodo Periphery (binary or container), bound only to the NetBird interface; set its passkey/TLS.
- Point its compose/stack dir at
/opt/parking_systems/(the existing files). - Keep
booth.sh+ a minimal local.env(no real secrets) as break-glass.
In Komodo Core:
- Add the booth as a Server, address = its NetBird IP (mesh, not LAN/WAN).
- Define the Stack = base +
docker-compose.prod.yml, env from Core's secret store, secrets per booth. - No deploy webhook on the booth Stack — deploys are manual; set
TAG=dev-<sha>explicitly. - Add Gitea registry creds so Periphery can pull.
- Sync the Stack/Server definitions from the repo's
komodo/directory (infra-as-code:komodo/resources.toml+ README) so the control plane is itself reviewable + version-controlled.
A non-booth stack (2026-09-06)
The Stack model turned out to fit a service that is not a booth: the Car Wash review
collector (vision-review-outbox) runs on the reviewer's GPU host as its own [[stack]]
(wash-collector, server = "art-docker-station", file_paths = ["docker-compose.collector.yml"]).
Same repo, branch and pinned TAG promotion, its own secret references, and — because a stack
names its compose files — nothing booth-side lands on that host and nothing of it on a booth.
The same stack carries the phase-B trainer as a compose profile (train,
bodytype-classifier-training): a deploy never starts it; the owner runs it by hand on the host
with docker compose … --profile train run --rm trainer …. Its two env lines (TRAINER_OUT, the
TRAINER_PUBLISH_TOKEN secret reference) stay commented in resources.toml until the first
publish.
Open / not yet done
- Per-booth secret generation + rotation flow — how a new site's unique
EVENT_SIGNING_KEYis generated and registered in Core (vs. on-siteopenssl rand). Tie-in: open-questions JWT-key item. - ATECC608 as the signer supersedes
EVENT_SIGNING_KEY-in-Core as the fraud root — until then central secrets carry the blast-radius noted above. - Periphery hardening checklist folded into disk-os-hardening (interface binding, passkey, TLS, agent as TCB).
- Staging vs production booth split — ✅ MODELLED 2026-06-29 (see "Promotion tiers" below).
park-buziis the first staging booth, tracking thestagebranch /:stageimage, deployed manual + pinned (TAG=stage-<sha>, no webhook — we hold the no-moving-tag-on-a-booth line even on staging, not the webhook-on-staging option the earlier sketch floated).komodo/resources.tomlkomodo/README.mdupdated.
- Core backup / DR — Core is now Tier-0; its loss = no fleet management (operation unaffected, per offline-first). Backup story TBD.
Supersedes / relates
- Supersedes the "SSH +
booth.shis the deploy mechanism" assumption in container-deployment (that page's build/tag/registry content stands; itsbooth.sh-as- primary-deploy framing is now the fallback). Cross-linked there. - Companion: the
komodo/infra-as-code sketch (in the repo, not the wiki), appliance-provisioning (what runs before Periphery), disk-os-hardening (the appliance's hardening surface).
park-lab — the lab bench joins the fleet (2026-07-07)
Second [[stack]] in komodo/resources.toml: park-lab (server = the lab box's Periphery
connect_as), the first non-booth member and the proof of the tier model in practice:
| Stack | compose branch | image tag | secrets |
|---|---|---|---|
| park-lab | dev |
moving dev (a lab may float) |
park_lab_* |
| park-buzi | stage |
pinned stage-<sha> |
park_buzi_* |
The three knobs are independent per stack — the ResourceSync's own branch only governs where the
FILE is read from, each stack's branch picks its compose files, TAG picks the image. Per-box
secrets even in the lab (blast radius). The lab box earned its keep immediately: it caught the
USB close-cancel truncation, the printer/controller wizard gate, and the Periphery v2.2.0
root_directory default before any of them reached a real booth (appliance-provisioning §7a).
ResourceSync branch drift — the exact gotcha this page already warned about (2026-09-03)
This page's own §"park-lab" note (2026-07-07) already spelled it out: "the ResourceSync's own
branch only governs where the FILE is read from" — independent of any [[stack]]'s own branch
field. It bit anyway. resource-sync-park-systems in Komodo Core was pointed at dev, while
park-buzi and park-2 are stage-tier Stacks (branch = "stage", pinned TAG=stage-<sha>, per
the promotion-tiers model above). resources.toml had been byte-identical on dev and stage
since park-buzi's Stack was first written, so this had zero observable effect for months — until
a desktop-app debugging session (see desktop-shell-tauri) landed 9 real commits on dev
(including a WS_ALLOWED_ORIGINS fix) that were never merged to stage, creating the first genuine
divergence between the two branches.
Symptom: merged dev → stage, pushed, bumped TAG in resources.toml on stage, committed,
pushed — then destroyed + recreated the park-2 Stack in Komodo Core and it STILL came back running
the old image. Every sync was silently re-reading resources.toml from dev (which still had the
stale TAG), overwriting the correct value just committed on stage. No error, no warning — the
sync just quietly did what it was configured to do, from the wrong branch.
Fix: pointed resource-sync-park-systems at stage in Komodo Core's UI (Sync config → branch
field), then re-synced + redeployed park-2 — confirmed via /api/version (previously 404,
proving a stale image; correctly 401-auth-gated after the fix, proving the new image + route exist).
Standing lesson, now written twice: a [[stack]]'s promotion tier (which branch its own
branch/TAG fields track) and the ResourceSync resource's own git branch are two independently
configured settings in Komodo Core — nothing enforces they agree, and a mismatch is invisible
until the two branches' resources.toml actually diverge. Check this FIRST whenever a
redeploy doesn't pick up an expected resources.toml change, before assuming the change itself,
the CI build, or the deploy step is broken.