Multi-tenant SaaS layered on the offline model: link-up monitoring of the signed ledger, device status, financials; one admin → many sites; per-site secret custody; recurring fee. Records the four tensions, the confirmed secrets boundary (sync creds + device-password escrow + app identity, NOT the signing key), and the two in-discussion corrections that stand (NetBird already solves booth isolation; remote barrier-open is pulseOpen-and-signed, driven by the unmanned future). status: open, postponed. Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
13 KiB
type, tags, sources, updated, status
| type | tags | sources | updated | status | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| decision |
|
2026-07-13 | open |
Cloud service — multi-tenant SaaS for fleet monitoring & control
Status: postponed (2026-07-13). Captured as context, not a commitment. This records an early requirements/architecture discussion so it isn't re-derived from scratch later. No app code, no schema. Two of the initial requirements were corrected in-discussion (see "Corrections" below) — read those before treating any first-pass answer as settled.
The idea
The offline backup model we ship today is the right tradeoff for offline sites and stays. On top of it, the user wants an online, multi-tenant SaaS — the "cloud service" — that subscribing park sites connect to for real-time (link-up) monitoring: the signed ledger, device status, financial reports, and whatever else a site reports. One admin owns more than one site (a portfolio). The cloud also custodies per-site secrets. Business model: recurring per-site monthly/yearly fee — a revenue line the offline appliance alone can't produce.
This is the customer-facing evolution of the off-site control plane that fleet-deployment-komodo already stood up (Komodo Core, NetBird mesh, Gitea registry). Much of the transport and Tier-0 reasoning there carries over directly; this page is about turning that internal ops plane into a multi-tenant product.
The four hard tensions (what makes a naïve SaaS wrong here)
The booth's two governing forces (offline-first, threat-model) plus the signed ledger (append-only-event-chain) make the "obvious" SaaS shape wrong. Four tensions dominate:
-
Offline-first vs. real-time monitoring. The cloud must never be in the critical path of entry/exit/payment/barrier (offline-first). It is a read-mostly mirror + control-plane, fed by the booth when the link is up, tolerant of hours/days offline, and unable to block booth operation by being down. "Real-time" = near-real-time when up, gracefully stale when not — and the UI must show staleness honestly (last-seen everywhere), never paint a dark site green.
-
The signed ledger must stay verifiable in the cloud, not merely displayed. If subscribers see "their ledger" in the cloud, the cloud copy must be re-verified server-side — re-check the hash chain + signatures on ingest, flag gaps/breaks/forks loudly. The [[threat-model|operator-as- adversary]] extends upward: an operator may want the cloud not to see certain events, so the sync must be gap-evident (sequence continuity). This is both the anti-tamper mechanism and a headline feature — "we can prove your revenue record wasn't altered, even by your own night shift." See reconciliation (this is reconciliation, productised).
-
Secrets for every site — the scariest requirement. A central secret store for hundreds of sites is a single juicy target. The custody boundary must be deliberate — see "Secrets boundary".
-
Multi-tenancy under operator-as-adversary — now at two levels. One admin, many sites ⇒ a new portfolio-owner role above the existing per-site roles (local-jwt-auth admin/operator/ cashier/readonly). Row-level tenant isolation must be airtight — a bug now leaks another company's revenue, not just an intra-site escalation. Every row carries
tenant_id+site_id, non-optional in the query path (not a filter someone can forget). Cloud identity is separate from booth-local auth — the booth keeps its offline JWT/bcrypt login untouched; a site never authenticates its users against the cloud (that would break offline-first).
Secrets boundary (settled-in-principle 2026-07-13)
User confirmed the cloud custodies three classes — and not the crown jewel:
| Class | Cloud custodies? | Notes |
|---|---|---|
| Sync/connection creds + ledger public (verify) key | ✅ yes | Per-site uplink credential + the public half to verify signatures. Smallest blast radius. |
Device/controller passwords (Dingtian relay_pw, camera creds, push tokens) |
✅ yes, as escrow | Solves the real pain: lost relay_pw after a DB reset (dingtian-relay). See escrow rules below. |
| App/admin identity (portfolio login) | ✅ yes | Cloud-side identity for portfolio admins. Separate from booth-local auth. |
| Ledger signing key / atecc608 private key, LUKS/TPM material | ❌ never | Centralising the signer kills the anti-fraud model (append-only-event-chain, hardware-signer-options). The user did not pick this. |
How device-password escrow must work (so it earns its keep instead of becoming the breach):
- Envelope encryption, per-tenant DEK, DEKs wrapped by a KMS master key; a DB dump is ciphertext, every decrypt is KMS-audited.
- The cloud is an escrow, not an operational credential store. Its job is "give the booth back
its
relay_pwafter a wipe," not "the cloud logs into the Dingtian." Decryption happens at the booth (booth fetches its own wrapped blob, unwraps locally); ideally the cloud never holds plaintext device secrets in memory. This keeps the [[dingtian-http-api-unauthenticated|unauthenticated- CGI]] exposure host-local. - The booth threat model applies upward: writes to escrow are append/version ops the operator can't silently rewrite; reads are logged where the operator can't scrub them.
- Sellable as: "your device credentials survive any wipe, encrypted so even we can't read them in bulk."
Corrections made in-discussion (2026-07-13) — read these
The first pass argued against the user's two boldest choices ("cloud reaches into the booth"; implicitly, "no remote barrier open"). The user corrected both, and the corrections stand.
Correction 1 — NetBird already solves the isolation objection
Initial worry: a cloud tunnel into the booth is a new inbound attack surface on every site. But park-buzi is already monitored remotely over a NetBird private mesh (WireGuard) — the same mesh fleet-deployment-komodo uses. The booth dials out to join the overlay; nothing is exposed on the booth PC. So "cloud reaches booth" is the booth-dialed reverse-channel pattern already in production, not a new hole. The objection is withdrawn. What it shifts rather than removes:
- Trust moves to the overlay's identity/ACL layer: "cloud can reach the booth" now means "any peer the mesh authorizes can reach the booth host." Mesh ACLs must enforce the same tenant isolation as the app layer — site A's admin never gets a route to site B's booth. Multi-tenant isolation in a different hat.
- The access-controller VLAN still holds: the mesh terminates at the host, not the controller segment. A cloud peer talks to the booth API; the booth talks to the Dingtian/UHPPOTE (network-isolation, access-direction-is-per-relay). The cloud never gets an L3 route to the UDP relay.
- NetBird's control plane joins the trust base (self-hosted = another service to harden; their SaaS = a third party who can authorize peers). A conscious call, not an architecture change.
Correction 2 — remote barrier-open is compatible with barrier-not-a-door, and the unmanned future requires it
Initial worry: the cloud must never open a barrier. The user's driver is the [[autonomous-direction| unmanned-site]] future — no operator on-site; if the exit reader or payment dies, someone must open the barrier remotely rather than trap people ("we can't take hostages because a stupid device is not responsive"). This is right, and it does not violate barrier-not-a-door:
- That rule was never "no remote open." It forbids driving the barrier as a timed auto-close
("open for N ms"); physical safety (loop-detector, anti-crush reversal) lives in the barrier
firmware. A remote human pressing "open" is an intent expression — exactly
pulseOpen. It's the fail-state-safety value, triggered by a remote human instead of a power-loss. - Constrain the how, not the whether (this is the command where [[threat-model|operator-as-
adversary]] bites hardest — a remote "let this car out free" is the classic fraud):
- Every remote open is a first-class signed ledger event (append-only-event-chain): appended, hash-chained, signed, with actor (which cloud identity), reason code, and site/relay. Control power and audit come as a pair — the same discipline setup-relay-test and booth-exit-flow already apply locally.
- A distinct, high-privilege capability, not bundled into "monitoring" — a readonly portfolio viewer can't open barriers.
- The booth stays the enforcer: cloud sends intent; the booth validates (for-me? authorized
peer? signed?) and issues
pulseOpento its own relay. Cloud never touches the relay. - Cloud can't be the sole egress path. A fully unattended site needs a local fail-open on host-loss + physical override too — offline-first means the cloud is a convenience remote-open path, not the only one, or you've recreated "device down = hostages" one layer up.
Emergent tenet: an unattended site is a higher safety bar than an attended one, not a lower one. Every local failure mode (barrier stuck, payment dead, network down) needs an answer that doesn't require the cloud; the cloud makes resolution nicer, not possible. Fold into autonomous-direction and fail-state-safety when this is picked up.
What looks straightforward (agreed quickly)
- Transport: the existing NetBird overlay (booth-dialed, nothing exposed) — not a bespoke channel. Reuses fleet-deployment-komodo.
- Sync: booth-push, verify-on-ingest — booth streams ledger + device telemetry (device-events) + snapshot metadata + financial data outbound; cloud re-verifies the chain + signatures and flags gaps.
- Staleness first-class in the UI: every site tile shows last-seen; a dark site is visibly stale.
- DB: almost certainly PostgreSQL — already the named deferred sync target (drizzle-orm, technology-stack); the Drizzle schemas are meant to port to it.
The genuinely open questions (postponed — pick up here)
- What does "real-time" mean to the buyer? Live-ish (seconds, streaming uplink → heavier booth) vs. every-few-minutes rollups (cheap, still sells "monitoring"). This gap is most of the engineering cost and drives how heavy the booth-side uplink must be.
- Financial reports computed where? Cloud re-derives revenue from the verified ledger → independently trustworthy ("we don't take the booth's word for it") but the cloud must implement the tariff pricing logic. Vs. booth sends pre-computed rollups (cheaper, but trusts the booth's math). Lean: cloud re-derives — the whole point of threat-model is not to trust the site's self-report (reporting-analytics is already "projections over the signed log").
- Hosting + licensing. The booth stack is deliberately all-MIT/Apache/BSD (technology-stack); a SaaS the user hosts has more freedom (like the fleet-deployment-komodo / vision-service self-host exceptions) — but anything that ever ships on-premise re-binds the constraint.
- Custodianship is leverage and liability. Holding other companies' financial records + device secrets is what makes the subscription sticky — and what pulls in backups, retention policy, breach disclosure, data-residency. A deliberate "yes, we want to be the custodian" call, with the obligations that implies. (Cloud/Core is a Tier-0 asset for the whole fleet — the same bar fleet-deployment-komodo already sets for Core.)
Relates
- fleet-deployment-komodo — the off-site control plane (Komodo Core + NetBird) this productises; Core-as-Tier-0 reasoning carries over.
- autonomous-direction — the unmanned future that drives remote barrier-open (Correction 2).
- reconciliation — the cloud is reconciliation, productised (verify-on-ingest, gap-evidence).
- append-only-event-chain / hardware-signer-options — why the signing key stays on the booth even as everything else centralises.
- threat-model / offline-first — the two forces every tension above traces back to.
- network-isolation / access-direction-is-per-relay — why the mesh terminates at the host.