Files
parking_solution/wiki/concepts/backup-recovery.md
T
julian 2910672b5a fix(backup): persist last-success/error status; wall-clock-based schedule
BackupService tracked last-success/last-error as plain in-process fields
and scheduled the daily backup via setInterval measured from process
start — so any server restart (deploy/crash/OOM/reboot, routine under
`restart: always`) silently reset the admin UI to "last successful
backup: Never" and drifted the actual cadence, independent of whether
backups were writing correctly to disk (they were — a real field
incident at park-buzi showed 7 valid rotating backups on disk with the
status stuck on "Never").

Persist last-success/error to new site_config columns (migration 0025)
and add BackupService.isDue(), computed from the persisted timestamp
instead of process uptime; server.ts now polls every 15 min and lets
isDue() gate the actual run. No API/UI contract change.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
2026-08-30 18:11:23 +02:00

275 lines
19 KiB
Markdown

---
type: concept
tags: [parking, durability, backup, recovery, security, crypto]
sources: []
updated: 2026-08-30
---
# Backup & Disaster Recovery
The appliance's [[sqlite]] DB **is** the signed [[append-only-event-chain]] — the whole
revenue/audit history. A disk failure or a stolen/destroyed PC currently means **total
loss** (this is [[open-questions]] #5). This page is the settled design for an on-site,
admin-driven backup that survives **total hardware loss** and restores to a fresh appliance
with the signed chain still verifying. (Designed 2026-06-29.)
## The recovery scenario it must satisfy
The driving scenario (the one that forces every decision below): **the PC is gone** — stolen
or destroyed. Its SSD is LUKS-encrypted and **TPM-sealed**, so the disk is unrecoverable *by
design* (a stolen disk won't unlock off its own TPM — see [[disk-os-hardening]], [[tpm]]). We do
**not** want the dead disk; we want to stand up a **new PC**, restore the backup, and continue
signing the **same** chain. For that to work, recovery must depend on **(a)** the backup file and
**(b) two keys held out-of-band** — never on the dead machine.
## Key custody — the load-bearing decision
This is the part the whole plan rests on, and it interacts with the secure-element question
([[open-questions]] #6). Three **independent** keys, three custodians:
| Key | Lives | Recoverable after PC loss? | Job |
| --- | --- | --- | --- |
| **`EVENT_SIGNING_KEY`** | [[fleet-deployment-komodo\|Komodo]] secret (`park_buzi_event_signing_key`), escrowed offsite | **Yes — by design** | Signs + verifies the ledger chain |
| **`park_buzi_backup_key`** *(new)* | Komodo secret, escrowed offsite, **separate** from the signing key | **Yes** | Encrypts/decrypts the backup file |
| **LUKS / TPM disk key** | The appliance's TPM only | **No — deliberately** | At-rest protection of the powered-off SSD |
- **The signing key is decoupled from the TPM** — kept an *extractable software HMAC secret*
([[append-only-event-chain]], `signer.ts`), held in Komodo and escrowed by the operator. This is a
**conscious trade**: a truly non-extractable TPM-sealed signing key (the #6 upgrade) would make the
ledger unforgeable even against a host-root attacker — but it would also make the **old ledger
permanently unverifiable after total hardware loss** (the sealed key dies with the machine;
`buildVerifier(keyId)` would return `undefined` forever). You cannot have *both* "key can never be
extracted" *and* "I can rescue the key after the machine dies" — they are the same property from two
sides. Against the [[threat-model|primary adversary]] (the **booth operator**, who has a UI login, not
host root) an escrowed software key is already tamper-evident, so the recoverable design is chosen
**today**; revisiting #6 means re-accepting the unverifiable-after-loss cost. See [[tpm]] "TPM vs.
ATECC608", [[fleet-deployment-komodo]] (the "EVENT_SIGNING_KEY-in-Core is a fraud-root blast radius"
caveat is the same trade).
- **Backup key is separate from the signing key** even though Komodo holds both — so they can be
managed independently. Rationale: (1) the **signing key must almost never rotate** (every rotation
fractures the chain into a new `keyId` segment — old events stay pinned to the old key forever),
whereas the **backup key may want routine rotation** (a USB went home, a target was decommissioned);
coupling them drags the cheap op into the expensive one. (2) The backup key **travels to every backup
destination** (USB, NAS, SFTP); the signing key should travel *nowhere* but Komodo → process memory —
sharing one key means every backup target conceptually exposes the signing key. (3) Keeping them
separate keeps the **#6 TPM-migration door open** without re-wiring backups. Decided 2026-06-29
(the "one fewer secret to escrow" simplicity of a shared key is real, but weakest here because Komodo
already holds both).
> **The keys are never inside the backup they unlock.** A key can't decrypt the file it's locked in.
> Recovery = backup file **+** both escrowed keys, supplied out-of-band. The runbook must say this
> plainly so nobody "helpfully" stores the keys next to the backups.
## What a backup contains
**Full SQLite DB, snapshots included** — one self-contained, restore-to-identical-appliance file
(ledger + sessions + config + subscriptions + the [[entry-exit-points|snapshot]] BLOBs). Chosen for
completeness over size.
> **Size caveat (interacts with [[open-questions]] #10).** Snapshot BLOBs **dominate** DB size and
> bloat *every* backup. They are unsigned, advisory, and already disk-pressure-pruned
> ([[entry-exit-points]]). A future **"exclude snapshots" toggle** (ledger/sessions/config only — much
> smaller, signed chain still fully preserved) is the obvious knob if backup size becomes a problem; the
> default is the complete picture.
The backup is produced via SQLite **online-backup / `VACUUM INTO`** (a consistent snapshot of the
live WAL-mode DB — **never a raw file copy**, which can capture a torn WAL), then encrypted with
`park_buzi_backup_key`. **Acceptance test:** a restored copy must still pass `verifyChain` — the
signed chain is the thing being protected, so an unverifiable restore is a failed backup.
## Triggers
- **Manual** — an admin-only **"Back up now"** button runs immediately to the configured target.
- **Periodic** — an **in-process daily timer** (same pattern as the snapshot-retention prune,
[[entry-exit-points]] / `snapshot-retention.ts`): runs only if the configured target is
reachable/mounted; surfaces last-success / last-error in the UI. No OS cron — it lives inside the
Fastify process, works inside the [[container-deployment|Docker container]], and is configured in
one place. ([[offline-first]]: the periodic path must tolerate a missing/unmounted target without
failing the app.)
## Destinations (admin-configurable)
All three supported in the first cut; the manual button and the periodic timer share them:
- **Local / USB / SATA disk** — a mounted path on an attached disk. Simplest, fully offline, matches
the air-gapped appliance. The strong first target.
- **Network drive (SMB/NFS)** — a mounted share on the isolated LAN (a site NAS). Still
local-network, no internet ([[network-isolation]]).
- **SFTP** — push to an SFTP endpoint, useful for an offsite copy. **FTP is excluded** (plaintext
credentials + data); SFTP is the safe equivalent.
### The target must be a bind-mounted host path — NOT a casually-plugged USB (2026-06-29)
The server runs **inside the `parking-server` container**, so it can only `stat()`/write paths that
are **bind-mounted into that container**. A USB stick the operator plugs in lands at a desktop
auto-mount path on the *host* (`/run/media/<user>/<UUID>`), which **does not exist inside the
container** — so the in-UI **Test target** correctly reports *"location does not exist."* This bit on
the first booth deploy (2026-06-29): `BACKUP_KEY` was finally injected, then the target test failed
because the USB path wasn't visible to the process.
**So a backup destination is provisioned by the ADMIN at the host level, not chosen ad hoc by the
operator.** The procedure:
1. Attach the disk (external HDD/SSD/USB) and mount it at a **stable host path** (e.g. `/mnt/backup`)
via **`/etc/fstab` by UUID** — *not* the desktop automounter, whose UUID-named path changes per
drive and vanishes on unplug.
2. **Bind-mount that host path into the container** in the prod compose (e.g.
`/mnt/backup:/mnt/backup` on the `server` service — same pattern as the `/dev/usb` printer
passthrough in [[container-deployment]]).
3. In the UI (Setup → Backup), set the **target directory to the in-container path** (`/mnt/backup`)
and **Test target** — now writable.
> **This is partly a feature, not just a limitation** ([[threat-model]]): because the destination is
> a host-provisioned bind-mount, the **booth operator cannot redirect backups to a removable stick
> they walk off with** — real destinations are an admin/host decision, on the trusted side of the
> [[trust-boundary]]. A network share (SMB/NFS) is the same shape: mount on the host, bind-mount in.
>
> **Limitation acknowledged:** the backup target is therefore **not operator-flexible** — you cannot
> just plug in a USB and back up from the UI. Adding a new destination = a host `fstab` + compose
> bind-mount change + redeploy. For the appliance model (single-purpose, admin-provisioned) this is
> the right trade; a future "back up to a freshly-plugged removable drive" flow would need host-level
> automount detection wired to the container, which is **deferred / not built**.
## Retention at the destination
**Keep last N + thinned dailies** (e.g. last 7 daily / last 4 weekly) — bounded disk use, and it
survives the "a bad/partial run clobbered the only good copy" failure. (A single rolling
overwrite-latest file was rejected for exactly that reason.)
## Threat-model fit — restore is the dangerous half
Writing a backup is benign; **restore is operator-adversary surface** ([[threat-model]]). A restored
DB *replaces* the live signed chain — so a malicious restore is a way to swap in a doctored history.
Therefore:
- **Restore is NOT a booth button.** It is an **admin-only, out-of-band runbook action** (new
appliance, deliberate provisioning step), not something reachable from the operator console.
- The backup **target configuration** and the **"Back up now"** action are admin-gated.
- Backups do **not** weaken the chain's tamper-evidence: a restored chain is re-verified with the
escrowed `EVENT_SIGNING_KEY`; a tampered restore fails `verifyChain` just as a tampered live DB
would. The backup is a **durability** control, not an integrity one — integrity stays with the
signed chain + [[reconciliation]].
## As-built (2026-06-29) — engine + local/mounted target
The first slice is **built and tested**: the backup **engine + a local/mounted target + the daily
timer + the manual route**. What landed:
- **`apps/server/src/backup.ts`** — the engine. Consistent online copy via better-sqlite3's native
`.backup()` (a transactionally-consistent snapshot of the live WAL DB — **not** a raw file copy),
then **AES-256-GCM** encryption with a **scrypt-derived** key from `BACKUP_KEY`. Self-describing
header (`magic | version | salt | iv | … | authTag`) so a restore tool needs only the key + the file
— **zero new dependencies** (Node `crypto`). The plaintext intermediate is written to **scratch**
(not the removable/network target) and **wiped in a `finally`**, success or fail. Retention =
**keep-last-N + one-per-day-within-N-days** (`pruneOldBackups`). Tested: round-trip decrypts to a
**byte-identical, queryable DB**; a flipped byte or wrong key **fails GCM auth**; short key rejected;
scratch plaintext always removed.
- **`backup-service.ts`** — the **target directory AND retention are admin-chosen** in the UI
(`site_config.backup_target_dir`, migration 0016; `backup_keep_last` + `backup_keep_daily_days`,
migration 0017) and read **fresh each run**, so changing them takes effect with no restart. Retention
columns are nullable → fall back to the code default (keep-last 7, keep-daily 30) per field. The
**encryption key is the ONLY backup env/Komodo secret** (`BACKUP_KEY`) — a key must never live in the
DB it backs up; target+retention are operational policy, not secrets. The service **serializes**
concurrent runs (single in-flight guard) and records last-success / last-error; `status()` exposes
`targetDir`, `keepLast`, `keepDailyDays` + `keyPresent` so the UI distinguishes "no target" from
"no key".
- **`routes/backup.ts`** — `GET /api/backup/status` (`backup:read`); `PUT /api/backup/config` to set/
clear the target (`backup:update`); `POST /api/backup/test` to probe a candidate path server-side —
exists / is-a-dir / writable (`backup:update`); `POST /api/backup/run` (`backup:create`), a clean
**409 `backup_not_configured`** when target+key aren't both set. New `backup` permission resource
(`backup:read/update/create`) in `@parking/shared`. **No restore route** — out-of-band by design.
- **`apps/web/src/BackupSettings.tsx`** — a Setup → **Backup** tab (gated `backup:read`): an editable
**target-path field** with a **Test target** probe (localized ok/missing/not-a-dir/not-writable),
**retention fields** (keep-last / keep-daily-days), one **Save**, the status panel (config state,
last-run size/pruned/error, a distinct amber **missing BACKUP_KEY** warning), a **Back up now**
button, and the restore-is-out-of-band note. Full i18n (sq + en).
- **Komodo wiring.** `BACKUP_KEY` is a **per-booth Komodo secret** (`[[park_buzi_backup_key]]` in
`komodo/resources.toml`; documented in `komodo/.env.komodo.example`), escrowed offsite alongside
`EVENT_SIGNING_KEY`. It is the *only* backup env var — target + retention are in the DB.
> **Gotcha — compose `environment:` is an ALLOWLIST (cost a full booth-deploy session, 2026-06-29).**
> Wiring `BACKUP_KEY` as a Komodo secret + Stack-env line is **necessary but not sufficient**:
> `docker-compose.yml`'s `server.environment:` block only forwards the variables it *names*. The key
> was wired everywhere (secret store, Stack env, `.env.example`, schema) but **never added to that
> compose block**, so the container came up *without* it — `docker inspect ...Config.Env` showed
> `JWT_SECRET`/`EVENT_SIGNING_KEY` present and `BACKUP_KEY` **absent (not empty)**, while the Backup
> screen correctly reported "BACKUP_KEY missing". Diagnosis was muddied by chasing Komodo (secret
> name, re-sync, destroy/redeploy, env-only-change-doesn't-recreate) before checking the compose
> allowlist. **Lesson: a new server env var needs a line in `docker-compose.yml` `server.environment:`
> too — that's the only place env reaches the container.** Fixed: `BACKUP_KEY: ${BACKUP_KEY:-}` next to
> `EVENT_SIGNING_KEY`. Quick check on a booth: `docker inspect <server> --format '{{range .Config.Env}}{{println .}}{{end}}' | grep -i backup`.
- **`server.ts`** — an **unref'd daily timer** (`backupService.runScheduled`), a **no-op until
configured**, and **deliberately NOT run at startup** (a just-power-cut booth shouldn't write to a
possibly-unmounted disk; the daily cadence + the manual button cover it).
- Env documented in `apps/server/.env.example` (with the escrow + separate-key notes).
**SMB/NFS already work** — they're just a mounted path the admin enters as the target. **Deferred to
follow-up slices:** an **SFTP** target and a **restore runbook / CLI**.
## Field bug — "last successful backup: Never" despite valid, rotating backups on disk (found + fixed 2026-08-30)
**Symptom (park-buzi):** the admin noticed the backup directory held 7 real, correctly-sized,
correctly-rotating encrypted backups (`parking-backup-*.sqlite.enc`, retention working exactly as
designed) — yet the Backup screen's "Kopja e fundit e suksesshme" (last successful backup) showed
**"Asnjëherë" (Never)**. Separately, the most recent file was 2 days old rather than ~1.
**Root cause — two independent, disconnected code paths, both traced to `setInterval`-since-
process-start:**
1. **Status was never persisted.** `BackupService` tracked `lastSuccessAt`/`lastResult`/
`lastErrorAt`/`lastError` as **plain in-process private fields** — set only inside `run()`,
read only by `status()` on the *same running instance*. Nothing wrote them to `site_config` or
anywhere else durable. The actual backup-writing engine (`backup.ts`: consistent copy → encrypt
→ `pruneOldBackups`) is a completely separate code path that only touches the filesystem and
has no notion of this status object. So "7 valid files on disk" and "status says Never" were
never contradictory — they were two unrelated signals, and **any** server restart (deploy,
crash, OOM, host reboot — all routine under `restart: always` in `docker-compose.prod.yml`)
silently reset the in-memory fields to `null` regardless of what had actually happened on disk.
2. **The schedule was measured from process start, not from the last real backup.** The daily
timer was `setInterval(() => backupService.runScheduled(), 24h)` — a fixed 24h period counted
from whenever the *process* last started, not from wall-clock time or from when a backup last
actually succeeded. The exact same restart that wiped the in-memory status also reset this
countdown, which is why the cadence can silently drift or skip past a day with no error ever
surfacing anywhere.
Both symptoms are one cause: **the server process restarted after the Aug 28 backup, and nothing
about this design was built to survive that.**
### Fix (2026-08-30)
- **`packages/db/src/schema.ts`** / migration `0025_backup_last_status.sql` — four new nullable
`site_config` columns: `backup_last_success_at`, `backup_last_result_json`,
`backup_last_error_at`, `backup_last_error`. Same table, same upsert pattern as
`backup_target_dir`/`backup_keep_last`/`backup_keep_daily_days` (migrations 0016/0017).
- **`backup-service.ts`** — `run()` now writes success/error outcomes to these columns (via a
`#persist` upsert helper) instead of private fields; `status()` reads them fresh from the DB on
every call. A brand-new `BackupService` instance (i.e. a fresh process) now sees exactly what
the previous instance last recorded — no more restart amnesia.
- **New `isDue(now, intervalMs = 24h)`** method: due iff `now - backupLastSuccessAt >= 24h` (or
immediately due if no success was ever recorded), computed from the **persisted** timestamp —
never from process uptime.
- **`server.ts`** — the daily `setInterval` was replaced with a **15-minute poll** calling
`runScheduled()`, which now itself no-ops unless `isDue()` is true. This makes the actual backup
cadence immune to restart timing entirely: however often the process happens to restart, the
next backup fires within 15 minutes of 24h having genuinely elapsed since the last real success
— not 24h after whatever moment the process most recently came back up.
- Covered by a new `backup-service.test.ts`: a fresh `BackupService` over the same DB handle
(simulating a restart) sees the prior instance's last success/error and its cleared-on-success
behavior; `isDue()` is exercised directly against injected timestamps rather than real sleeps.
No change to the `BackupStatus` shape returned by `GET /api/backup/status` or to
`BackupSettings.tsx` — this was purely a durability fix underneath the same contract.
## Status
Design settled 2026-06-29; **engine + admin-configured local/mounted target + admin UI BUILT
2026-06-29** (SFTP + restore tooling pending). The target directory is **admin-chosen in the UI**
(`site_config`, migration 0016), not an env var — the on-site admin picks where backups land; only
`BACKUP_KEY` stays a server secret. **Last-success/last-error status + the scheduling cadence are
now restart-durable (migration 0025, 2026-08-30)** — see field bug above. Resolves the *design*
half of [[open-questions]] #5 and the first build slices; records the key-custody stance that bears
on #6 (signing stays decoupled from the TPM) and #10 (snapshots bloat backups → future exclude
toggle). See [[append-only-event-chain]], [[disk-os-hardening]], [[tpm]], [[fleet-deployment-komodo]],
[[reconciliation]].