docs(wiki): settle on-site encrypted backup + disaster-recovery design
New concept page backup-recovery.md resolving the design half of open-question #5. Driving scenario: a stolen/destroyed PC whose LUKS+TPM disk is unrecoverable by design — recovery stands up a NEW PC, restores a backup, and keeps signing the SAME chain. Settled: admin-driven encrypted full-DB backup (SQLite online-backup/VACUUM INTO, snapshots included) to local/USB, SMB/NFS, or SFTP targets; manual button + an in-process daily timer; keep-last-N + dailies retention; restore is admin-only / out-of-band (operator-adversary surface). A restored copy must still verifyChain. Key custody (the load-bearing decision, bears on #6): three independent keys — EVENT_SIGNING_KEY kept an extractable, escrowed software key DECOUPLED from the TPM so the ledger survives total hardware loss (the conscious trade: a TPM-sealed signing key would be unforgeable but permanently unverifiable after the machine dies); a NEW dedicated park_buzi_backup_key in Komodo for backup encryption, separate from the signing key; the LUKS/TPM disk key, appliance-only and deliberately non-recoverable. Keys are never inside the backup they unlock. Updated open-questions #5 (design SETTLED) + #10 note; disk-os-hardening deploy runbook (why the signing key is not sealed + park_buzi_backup_key); index catalog + concept count. Design only — not yet built. Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
This commit is contained in:
@@ -0,0 +1,128 @@
|
||||
---
|
||||
type: concept
|
||||
tags: [parking, durability, backup, recovery, security, crypto]
|
||||
sources: []
|
||||
updated: 2026-06-29
|
||||
---
|
||||
|
||||
# Backup & Disaster Recovery
|
||||
|
||||
The appliance's [[sqlite]] DB **is** the signed [[append-only-event-chain]] — the whole
|
||||
revenue/audit history. A disk failure or a stolen/destroyed PC currently means **total
|
||||
loss** (this is [[open-questions]] #5). This page is the settled design for an on-site,
|
||||
admin-driven backup that survives **total hardware loss** and restores to a fresh appliance
|
||||
with the signed chain still verifying. (Designed 2026-06-29.)
|
||||
|
||||
## The recovery scenario it must satisfy
|
||||
|
||||
The driving scenario (the one that forces every decision below): **the PC is gone** — stolen
|
||||
or destroyed. Its SSD is LUKS-encrypted and **TPM-sealed**, so the disk is unrecoverable *by
|
||||
design* (a stolen disk won't unlock off its own TPM — see [[disk-os-hardening]], [[tpm]]). We do
|
||||
**not** want the dead disk; we want to stand up a **new PC**, restore the backup, and continue
|
||||
signing the **same** chain. For that to work, recovery must depend on **(a)** the backup file and
|
||||
**(b) two keys held out-of-band** — never on the dead machine.
|
||||
|
||||
## Key custody — the load-bearing decision
|
||||
|
||||
This is the part the whole plan rests on, and it interacts with the secure-element question
|
||||
([[open-questions]] #6). Three **independent** keys, three custodians:
|
||||
|
||||
| Key | Lives | Recoverable after PC loss? | Job |
|
||||
| --- | --- | --- | --- |
|
||||
| **`EVENT_SIGNING_KEY`** | [[fleet-deployment-komodo\|Komodo]] secret (`park_buzi_event_signing_key`), escrowed offsite | **Yes — by design** | Signs + verifies the ledger chain |
|
||||
| **`park_buzi_backup_key`** *(new)* | Komodo secret, escrowed offsite, **separate** from the signing key | **Yes** | Encrypts/decrypts the backup file |
|
||||
| **LUKS / TPM disk key** | The appliance's TPM only | **No — deliberately** | At-rest protection of the powered-off SSD |
|
||||
|
||||
- **The signing key is decoupled from the TPM** — kept an *extractable software HMAC secret*
|
||||
([[append-only-event-chain]], `signer.ts`), held in Komodo and escrowed by the operator. This is a
|
||||
**conscious trade**: a truly non-extractable TPM-sealed signing key (the #6 upgrade) would make the
|
||||
ledger unforgeable even against a host-root attacker — but it would also make the **old ledger
|
||||
permanently unverifiable after total hardware loss** (the sealed key dies with the machine;
|
||||
`buildVerifier(keyId)` would return `undefined` forever). You cannot have *both* "key can never be
|
||||
extracted" *and* "I can rescue the key after the machine dies" — they are the same property from two
|
||||
sides. Against the [[threat-model|primary adversary]] (the **booth operator**, who has a UI login, not
|
||||
host root) an escrowed software key is already tamper-evident, so the recoverable design is chosen
|
||||
**today**; revisiting #6 means re-accepting the unverifiable-after-loss cost. See [[tpm]] "TPM vs.
|
||||
ATECC608", [[fleet-deployment-komodo]] (the "EVENT_SIGNING_KEY-in-Core is a fraud-root blast radius"
|
||||
caveat is the same trade).
|
||||
|
||||
- **Backup key is separate from the signing key** even though Komodo holds both — so they can be
|
||||
managed independently. Rationale: (1) the **signing key must almost never rotate** (every rotation
|
||||
fractures the chain into a new `keyId` segment — old events stay pinned to the old key forever),
|
||||
whereas the **backup key may want routine rotation** (a USB went home, a target was decommissioned);
|
||||
coupling them drags the cheap op into the expensive one. (2) The backup key **travels to every backup
|
||||
destination** (USB, NAS, SFTP); the signing key should travel *nowhere* but Komodo → process memory —
|
||||
sharing one key means every backup target conceptually exposes the signing key. (3) Keeping them
|
||||
separate keeps the **#6 TPM-migration door open** without re-wiring backups. Decided 2026-06-29
|
||||
(the "one fewer secret to escrow" simplicity of a shared key is real, but weakest here because Komodo
|
||||
already holds both).
|
||||
|
||||
> **The keys are never inside the backup they unlock.** A key can't decrypt the file it's locked in.
|
||||
> Recovery = backup file **+** both escrowed keys, supplied out-of-band. The runbook must say this
|
||||
> plainly so nobody "helpfully" stores the keys next to the backups.
|
||||
|
||||
## What a backup contains
|
||||
|
||||
**Full SQLite DB, snapshots included** — one self-contained, restore-to-identical-appliance file
|
||||
(ledger + sessions + config + subscriptions + the [[entry-exit-points|snapshot]] BLOBs). Chosen for
|
||||
completeness over size.
|
||||
|
||||
> **Size caveat (interacts with [[open-questions]] #10).** Snapshot BLOBs **dominate** DB size and
|
||||
> bloat *every* backup. They are unsigned, advisory, and already disk-pressure-pruned
|
||||
> ([[entry-exit-points]]). A future **"exclude snapshots" toggle** (ledger/sessions/config only — much
|
||||
> smaller, signed chain still fully preserved) is the obvious knob if backup size becomes a problem; the
|
||||
> default is the complete picture.
|
||||
|
||||
The backup is produced via SQLite **online-backup / `VACUUM INTO`** (a consistent snapshot of the
|
||||
live WAL-mode DB — **never a raw file copy**, which can capture a torn WAL), then encrypted with
|
||||
`park_buzi_backup_key`. **Acceptance test:** a restored copy must still pass `verifyChain` — the
|
||||
signed chain is the thing being protected, so an unverifiable restore is a failed backup.
|
||||
|
||||
## Triggers
|
||||
|
||||
- **Manual** — an admin-only **"Back up now"** button runs immediately to the configured target.
|
||||
- **Periodic** — an **in-process daily timer** (same pattern as the snapshot-retention prune,
|
||||
[[entry-exit-points]] / `snapshot-retention.ts`): runs only if the configured target is
|
||||
reachable/mounted; surfaces last-success / last-error in the UI. No OS cron — it lives inside the
|
||||
Fastify process, works inside the [[container-deployment|Docker container]], and is configured in
|
||||
one place. ([[offline-first]]: the periodic path must tolerate a missing/unmounted target without
|
||||
failing the app.)
|
||||
|
||||
## Destinations (admin-configurable)
|
||||
|
||||
All three supported in the first cut; the manual button and the periodic timer share them:
|
||||
|
||||
- **Local / USB / SATA disk** — a mounted path on an attached disk. Simplest, fully offline, matches
|
||||
the air-gapped appliance. The strong first target.
|
||||
- **Network drive (SMB/NFS)** — a mounted share on the isolated LAN (a site NAS). Still
|
||||
local-network, no internet ([[network-isolation]]).
|
||||
- **SFTP** — push to an SFTP endpoint, useful for an offsite copy. **FTP is excluded** (plaintext
|
||||
credentials + data); SFTP is the safe equivalent.
|
||||
|
||||
## Retention at the destination
|
||||
|
||||
**Keep last N + thinned dailies** (e.g. last 7 daily / last 4 weekly) — bounded disk use, and it
|
||||
survives the "a bad/partial run clobbered the only good copy" failure. (A single rolling
|
||||
overwrite-latest file was rejected for exactly that reason.)
|
||||
|
||||
## Threat-model fit — restore is the dangerous half
|
||||
|
||||
Writing a backup is benign; **restore is operator-adversary surface** ([[threat-model]]). A restored
|
||||
DB *replaces* the live signed chain — so a malicious restore is a way to swap in a doctored history.
|
||||
Therefore:
|
||||
|
||||
- **Restore is NOT a booth button.** It is an **admin-only, out-of-band runbook action** (new
|
||||
appliance, deliberate provisioning step), not something reachable from the operator console.
|
||||
- The backup **target configuration** and the **"Back up now"** action are admin-gated.
|
||||
- Backups do **not** weaken the chain's tamper-evidence: a restored chain is re-verified with the
|
||||
escrowed `EVENT_SIGNING_KEY`; a tampered restore fails `verifyChain` just as a tampered live DB
|
||||
would. The backup is a **durability** control, not an integrity one — integrity stays with the
|
||||
signed chain + [[reconciliation]].
|
||||
|
||||
## Status
|
||||
|
||||
Design settled 2026-06-29; **not yet built**. Resolves the *design* half of [[open-questions]] #5
|
||||
(implementation pending), and records the key-custody stance that bears on #6 (signing stays
|
||||
decoupled from the TPM) and #10 (snapshots bloat backups → future exclude toggle). See
|
||||
[[append-only-event-chain]], [[disk-os-hardening]], [[tpm]], [[fleet-deployment-komodo]],
|
||||
[[reconciliation]].
|
||||
@@ -2,7 +2,7 @@
|
||||
type: concept
|
||||
tags: [parking, security, platform]
|
||||
sources: [parking-system-architecture]
|
||||
updated: 2026-06-21
|
||||
updated: 2026-06-29
|
||||
---
|
||||
|
||||
# Disk / OS Hardening
|
||||
@@ -44,7 +44,13 @@ Env in `apps/server/.env` on the appliance (see `apps/server/.env.example`). The
|
||||
- **`JWT_SECRET`** — ≥32 random chars; the server refuses to boot without a strong one (no
|
||||
insecure default). `openssl rand -hex 32`. See [[local-jwt-auth]].
|
||||
- **`EVENT_SIGNING_KEY`** — dedicated HMAC key for the signed ledger; ≥16 chars. Falls back to
|
||||
`JWT_SECRET` with a warning if unset — set a dedicated one before production.
|
||||
`JWT_SECRET` with a warning if unset — set a dedicated one before production. Deliberately kept an
|
||||
**extractable, escrowed software key (NOT TPM-sealed)** so the ledger survives total hardware loss —
|
||||
see [[backup-recovery]] for the custody trade vs. [[open-questions]] #6.
|
||||
- **`park_buzi_backup_key`** *(planned)* — dedicated key for encrypting [[backup-recovery|DB backups]],
|
||||
**separate** from `EVENT_SIGNING_KEY` (independent rotation; backups travel, the signing key
|
||||
shouldn't). Both escrowed offsite in [[fleet-deployment-komodo|Komodo]]; recovery needs both, held
|
||||
out-of-band.
|
||||
- **`COOKIE_SECURE=0`** — **REQUIRED on the plain-HTTP LAN appliance.** Auth/CSRF cookies are
|
||||
`Secure` by **default** (fail-safe). The appliance serves the SPA same-origin over **plain
|
||||
http** on the booth LAN, where a `Secure` cookie is **never sent** — so without this opt-out
|
||||
|
||||
Reference in New Issue
Block a user