ee61c24bb9
Lab box (park-test) crash-looped: panic writing the agent key to /etc/komodo/keys/periphery.key (Permission denied). Gotcha #9 was framed as a hand-config hazard; v2.2.0's installer now writes the system-style default even with --user. §7a: verify root_directory after every install + sed fix + reset-failed/restart; user-unit vs sudo note; the onboarding key survives a pre-connect crash. Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
487 lines
29 KiB
Markdown
487 lines
29 KiB
Markdown
---
|
||
type: reference
|
||
tags: [parking, deployment, appliance, hardening, runbook, offline-first]
|
||
sources: []
|
||
updated: 2026-07-06
|
||
status: settled
|
||
---
|
||
|
||
# Appliance provisioning runbook (booth PC)
|
||
|
||
Step-by-step to take a booth PC from factory Windows to a hardened, encrypted, container-running
|
||
parking appliance. Written from the **first real provisioning, 2026-06-23** (hardening) +
|
||
**first Komodo deploy, 2026-06-27** (the runtime) — every command here was run and verified on the
|
||
actual hardware, including the firmware-specific workaround. Companion to [[disk-os-hardening]] (the
|
||
*why*), [[tpm]] (TPM analysis), [[container-deployment]] (the images), and
|
||
[[fleet-deployment-komodo]] (the deploy control plane this runbook's §7 uses).
|
||
|
||
> ⚠ This box is the [[threat-model|outsider-with-the-box]] defence. The load-bearing anti-fraud
|
||
> control is still [[reconciliation]] over the [[append-only-event-chain|signed chain]] — disk
|
||
> encryption + Secure Boot raise the cost of offline tamper, they don't replace reconciliation.
|
||
|
||
## Reference hardware (first unit, 2026-06-23)
|
||
|
||
- **Dell OptiPlex 7070**, **Intel Core i5-8500** (Coffee Lake), 238 GB SATA SSD (`/dev/sda`).
|
||
- **TPM 2.0 — discrete Nuvoton** (`Get-Tpm` → `ManufacturerIdTxt NTC`, fw 7.2.1.0). NOT Intel
|
||
PTT/fTPM. Discrete ⇒ an external LPC/SPI bus exists (bus-sniff is a theoretical physical attack on
|
||
PCR-only sealing — accepted; see [[tpm]]). Used PC — previous owner irrelevant.
|
||
- Shipped Windows 11; formatted to **Ubuntu 26.04 LTS** (the decided platform — [[desktop-shell-tauri]]).
|
||
|
||
## 1. BIOS (F2 at the Dell logo)
|
||
|
||
- **TPM**: leave **On**. Used PC → **Clear the TPM once** (Security → TPM → Clear) so the prior
|
||
owner's keys are wiped before LUKS enrollment. (PPI "Bypass for Clear" was unchecked → it asks for
|
||
physical confirmation at next boot; that's normal.)
|
||
- **Secure Boot**: **Enabled**, **Deployed Mode** (not Audit). **"Enable Custom Mode" UNCHECKED** =
|
||
Standard Mode with stock Microsoft keys — this is what Ubuntu's signed shim needs. Do NOT touch
|
||
PK/KEK/db/dbx. NB: the 7070's Expert Key Management is **edit-only** (Save/Replace/Append/Delete —
|
||
no read-only "View Key"), so you **cannot inspect db from BIOS**; verify via the live USB instead
|
||
(step 2).
|
||
- **Boot**: UEFI only (no CSM/Legacy — a Legacy install has no Secure Boot / TPM-seal path).
|
||
- Set a **BIOS admin password**.
|
||
|
||
## 2. Boot the Ubuntu 26.04 USB (Secure Boot ON)
|
||
|
||
- **Flash the ISO DIRECTLY** (Rufus GPT/UEFI, Etcher, or `dd`). **NOT Ventoy** — Ventoy's own
|
||
bootloader isn't in `db`, so Secure Boot rejects it with **`Verification failed: (0x1A) Security
|
||
Violation`** (this is Secure Boot working correctly, not a fault). A directly-flashed Ubuntu USB
|
||
boots the Microsoft-signed shim, which stock `db` trusts.
|
||
- **F12** at the Dell logo → pick the USB under **UEFI BOOT**.
|
||
- Reaching the installer with Secure Boot ON = positive proof the MS third-party UEFI CA is in `db`
|
||
(the verification the BIOS couldn't show us).
|
||
|
||
## 3. Encrypted install — the firmware workaround (IMPORTANT)
|
||
|
||
The 26.04 installer disk page offers: No Encryption / **Encrypt with a passphrase** / **Use
|
||
hardware-backed encryption** (+ advanced LVM/ZFS, both ZFS experimental).
|
||
|
||
- **"Use hardware-backed encryption" FAILS on this 7070** with:
|
||
`PCR_UNUSABLE … error with secure boot policy (PCR7) measurements: generating secure boot profiles
|
||
for systems with timestamp revocation (dbt) support is currently not supported.`
|
||
→ Ubuntu's *automated* FDE profiler can't model PCR7 on Dell firmware carrying a `dbt` (UEFI
|
||
timestamp revocation list). It is NOT a TPM or Secure-Boot fault — both are fine.
|
||
- **So: choose "Encrypt with a passphrase".** Set a strong passphrase and **SAVE IT OFF-MACHINE**
|
||
(phone / password manager). It is both the boot unlock (until TPM sealing) AND the permanent
|
||
recovery slot. Finish the install.
|
||
- Result (verify with `lsblk`): `sda1` vfat `/boot/efi`, `sda2` ext4 `/boot`, `sda3` `crypto_LUKS`
|
||
→ `dm_crypt-0` (LVM2) → `ubuntu--vg-ubuntu--lv` ext4 `/`.
|
||
|
||
## 4. Seal LUKS to the TPM (manual — PCR 7 only)
|
||
|
||
Do this AFTER first boot. Manual enrollment sidesteps the installer's dbt profiler and lets us pick
|
||
PCRs. **Bind to PCR 7 only** (Secure Boot state): it catches the attack that matters (disabling
|
||
Secure Boot to boot a tampered kernel) WITHOUT breaking on routine kernel/GRUB updates (which churn
|
||
PCRs 4/8/9 → would otherwise drop every boot to the passphrase). Firmware-only PCR 0 is the fallback
|
||
if PCR 7 ever errors.
|
||
|
||
```bash
|
||
sudo apt update && sudo apt install -y tpm2-tools
|
||
sudo tpm2_pcrread sha256 # sanity: PCRs 0-10 populated, PCR 7 has a real value
|
||
|
||
# Enroll the TPM (prompts for the EXISTING install passphrase to authorize the new slot):
|
||
sudo systemd-cryptenroll --tpm2-device=auto --tpm2-pcrs=7 /dev/sda3
|
||
|
||
# Verify TWO slots — keep BOTH (slot 0 password = recovery, slot 1 tpm2 = auto-unlock):
|
||
sudo systemd-cryptenroll /dev/sda3
|
||
# SLOT TYPE
|
||
# 0 password
|
||
# 1 tpm2
|
||
```
|
||
|
||
Wire it into boot (back up first; the mapping is `dm_crypt-0`, the LUKS UUID is in `/etc/crypttab`):
|
||
|
||
```bash
|
||
sudo cp /etc/crypttab /etc/crypttab.bak
|
||
sudo sed -i 's/none luks$/none luks,tpm2-device=auto/' /etc/crypttab
|
||
cat /etc/crypttab # → dm_crypt-0 UUID=… none luks,tpm2-device=auto
|
||
sudo update-initramfs -u
|
||
sudo reboot
|
||
```
|
||
|
||
- **Boots straight to login, no passphrase prompt** = ✅ TPM auto-unlock works (unattended reboot
|
||
achieved — VERIFIED on this unit 2026-06-23).
|
||
- Still prompts = PCR mismatch; type the passphrase (NOT locked out), then retry with
|
||
`--tpm2-pcrs=0`. The `password` slot + `crypttab.bak` make this fully reversible.
|
||
|
||
> **Re-seal runbook:** a BIOS update / Secure Boot change / **UEFI dbx (revocation list) update**
|
||
> alters PCR 7 → the TPM refuses → boot falls back to the passphrase prompt (not a brick). After
|
||
> such a change, re-run step 4's
|
||
> `systemd-cryptenroll --wipe-slot=tpm2 --tpm2-device=auto --tpm2-pcrs=7 /dev/sda3` to re-bind, then
|
||
> reboot to confirm unattended unlock returned.
|
||
|
||
### 4a. Firmware / UEFI dbx updates break PCR 7 — and are an OPERATOR threat (VERIFIED 2026-06-30)
|
||
|
||
The PCR-7 re-seal hazard above is **not** a rare event — the most common trigger is a **UEFI `dbx`
|
||
(Secure Boot revocation database) update**, and it bit the real `park-buzi` booth on 2026-06-28:
|
||
|
||
- **What `dbx` is:** the Secure Boot blocklist of known-vulnerable bootloader/shim hashes
|
||
(vendor = Microsoft). It is delivered by **`fwupd`/LVFS — a channel SEPARATE from APT** (the GNOME
|
||
"Firmware Updater", which on Ubuntu is the **`firmware-updater` snap**, surfaces it). `apt list
|
||
--upgradable` being clean does NOT mean a firmware/dbx update isn't pending.
|
||
- **The GRUB panic (root cause):** applying a *new* dbx against a *stale* GRUB/shim revokes the
|
||
installed bootloader → Secure Boot refuses to load it → **unbootable / GRUB "panic"**. The fix is
|
||
ordering: `apt full-upgrade` (current `grub-efi`/`shim-signed`) FIRST, *then* dbx. A fresh reinstall
|
||
ships a current GRUB, so reinstalling recovers it.
|
||
- **It moves PCR 7:** even with a current GRUB, applying dbx changes the Secure-Boot-policy
|
||
measurement → the TPM (slot 1) refuses to release the key → next boot **drops to the slot-0
|
||
passphrase prompt**. Recover with the re-seal runbook above. VERIFIED: on `park-buzi` the dbx
|
||
update went through, the box rebooted to a passphrase prompt, the slot-0 passphrase unlocked it,
|
||
and `systemd-cryptenroll --wipe-slot=tpm2 … --tpm2-pcrs=7` restored silent auto-unlock.
|
||
|
||
**Threat-model consequence ([[threat-model]]: the operator is the adversary).** A firmware/dbx update
|
||
on a TPM-sealed booth → the booth won't boot unattended and needs the slot-0 passphrase. So the
|
||
operator must be unable to *trigger* a firmware update, and must never hold the passphrase. Lock it
|
||
down (DONE on `park-buzi` 2026-06-30):
|
||
|
||
```bash
|
||
# 1. Kill the firmware-update DAEMON (the GUI "Update" button then fails with no daemon):
|
||
sudo systemctl mask fwupd.service fwupd-refresh.timer
|
||
systemctl is-enabled fwupd.service fwupd-refresh.timer # → masked / masked (persists across reboot)
|
||
|
||
# 2. Remove the operator-facing GUI so the screen is never even presented (Ubuntu = a snap):
|
||
sudo snap remove firmware-updater
|
||
snap list | grep -i firmware # → no output (re-check: seeded snaps can re-install)
|
||
```
|
||
|
||
Plus: the **BIOS admin password** (§1) must gate *entering setup / changing settings* (a
|
||
supervisor/admin password, not just a boot password) so the operator can't disable Secure Boot or
|
||
change boot order — either of which also breaks the seal. And the **slot-0 passphrase stays
|
||
off-machine / escrowed** (same custody as `EVENT_SIGNING_KEY` / `BACKUP_KEY`); it is an admin-only
|
||
recovery secret, used on-site during a maintenance window, never known to operators.
|
||
|
||
> **Net:** firmware/dbx updates become an **admin-only, on-site, deliberate** action. The booth is
|
||
> unattended-bootable only while the firmware/Secure-Boot state is frozen — that is the security
|
||
> property, not a bug. Legitimate firmware maintenance now costs: physical presence + the slot-0
|
||
> passphrase + a PCR-7 re-enroll.
|
||
|
||
> **⚠ Gotcha — `cryptsetup … --test-passphrase` SILENTLY passes via the TPM.** Before any
|
||
> firmware/dbx change, you must *prove a typed passphrase still unlocks the disk* (the TPM-independent
|
||
> safety net). But `sudo cryptsetup open --test-passphrase /dev/sda3` with a TPM2 token enrolled will
|
||
> succeed **without prompting** — the TPM auto-answers (it unlocks the tpm2 *slot*, e.g. slot 1), a
|
||
> FALSE positive that proves nothing about a human-typeable key. Force a real test with
|
||
> `--disable-external-tokens` (→ `No usable token is available.` then it prompts; success on slot 0 =
|
||
> the passphrase genuinely works):
|
||
> ```bash
|
||
> sudo cryptsetup open --test-passphrase /dev/sda3 --disable-external-tokens --verbose
|
||
> ```
|
||
|
||
## 5. GRUB password — EDIT-ONLY (VERIFIED 2026-06-23)
|
||
|
||
Closes the `init=/bin/bash` / `systemd.unit=rescue.target` local-root hole: without it, anyone at
|
||
the keyboard presses `e` at the GRUB menu, edits the kernel cmdline, and boots to a **root shell with
|
||
no login**. **The PCR-7 TPM seal does NOT cover this** — editing the GRUB cmdline doesn't change
|
||
PCR 7 (Secure Boot policy), so the TPM still releases the key and the attacker lands on the decrypted
|
||
disk. This is the specific countermeasure for the [[threat-model|operator-at-the-booth]]. Use
|
||
**edit-only** mode (`--unrestricted`) so the box still boots UNATTENDED — the password is required
|
||
only to EDIT entries, never to boot.
|
||
|
||
```bash
|
||
grub-mkpasswd-pbkdf2 # enter a password (twice) → copy the grub.pbkdf2.sha512.* hash
|
||
```
|
||
|
||
Add the superuser (paste YOUR hash) to the end of `/etc/grub.d/40_custom`:
|
||
```
|
||
set superusers="admin"
|
||
password_pbkdf2 admin grub.pbkdf2.sha512.10000.<YOUR_HASH>
|
||
```
|
||
|
||
Make menu entries bootable WITHOUT the password (edit-only) — in `/etc/grub.d/10_linux`, set the
|
||
active `CLASS=` line to include `--unrestricted`:
|
||
```
|
||
CLASS="--class gnu-linux --class gnu --class os --unrestricted"
|
||
```
|
||
|
||
Regenerate + VERIFY BOTH HALVES landed in the real config BEFORE rebooting (a GRUB misconfig means a
|
||
rescue-USB recovery):
|
||
```bash
|
||
sudo update-grub
|
||
sudo grep -c "password_pbkdf2" /boot/grub/grub.cfg # want ≥1 (password present)
|
||
sudo grep -c "unrestricted" /boot/grub/grub.cfg # want ≥1 (entries bootable w/o password)
|
||
sudo reboot
|
||
```
|
||
|
||
✅ VERIFIED on this unit: boots straight to login (no GRUB prompt, TPM still auto-unlocks) AND
|
||
pressing `e` at the menu prompts for `admin` + password. Store the GRUB password off-machine
|
||
(alongside the LUKS passphrase).
|
||
|
||
> OS hardening on the first unit is now COMPLETE: LUKS FDE + TPM auto-unlock (PCR 7) + Secure Boot
|
||
> (Deployed) + GRUB edit-lock.
|
||
|
||
## 5c. OS user model — admin vs operator (VERIFIED 2026-06-23)
|
||
|
||
The OS has TWO roles and they must be different identities ([[threat-model]]: the operator is the
|
||
adversary). Create a dedicated **admin** (real password, sudo, NO auto-login) and keep the
|
||
**operator** as an auto-login, UNPRIVILEGED account.
|
||
|
||
```bash
|
||
sudo adduser admin && sudo usermod -aG sudo admin
|
||
# VERIFY in a second session: log in as admin → `sudo whoami` prints root — BEFORE the next step:
|
||
sudo gpasswd -d <operator> sudo # demote the auto-login operator
|
||
groups <operator> # confirm: no 'sudo'
|
||
```
|
||
|
||
> Use **`gpasswd -d`**, not `deluser <user> <group>`: on this Ubuntu the perl adduser tooling
|
||
> rejects hyphenated usernames (`sanitize_string: invalid characters in 'park-operator'` —
|
||
> VERIFIED on park-buzi 2026-07-06). And group removal applies at **next login** — the auto-login
|
||
> operator session keeps its old memberships until the box reboots (or the session relogs);
|
||
> re-verify `groups` from inside the operator session afterwards.
|
||
|
||
⚠ Order matters: confirm the new admin's sudo works **before** demoting the operator, or you lock
|
||
yourself out. Keep auto-login on the OPERATOR, not admin. **Leave root password disabled** (Ubuntu
|
||
default) — `admin`+sudo IS the root path; enabling root adds risk, no gain.
|
||
|
||
> Strip latent escalation groups from the operator: **`sudo gpasswd -d <operator> lxd`** (lxd group
|
||
> = launch a privileged container that mounts host `/` as root — undoes the no-sudo hardening) and
|
||
> `sudo gpasswd -d <operator> lpadmin` (printer admin, unneeded). And NEVER add the operator to
|
||
> `docker` (also root-equivalent).
|
||
|
||
## 5b. Further hardening (TODO — not yet done)
|
||
|
||
- **Key-based SSH only** (disable password auth) if SSH is enabled at all. Routine ops no longer
|
||
need SSH — Komodo Periphery (§7) drives deploys + gives a container terminal over the mesh — so
|
||
SSH can be locked down hard or disabled, leaving the mesh + Komodo as the management path.
|
||
- **No/locked-down desktop + kiosk autostart** — single-purpose; the operator never reaches a shell
|
||
([[desktop-shell-tauri]]).
|
||
- Consider moving the host **event-signing key into the TPM** (non-extractable) — [[tpm]], [[open-questions]] #12.
|
||
- `sudo apt autoremove` the leftover old kernel once the new one is proven.
|
||
|
||
## 6. Runtime — Docker engine (VERIFIED 2026-06-23)
|
||
|
||
Install Docker Engine + compose (as `admin`). NB Ubuntu 26.04 codename is **`resolute`**, which
|
||
download.docker.com may not yet publish — pin the repo line to `noble`, OR use Ubuntu's `docker.io`.
|
||
Add only `admin` to the `docker` group (root-equivalent — NEVER the operator).
|
||
|
||
This gives the appliance the engine. **How the stack gets ONTO it is step 7** — and as of
|
||
2026-06-27 the primary path is **Komodo (remote, no-SSH)**, not a hand-copied dir. The manual
|
||
`docker compose` flow survives as a **break-glass fallback** (§7c).
|
||
|
||
## 7. Deploy the stack — Komodo Periphery (PRIMARY, 2026-06-27)
|
||
|
||
The booth is driven by a central **Komodo Core** over the **NetBird** mesh. The appliance runs a
|
||
small **Periphery** agent that *dials out* to Core; Core then deploys the same compose files. No
|
||
inbound port on the booth, no SSH for routine ops. Full rationale + threat model:
|
||
[[fleet-deployment-komodo]]. Verified end-to-end on the first booth (`park-buzi`) 2026-06-27.
|
||
|
||
### 7a. Install Periphery (on the booth, as `admin`)
|
||
|
||
Prereq: the booth is on the **NetBird** mesh and can reach Core's reverse-proxy URL
|
||
(`https://komodo.infra.msai.al`).
|
||
|
||
1. In Core: **Settings → Onboarding → + New Onboarding Key** (Name = the booth, e.g. `park-buzi`;
|
||
Expiry ~1 day; Pre-Existing Key empty). Copy the one-time `O-…` key. **Single-use** — delete it
|
||
after the agent connects.
|
||
2. On the booth, install Periphery in **user mode** (runs as `admin`, who is in `docker`; NO root
|
||
daemon; **outbound** → opens no inbound port):
|
||
|
||
```bash
|
||
curl -sSL https://raw.githubusercontent.com/moghtech/komodo/main/scripts/setup-periphery.py | python3 - --user \
|
||
--core-address="https://komodo.infra.msai.al" \
|
||
--connect-as="park-buzi" \
|
||
--onboarding-key="O-…"
|
||
sudo loginctl enable-linger admin # so the user service starts at boot without a login
|
||
```
|
||
|
||
- `--connect-as` is the **Server name in Core** — unique, stable, site-meaningful (the fleet's
|
||
primary key). Booth #2 = a different name (e.g. `park-durres`); never reuse one.
|
||
- `--core-address` is Core's **reverse-proxy URL** (the URL you load the Core UI at over the mesh),
|
||
NOT `:9120` — Core's container port `9120` is exposed-not-published; the agent reaches it through
|
||
the proxy. (Gotcha #7 below.)
|
||
- Config lands at `~/.config/komodo/periphery.config.toml`. The key field is **`core_address`**
|
||
(singular); `root_directory` must be a path `admin` can write. **⚠ VERIFY THIS after install —
|
||
Periphery v2.2.0's installer writes `root_directory = "/etc/komodo"` even with `--user`**
|
||
(bit the lab box 2026-07-07: panic `Failed to write private key pem to "/etc/komodo/keys/
|
||
periphery.key" … Permission denied`, crash-loop until systemd gives up). Fix + restart:
|
||
```bash
|
||
sed -i 's|^root_directory = .*|root_directory = "'"$HOME"'/.komodo"|' ~/.config/komodo/periphery.config.toml
|
||
systemctl --user reset-failed periphery && systemctl --user restart periphery
|
||
```
|
||
NB `sudo systemctl restart periphery` says *unit not found* — it's a USER unit; always
|
||
`systemctl --user …`. The onboarding key survives a pre-connect crash (unused until first dial).
|
||
|
||
Verify: `systemctl --user status periphery` → active; the server **`park-buzi`** appears and goes
|
||
**OK/green** in Core → Servers. Then **delete the onboarding key**.
|
||
|
||
### 7b. Deploy the Stack (in Core — by hand once, then code)
|
||
|
||
Add **Registry Account** + **Git Account** for `git.infra.msai.al` (user `komodo`, tokens) in Core
|
||
so Periphery can clone the repo AND pull the private images. Two distinct credential types — the
|
||
git clone working does NOT imply the image pull is authed (gotcha #8). Per-booth secrets
|
||
(`park_<booth>_jwt_secret`, `park_<booth>_event_signing_key` — distinct values, `openssl rand -hex
|
||
32`) live in Core's **Variables/Secrets** store, referenced from the Stack as `[[…]]`.
|
||
|
||
Create a **Stack** (UI → Stacks → New), name = the booth (`park-buzi`):
|
||
|
||
- **Server:** `park-buzi` · **Source:** repo `mca/parking_solution`, branch `dev`, files
|
||
`docker-compose.yml` + `docker-compose.prod.yml` · **Registry account:** `komodo` (else the pull
|
||
is anonymous → `no basic auth credentials`).
|
||
- **Environment** (Komodo writes this to a `.env` on the booth at deploy, substituting `[[…]]`):
|
||
|
||
```
|
||
REGISTRY=git.infra.msai.al/mca/parking_solution
|
||
TAG=dev # moving tag (staging). PIN to dev-<sha> for a live booth.
|
||
COOKIE_SECURE=0 # CRITICAL on plain-http or the auth cookie never sends → no login
|
||
VISION_ENABLED=1
|
||
WS_ALLOWED_ORIGINS= # browser at the booth URL is same-origin; leave empty (the
|
||
# Tauri desktop app needs its origin here — separate task)
|
||
JWT_SECRET=[[park_buzi_jwt_secret]]
|
||
EVENT_SIGNING_KEY=[[park_buzi_event_signing_key]]
|
||
```
|
||
|
||
Deploy → Periphery pulls + `compose up`s. All containers (`proxy`/Caddy, `server`, `vision`) green.
|
||
Seed the FIRST admin (DB starts empty → nobody can log in until this runs; idempotent) **via
|
||
Komodo's terminal on the `server` container** (no SSH):
|
||
|
||
```bash
|
||
docker exec -it -e ADMIN_USER=admin -e ADMIN_PASS='<strong-pw>' \
|
||
park-buzi-server-1 node scripts/seed-admin.mjs
|
||
```
|
||
|
||
> **Secrets-on-disk note.** The generated `.env` lands on the booth with **cleartext** secrets
|
||
> (compose needs real values). That's why the disk is LUKS-encrypted (§3–4) and keys are per-booth
|
||
> — the encryption is the control, and a single-booth compromise leaks only that booth's key. See
|
||
> [[fleet-deployment-komodo]] (the `EVENT_SIGNING_KEY`-in-Core blast-radius caveat; ATECC608 is the
|
||
> intended long-term signer).
|
||
|
||
### 7b-bis. Fleet-as-code (`resources.toml`) — optional but recommended
|
||
|
||
The repo's `komodo/resources.toml` mirrors the working Stack. Pointing a Core **ResourceSync** at
|
||
it makes the fleet **git-managed**: booth #N is a copy-pasted `[[stack]]` block; an image bump is a
|
||
one-line `TAG=` edit + push + Execute; every change is an auditable commit; a rebuilt Core
|
||
re-creates everything from the file. Keep the sync **Unmanaged** + **Delete-Unmatched OFF** until
|
||
trusted. An **empty diff / disabled Execute = the file already matches the live Stack** (success,
|
||
not an error). See `komodo/README.md` and [[fleet-deployment-komodo]].
|
||
|
||
### 7c. Break-glass — manual compose (mesh/Core down)
|
||
|
||
When the mesh or Core is unreachable, the same compose files run locally via `scripts/booth.sh`
|
||
(or raw `docker compose`). Needs a local `.env` and a `docker login git.infra.msai.al` (a
|
||
read-only package token). This is the FALLBACK, not the routine path:
|
||
|
||
```bash
|
||
docker login git.infra.msai.al
|
||
ENV=prod ./booth.sh config # dry-run the merged env
|
||
ENV=prod ./booth.sh up
|
||
```
|
||
|
||
`booth.sh` runs from wherever it sits next to the compose files (the booth deploys them flat, e.g.
|
||
`/opt/parking_systems/`). See [[container-deployment]].
|
||
|
||
### 7d. Reset the DB for TRAINING/DEMO — `docker exec`, not `pnpm` (2026-06-30)
|
||
|
||
A site is sometimes run live to **train** operators/admins on the real app; afterwards the demo data
|
||
must go without leaving an obvious self-serve button (the [[threat-model|operator must not be able to
|
||
wipe history]]). The reset is a **CLI script** (`packages/db/scripts/reset-db.mjs`), and on the booth
|
||
there is **no `pnpm`** — only the running containers. So run it the same way as the seed-admin step in
|
||
§7b: **`docker exec` into the `server` container**, where the script ships inside the deploy bundle at
|
||
`node_modules/@parking/db/scripts/reset-db.mjs` (the same place the boot migrator lives — see the
|
||
entrypoint). `DATABASE_URL` in-container is **`/data/parking.sqlite`** (the `parking-data` volume).
|
||
|
||
```bash
|
||
# On the booth (or via Komodo's terminal on the server container). Category flags:
|
||
# --financial ledger (entry/exit/payment/void/shift/cash/anomaly) + device_events + snapshots +
|
||
# subscription INSTANCES/credentials/plates + blocklist. KEEPS users/devices/config/
|
||
# tariffs/subscription PLANS.
|
||
# --config site_config, devices, setup_state (re-runs first-run setup), tariffs + versions, plans.
|
||
# --users users, roles, role_permissions, auth sessions. --all every table.
|
||
docker exec -it \
|
||
-e RESET_ALLOWED=1 \
|
||
-e DATABASE_URL=/data/parking.sqlite \
|
||
park-buzi-server-1 \
|
||
node node_modules/@parking/db/scripts/reset-db.mjs --financial
|
||
```
|
||
|
||
> **⚠ `--financial`/`--all` TRUNCATE the append-only, signed [[append-only-event-chain|ledger]]** —
|
||
> the anti-fraud record. A *partial* delete would break the hash chain, so a financial reset wipes the
|
||
> whole ledger back to empty (re-seeding starts a NEW chain under the **same** `EVENT_SIGNING_KEY`/
|
||
> `BACKUP_KEY` — the keys are **not** touched). This is the opposite of how the ledger is meant to
|
||
> behave, hence the two gates: it refuses unless **`RESET_ALLOWED=1`** is set (a real booth never sets
|
||
> it) **and** you type the DB filename to confirm (`parking.sqlite`; `--yes` skips that for scripted
|
||
> setup only). It is a **training/demo** tool — never run on a production booth's data.
|
||
|
||
After `--users`/`--all` (users cleared), re-seed the first admin exactly as in §7b
|
||
(`docker exec … node scripts/seed-admin.mjs`) so someone can log back in. Since 2026-07-06 the seed
|
||
script **self-heals the built-in `admin` role row** that this reset also wipes — before that fix the
|
||
documented re-seed died on a `role_id` FOREIGN KEY error (field failure on `park-buzi`). For dev
|
||
(where `pnpm` exists) the same script is `pnpm db:reset --financial` — see [[local-dev-workflow]].
|
||
|
||
### 7e. Lost APP admin password — reset from the Linux admin account (2026-07-06)
|
||
|
||
The app's admin password lives only as a bcrypt hash in the booth DB; there is no in-app recovery
|
||
(nobody above the admin exists to send a reset). The recovery path is the **Linux `admin` account**
|
||
(the only user in the `docker` group): the seed script doubles as the password-reset tool via
|
||
`FORCE=1` — on an existing username it RESETS that user's password (and restores `roleId: admin`,
|
||
so it also rescues a demoted admin).
|
||
|
||
```bash
|
||
# Interactive (preferred — the password never lands in shell history):
|
||
docker exec -it -e FORCE=1 park-buzi-server-1 node scripts/seed-admin.mjs
|
||
# → prompts: username (Enter = admin), new password (min 8 chars)
|
||
|
||
# Non-interactive (scripted; NB the password enters the HOST's shell history):
|
||
docker exec -e FORCE=1 -e ADMIN_USER=admin -e ADMIN_PASS='new-strong-pass' \
|
||
park-buzi-server-1 node scripts/seed-admin.mjs
|
||
```
|
||
|
||
- **Attributable, not gated.** Whoever holds Linux root owns the DB file — the app cannot defend
|
||
against that actor and doesn't pretend to. What it CAN do: the script appends a **signed
|
||
`config_change` ledger event** (`admin.passwordReset` / `admin.seeded` on first seed, operator
|
||
`console:seed-admin`) so a console reset stays visible in the chain afterwards. If the signing key
|
||
is unavailable (e.g. a dev shell), it warns loudly and proceeds — locking an admin out to protect
|
||
an audit line would invert the priority. The [[threat-model]] adversary remains the *operator*,
|
||
who has no Linux account at all.
|
||
- **Sessions are NOT revoked** by a password reset — issued JWT cookies ride to expiry. A *forgotten*
|
||
password needs nothing more; a *suspected-stolen* one should also rotate the booth's `JWT_SECRET`
|
||
(Komodo Variables → redeploy), which invalidates every session instantly.
|
||
- Works on a fresh/reset DB too (the role-row self-heal above), so §7b first-seed, §7d post-reset
|
||
re-seed, and this recovery are all the same one command.
|
||
|
||
### Healthy startup + web-access
|
||
|
||
Healthy logs: vision `Initialized LicensePlateDetector …` with NO "Downloading" (baked weights),
|
||
server `[migrate] done` → `SPA static serving enabled` → `Server listening`. The transient
|
||
`vision-service -> offline` at boot then `-> ready (fast_alpr)` ~8s later is normal (monitor polls
|
||
before vision finishes loading). Reach the UI at **`http://<name-or-ip>/`** (Caddy on :80).
|
||
|
||
**Web-access gotchas (all fixed in the images/compose — see [[container-deployment]] "Web access"):**
|
||
the SPA uses a RELATIVE `/api` base (works from any host; do NOT bake a domain) + a Caddy proxy gives
|
||
the clean port-80 URL; the domain (`parksystems.msai.al`) is pointed at the booth's LAN IP via
|
||
`hosts`/DNS ON-SITE, never an image rebuild. The **Tauri desktop app** is hardcoded to
|
||
`localhost:3000` (CSP + endpoints) and can't reach a remote booth without code changes — a browser
|
||
works; the desktop app is a separate workstream.
|
||
|
||
## Quick-reference: the gotchas, in order they bit us
|
||
|
||
1. Ventoy USB → `0x1A` Security Violation under Secure Boot → flash the ISO directly instead.
|
||
2. 7070 BIOS has no "View Key" → can't inspect db; the live-USB boot IS the verification.
|
||
3. Installer "hardware-backed encryption" → `PCR_UNUSABLE`/dbt → use passphrase LUKS + manual seal.
|
||
4. Bind TPM to **PCR 7 only**, not a multi-PCR set (kernel updates churn 4/8/9 → passphrase every boot).
|
||
5. Always keep the **password slot** + an off-machine copy of the passphrase (TPM is never the only key).
|
||
6. GRUB password MUST be **edit-only** (`--unrestricted` on entries) or it prompts on EVERY boot →
|
||
breaks unattended reboot. Verify `grep -c unrestricted /boot/grub/grub.cfg` ≥1 before rebooting.
|
||
|
||
### Firmware / dbx gotchas (2026-06-30, §4a)
|
||
|
||
12. **UEFI dbx ships via `fwupd`/LVFS, NOT APT.** `apt list --upgradable` clean ≠ no firmware update
|
||
pending. A new dbx vs a stale GRUB → revoked bootloader → **unbootable / GRUB panic** (`apt
|
||
full-upgrade` first, then dbx). And dbx **moves PCR 7** → breaks TPM auto-unlock → passphrase
|
||
prompt → re-seal (§4 runbook). Mask `fwupd` + remove the `firmware-updater` snap so the operator
|
||
can't trigger it.
|
||
13. `cryptsetup … --test-passphrase` **silently passes via the TPM token** (false safety signal). Use
|
||
`--disable-external-tokens` to actually force a typed-passphrase test before any firmware change.
|
||
|
||
### Komodo deploy gotchas (2026-06-27)
|
||
|
||
7. Periphery `core_address` is **Core's reverse-proxy URL** (`https://komodo.infra.msai.al`), NOT
|
||
`100.x:9120`. Core's `9120` is exposed-not-published (`docker ps` shows `9120/tcp` with no `->`)
|
||
→ a direct dial gets `Connection refused`. Ping/SSH working over the mesh does NOT mean `:9120`
|
||
is reachable.
|
||
8. **Git auth ≠ registry auth.** The repo cloning fine does not mean image pull is authed — they're
|
||
separate Komodo credentials. A blank registry account on the Stack → anonymous pull →
|
||
`no basic auth credentials`. Set the Stack's **Registry Account** (`komodo`).
|
||
9. **User-mode Periphery + `/etc/komodo` `root_directory` = `Permission denied`** writing the agent
|
||
key. User-mode (runs as `admin`, no root daemon) must keep `root_directory` under `$HOME`.
|
||
10. The config key is **`core_address`** (singular). And `--core-address` derives `wss://` from
|
||
`https://` — if Core were plain-HTTP you'd need `http://` (→ `ws://`).
|
||
11. ResourceSync **Execute disabled + file shown clean in Info = empty diff = already in sync**
|
||
(success). Execute only enables when the file and Core diverge (e.g. you edit `TAG`).
|