fix(collector,trainer): migrate an existing collector DB on open; trainer handlers answer 500 JSON

The reviewer host's collector.sqlite was created by an earlier build, before the
`kind` column. CREATE TABLE IF NOT EXISTS shapes only a new database, so every query
naming the column failed: the collector's /health (container unhealthy), every
booth ingest, and the trainer's readiness — whose stdlib server printed the
traceback and dropped the socket, which the collector could only render as
"trainer not reachable: fetch failed". Nine days like that.

- CollectorDb.#migrate(): PRAGMA table_info against the list of columns added
  since the first deploy; ALTER TABLE ADD COLUMN for each missing one (all
  nullable or defaulted). Append to that list whenever a column joins the CREATE.
  Test replays the original schema: health, ingest, stats, a legacy row reads
  back with the defaults.
- Trainer Handler._guarded(): any unexpected exception → 500 JSON naming it,
  never a dropped connection; /health keeps answering. Test drives readiness
  against an old-schema DB.
- The collector's training status proxy includes the trainer's error text.

Wiki: the incident and the schema rule (vision-review-outbox), what the message
means (bodytype-classifier-training), log. Deploy: the new collector migrates on
start; nothing manual.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
This commit is contained in:
2026-09-16 10:33:47 +02:00
parent fe3b12a60d
commit f4b806a538
8 changed files with 186 additions and 2 deletions
+14
View File
@@ -3308,3 +3308,17 @@ Live after the fix: degraded "cover open, paper out, printer off-line"; closed
"we are good using the network with this printer." Also: park-lab Periphery was "not loaded"
after a reboot — the unit had never been enabled; `systemctl --user enable --now periphery`
recorded as a §7a gotcha on [[appliance-provisioning]]. Pages: [[k200l-printer]].
## [2026-09-16] fix | Collector DB schema migration; trainer handlers answer 500 JSON; the "trainer not reachable" incident
User: "Training — trainer not reachable: fetch failed". Read-only look at art-docker-station over
SSH: collector stage-2d9bb15 unhealthy (`/health` 500 "no such column: kind"), trainer healthy
but `/readiness` tracebacks on the same column; the volume's collector.sqlite (2026-09-07, 0
items, original columns) predates `kind` — CREATE TABLE IF NOT EXISTS never migrates an existing
table. No `/ingest` request in the container's 9-day log at all (park-2 either not sending or not
reaching the host — to check on the booth). Built: `CollectorDb.#migrate()` (PRAGMA table_info vs
the list of columns added since the first deploy → ALTER TABLE ADD COLUMN; test replays the old
schema: health, ingest, stats, legacy row reads back with defaults); trainer `Handler._guarded`
(any exception → 500 JSON naming it; test: readiness on an old-schema DB → 500 "no such column:
kind", /health still 200); the collector's training proxy includes the trainer's error text. Pages:
[[vision-review-outbox]] (incident + rule), [[bodytype-classifier-training]] (what the message
means). Deploy: nothing manual — the new collector migrates on start.