fix(collector,trainer): migrate an existing collector DB on open; trainer handlers answer 500 JSON

The reviewer host's collector.sqlite was created by an earlier build, before the
`kind` column. CREATE TABLE IF NOT EXISTS shapes only a new database, so every query
naming the column failed: the collector's /health (container unhealthy), every
booth ingest, and the trainer's readiness — whose stdlib server printed the
traceback and dropped the socket, which the collector could only render as
"trainer not reachable: fetch failed". Nine days like that.

- CollectorDb.#migrate(): PRAGMA table_info against the list of columns added
  since the first deploy; ALTER TABLE ADD COLUMN for each missing one (all
  nullable or defaulted). Append to that list whenever a column joins the CREATE.
  Test replays the original schema: health, ingest, stats, a legacy row reads
  back with the defaults.
- Trainer Handler._guarded(): any unexpected exception → 500 JSON naming it,
  never a dropped connection; /health keeps answering. Test drives readiness
  against an old-schema DB.
- The collector's training status proxy includes the trainer's error text.

Wiki: the incident and the schema rule (vision-review-outbox), what the message
means (bodytype-classifier-training), log. Deploy: the new collector migrates on
start; nothing manual.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
This commit is contained in:
2026-09-16 10:33:47 +02:00
parent fe3b12a60d
commit f4b806a538
8 changed files with 186 additions and 2 deletions
+20
View File
@@ -154,6 +154,26 @@ its own repo the day it needs its own cadence. Deploy the collector BEFORE a boo
package kind it does not know (a 422 is abandoned, not retried). The export neutralises cells
that start like a spreadsheet formula (category/service names are booth-supplied text).
> **Incident 2026-09-16 — the collector's DB predated the `kind` column; nothing worked for 9
> days and nothing said so.** The reviewer opened /review: *Training — trainer not reachable:
> fetch failed*. On the host: collector `stage-2d9bb15` **unhealthy** (`/health` → 500 *no such
> column: kind*), trainer healthy but every `/readiness` a Python traceback; the volume's
> `collector.sqlite` (created 2026-09-07 by the previous build, **0 items**) had the original
> column set. `CREATE TABLE IF NOT EXISTS` shapes only a NEW database — an existing volume keeps
> its old columns, so every query naming `kind` failed: the collector's health, **every ingest**
> (booths would have got 500s and kept retrying — the log shows none ever arrived, a separate
> question), and the trainer's readiness. The trainer's stdlib server printed the traceback and
> dropped the socket, which the collector could only render as "fetch failed".
>
> Fixes (same day): `CollectorDb` now **migrates on open** — `PRAGMA table_info` vs a list of the
> columns added since the first deploy, `ALTER TABLE … ADD COLUMN` for each missing one (all
> nullable or defaulted; **append to that list whenever a column joins the CREATE**); the trainer's
> handlers are guarded — an unexpected exception is a **500 JSON** naming the error, never a
> dropped connection; the collector's status proxy surfaces the trainer's error text. Rule going
> forward: the collector owns the schema; the trainer only reads; a deploy that changes the table
> must be accompanied by a migration entry, and the Training section is the first place a
> schema/DB mismatch shows — read its error text before suspecting the network.
**Status (2026-09-07).** Live: the collector runs on `art-docker-station` and park-2 is wired to
it (`stage-dbbb051` on both stacks, every entry sampled). The review screen at
`http://docker-station.nb.infra:8090/review` is filling; no labels reviewed yet.