feat(trainer): phase-B body-type classifier — trainer job on the collector host + the classifier stage on the booth
Build & push images / images (push) Successful in 6m31s

apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash
collector's SQLite + crops read-only off its volume; time split (validation = newest
slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet
backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light
augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against
the torch model; NO model file below the validation floor (exit 3, report still written);
exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after
training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package.
Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip.

apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract:
crop margin, input size, RGB 0-255, normalisation inside the graph) and
RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained
on, min-confidence, `detector_class` on the result; path set but no file = phase B off
without an error; a broken file is a health detail. models/bodytype.version (tracked,
empty) pins the published version the Dockerfile fetches at build (BuildKit secret;
a pin that cannot be fetched fails the build). Verified: a trainer model gives identical
probabilities inside the vision service; both images built and smoke-tested.

Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the
collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in
the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI.

Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site,
secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox,
vision-service-packaging, fleet-deployment-komodo, index, log.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
This commit is contained in:
2026-09-07 11:14:50 +02:00
parent f9cb973fe9
commit f7a262ac9a
41 changed files with 3797 additions and 76 deletions
+12 -4
View File
@@ -128,10 +128,18 @@ Three surfaces, nothing else — it must not grow into a fleet console:
/ fraud rate.
- **`GET /export/labels.csv`** — reviewed, usable rows: item, booth, crop path, the reviewer's
label, the operator's category + classes, the camera's class + confidence, downgraded, at.
Crops are not packaged: the phase-B trainer runs **on the same host** (its GPU) and reads them
off the volume ([[bodytype-classifier-training]]: CPU-only, the Xeon is enough) —
`docker-compose.collector.yml` carries the `trainer` seam as a commented
`profiles: [train]` one-off job (next increment).
Crops are not packaged: the phase-B trainer runs **on the same host** and reads the SQLite
+ crops straight off the volume, read-only ([[bodytype-classifier-training]]: CPU-only, the
Xeon is enough) — `docker-compose.collector.yml` carries it as the `trainer` service under
`profiles: ["train"]`, a one-off job never started by a deploy (built 2026-09-07; the CSV
export stays for a human with a spreadsheet).
**Where the data lives.** The collector writes to `/data` in its container: `collector.sqlite`
and one JPEG per item at `crops/<booth-id>/<item-id>.jpg`. `/data` is the named Docker volume
`collector-data` (compose), on the host under Docker's volume directory — normally
`/var/lib/docker/volumes/wash-collector_collector-data/_data/` (`docker volume inspect
wash-collector_collector-data` confirms). The trainer mounts the same volume read-only at its
own `/data`; nothing is copied or exported for training.
**Deploy notes.** Bind the published port to the host's **Netbird address** (`COLLECTOR_BIND`),
never `0.0.0.0` on a host with a public interface; Netbird policy: booths → this host:8090 and