feat(trainer): phase-B body-type classifier — trainer job on the collector host + the classifier stage on the booth
Build & push images / images (push) Successful in 6m31s
Build & push images / images (push) Successful in 6m31s
apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash collector's SQLite + crops read-only off its volume; time split (validation = newest slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against the torch model; NO model file below the validation floor (exit 3, report still written); exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package. Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip. apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract: crop margin, input size, RGB 0-255, normalisation inside the graph) and RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained on, min-confidence, `detector_class` on the result; path set but no file = phase B off without an error; a broken file is a health detail. models/bodytype.version (tracked, empty) pins the published version the Dockerfile fetches at build (BuildKit secret; a pin that cannot be fetched fails the build). Verified: a trainer model gives identical probabilities inside the vision service; both images built and smoke-tested. Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI. Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site, secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox, vision-service-packaging, fleet-deployment-komodo, index, log. Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
This commit is contained in:
@@ -92,6 +92,14 @@ The skeleton is **built and wired** (no recognizer models yet):
|
||||
that has `alpr`. Rule (2026-09-07, after three red runs): pure post-processing tests get numpy
|
||||
from the **dev group**; anything needing OpenCV uses `pytest.importorskip("cv2")`; the service
|
||||
itself imports both lazily inside functions.
|
||||
- **The same pattern, second package (2026-09-07):** `apps/trainer` (`@parking/trainer`,
|
||||
[[bodytype-classifier-training]]) — light core (numpy, opencv-headless, onnxruntime) + a
|
||||
`train` extra (CPU-only torch/torchvision/onnx/onnxscript from PyTorch's wheel index via
|
||||
`tool.uv.index`); CI syncs without it, torch tests `importorskip("torch")`, the module that
|
||||
imports torch is imported lazily by the `train` command only. Its own image
|
||||
(`parking-trainer`, context `apps/trainer`, uv base image, bakes `--extra train` + the
|
||||
ImageNet backbone weights) is built by build-images.yml beside the other three; both Python
|
||||
contexts now carry a `.dockerignore` (venv/caches/weights out). Workspace count 7→8.
|
||||
- **Light-core, heavy-optional:** core deps boot in **stub mode** (no model download) so `uv sync` +
|
||||
tests work offline; the real stack is the `alpr` extra (`uv sync --extra alpr` →
|
||||
fast-alpr + onnxruntime). `VISION_RECOGNIZER=fast_alpr` switches it on.
|
||||
|
||||
Reference in New Issue
Block a user