f7a262ac9a
Build & push images / images (push) Successful in 6m31s
apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash collector's SQLite + crops read-only off its volume; time split (validation = newest slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against the torch model; NO model file below the validation floor (exit 3, report still written); exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package. Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip. apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract: crop margin, input size, RGB 0-255, normalisation inside the graph) and RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained on, min-confidence, `detector_class` on the result; path set but no file = phase B off without an error; a broken file is a health detail. models/bodytype.version (tracked, empty) pins the published version the Dockerfile fetches at build (BuildKit secret; a pin that cannot be fetched fails the build). Verified: a trainer model gives identical probabilities inside the vision service; both images built and smoke-tested. Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI. Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site, secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox, vision-service-packaging, fleet-deployment-komodo, index, log. Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
85 lines
4.9 KiB
Docker
85 lines
4.9 KiB
Docker
# syntax=docker/dockerfile:1.7
|
|
# Parking VISION image: the Python/uv ANPR microservice. Build CONTEXT is apps/vision
|
|
# (self-contained Python package; no monorepo deps). Ships WITH the `alpr` extra (real
|
|
# fast-alpr/onnxruntime stack) but the engine is env-selected: VISION_RECOGNIZER=stub
|
|
# (default, boots anywhere) or fast_alpr (prod). See wiki/decisions/container-deployment.md,
|
|
# wiki/decisions/vision-service-packaging.md.
|
|
|
|
# uv-provided Python 3.12 (matches apps/vision/.python-version).
|
|
FROM ghcr.io/astral-sh/uv:python3.12-bookworm-slim AS base
|
|
WORKDIR /app
|
|
ENV UV_LINK_MODE=copy \
|
|
UV_COMPILE_BYTECODE=1 \
|
|
PYTHONUNBUFFERED=1
|
|
|
|
# System libs the recognizer stack needs (opencv/onnxruntime): GL + glib. Kept minimal.
|
|
RUN apt-get update \
|
|
&& apt-get install -y --no-install-recommends libgl1 libglib2.0-0 curl \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
# ---- deps: resolve + install the venv from the lockfile (cache-friendly) ----
|
|
# Manifests first so the heavy `uv sync` layer caches across source edits.
|
|
COPY pyproject.toml uv.lock .python-version ./
|
|
RUN --mount=type=cache,target=/root/.cache/uv \
|
|
uv sync --frozen --no-install-project --extra alpr
|
|
|
|
# ---- project source ----
|
|
COPY vision_service/ ./vision_service/
|
|
COPY README.md ./
|
|
# Vehicle stage weights (phase A): YOLOX-S, Apache-2.0, ~36 MB, baked into the image so the
|
|
# air-gapped appliance never fetches at runtime and no operator-writable path holds a model
|
|
# (vision-service-hardening.md). Best-effort at build: without network the stage stays off.
|
|
ARG YOLOX_URL=https://github.com/Megvii-BaseDetection/YOLOX/releases/download/0.1.1rc0/yolox_s.onnx
|
|
RUN mkdir -p /app/models \
|
|
&& (curl -fsSL -o /app/models/yolox_s.onnx "$YOLOX_URL" \
|
|
|| (echo "[build] yolox weights not fetched (no network) — vehicle stage off" && rm -f /app/models/yolox_s.onnx))
|
|
# Phase B body-type classifier (apps/trainer output, published to the Gitea generic package
|
|
# registry — weights are not code, they never live in git). models/bodytype.version PINS the
|
|
# version this image carries: empty = no classifier (phase B off). A pinned version that
|
|
# cannot be fetched FAILS the build — the image must carry what git says it carries. The
|
|
# registry may need auth: pass a BuildKit secret `bodytype_auth` holding "user:token".
|
|
ARG BODYTYPE_BASE_URL=https://git.infra.msai.al/api/packages/mca/generic/parking-bodytype
|
|
COPY models/bodytype.version ./models/bodytype.version
|
|
RUN --mount=type=secret,id=bodytype_auth \
|
|
v="$(tr -d '[:space:]' < /app/models/bodytype.version)"; \
|
|
if [ -n "$v" ]; then \
|
|
cfg=/tmp/curl.cfg; : > "$cfg"; \
|
|
[ -f /run/secrets/bodytype_auth ] && printf 'user = "%s"\n' "$(cat /run/secrets/bodytype_auth)" > "$cfg"; \
|
|
curl -fsSL -K "$cfg" -o /app/models/bodytype.onnx "$BODYTYPE_BASE_URL/$v/bodytype.onnx" \
|
|
&& curl -fsSL -K "$cfg" -o /app/models/bodytype.json "$BODYTYPE_BASE_URL/$v/bodytype.json" \
|
|
&& echo "[build] bodytype classifier $v baked" \
|
|
|| { echo "[build] bodytype classifier $v could not be fetched"; rm -f "$cfg"; exit 1; }; \
|
|
rm -f "$cfg"; \
|
|
else echo "[build] no bodytype version pinned — phase B off"; fi
|
|
RUN --mount=type=cache,target=/root/.cache/uv \
|
|
uv sync --frozen --extra alpr
|
|
|
|
# Non-root runtime user, created BEFORE the model pre-warm so the weights cache lands in
|
|
# this user's HOME (~/.cache) — the SAME path the runtime reads. (fast-alpr's
|
|
# open-image-models caches under $HOME/.cache/open-image-models keyed to HOME, ignoring
|
|
# HF_HOME/XDG_CACHE_HOME — so the pre-warm MUST run as the runtime user, not root.)
|
|
RUN useradd --system --create-home --uid 999 vision \
|
|
&& chown -R vision:vision /app
|
|
USER vision
|
|
|
|
# Pre-warm the fast-alpr model weights INTO the image (as the vision user → /home/vision/
|
|
# .cache) so the prod recognizer is OFFLINE-first: ALPR() downloads weights on first
|
|
# construction, which would otherwise need network on the appliance's first scan. Best-effort
|
|
# — if the build host has no network this is skipped and weights fetch lazily at runtime.
|
|
# NB: NO --mount=type=cache here — a BuildKit cache mount at ~/.cache is NOT committed to the
|
|
# image layer, so the downloaded weights would vanish. They must write to the real layer.
|
|
RUN uv run python -c "from fast_alpr import ALPR; ALPR()" \
|
|
|| echo "[build] model pre-warm skipped (no network) — weights fetch at runtime"
|
|
|
|
# Default to the stub recognizer (offline, no model load); override to fast_alpr in prod.
|
|
ENV VISION_RECOGNIZER=stub \
|
|
VISION_HOST=0.0.0.0 \
|
|
VISION_PORT=8089 \
|
|
VISION_VEHICLE_MODEL_PATH=/app/models/yolox_s.onnx \
|
|
VISION_VEHICLE_CLASSIFIER_PATH=/app/models/bodytype.onnx
|
|
EXPOSE 8089
|
|
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
|
|
CMD python -c "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://localhost:8089/health').status==200 else 1)" || exit 1
|
|
|
|
CMD ["uv", "run", "uvicorn", "vision_service.app:app", "--host", "0.0.0.0", "--port", "8089"]
|