Files
julian ea8fe22969
Build desktop / desktop (push) Successful in 5m14s
Build & push images / images (push) Successful in 3m1s
CI / check (push) Successful in 43s
docs(wiki): USB printer cover-open field bug writeup; add art-docker-station lab box
Printer investigation (park-buzi): cover-open on the USB thermal
printer wedges its status offline/faulty, surviving a full reboot,
recoverable only via `docker restart server`. Traced sendRawUsb/
probeUsb end-to-end — no persistent handle in the app layer, so the
leading theory is the container's /dev/usb directory bind-mount
retaining a stale view across the printer's physical re-enumeration.
Not yet confirmed on hardware; documented with repro/confirmation
commands and ranked candidate fixes.

Also registers a new lab bench box, "art-docker-station", as a Komodo
Stack (dev tier, same shape as park-lab, its own isolated secret refs).

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
2026-08-30 18:11:34 +02:00

14 KiB
Raw Permalink Blame History

type, tags, sources, updated, status
type tags sources updated status
concept
parking
device
printer
transport
usb
escpos
provisioning
2026-08-30 settled

Printer USB transport (kernel usblp, behind the ESC/POS layer)

The ESC/POS printer drivers (rongta-printer, cashino) can deliver their byte stream over either a raw TCP socket (port 9100) or a local USB character device (/dev/usb/lp0), selected per device by config.transport ("tcp-ip" | "usb"). The original architecture always intended one ESC/POS adapter to cover "USB or network" (parking-system-architecture §BOM); the first implementation shipped TCP-only, and this closes that gap.

The seam — render once, dispatch the transport

Every render*() function in packages/devices/src/drivers/printer-escpos.ts produces a transport-independent ESC/POS Buffer. Only delivery differs. The transport is resolved once per driver from config and every print/probe call site stays transport-blind:

  • transportFromConfig(config) → a discriminated Transport ({ kind: "tcp", host, port } or { kind: "usb", devicePath }). Anything other than transport: "usb" is TCP, so existing host-only configs keep working unchanged (no migration).
  • sendTo(t, payload, timeoutMs) / probeTo(t, timeoutMs) dispatch to the TCP pair (sendRaw/probe) or the USB pair (sendRawUsb/probeUsb).

Adding a transport = one more arm in the dispatcher; not a single rendered byte changes. This is why the CP852 map, the Code128/QR builders, roles/failover, and the receipt/ticket/voucher layouts are all untouched by USB support.

USB transport = the in-box usblp char device

A USB ESC/POS printer plugged into the appliance enumerates as a character device (e.g. /dev/usb/lp0) via the kernel's in-box usblp driver. We just open it O_WRONLY and write the same bytes:

  • No native dependency. A plain fs write — no libusb, no CUPS, no native addon. This keeps the MIT/Apache/BSD-only dependency constraint and the offline-first, minimal-deps appliance posture (see technology-stack, offline-first).
  • usblp is raw. Unlike the TCP path there is no FIN/half-close dance (the graceful-close fix was a TCP concern — an early destroy() could RST-truncate the stream; see rongta-printer). A single open + write delivers the job; we always close the handle.
  • Bounded by a timeout. A wedged USB printer can block the write (or the open) indefinitely; a stuck print must surface as a failure, not hang the entry flow. withTimeout rejects after timeoutMs.

Status over USB — reachability only (honesty rule)

probeUsb is "does the char device exist and open writable" — the USB analogue of the TCP connect probe. A present, openable /dev/usb/lp0 means usblp bound a powered, enumerated printer.

  • The cashino driver is reachability-only on both transports (it never had a status page).
  • The rongta driver's rich readStatus() scrapes the board's HTTP /prn_stat.htm — a network feature. Over USB there is no such page, so readStatus() degrades to the reachability floor (ready/offline only, never a guessed paper/cover state). A USB Rongta is effectively a Cashino for monitoring. This preserves the standing honesty rule from printer-status-monitoring: never report a paper/cover verdict the transport can't actually sense.

Threat model

The USB path is a local character device the booth operator (the threat model's adversary) cannot reach over the network — narrower attack surface than the unauthenticated TCP print socket on the VLAN. Printers are advisory output; nothing about the signed append-only-event-chain or barrier control is touched.

Provisioning dependency (NOT app code) — see open-questions #14

Driving a USB printer depends on the appliance image:

  1. the usblp kernel module is loaded (it is in-box on Ubuntu 26.04; CUPS can claim the interface first — may need usblp to win, or CUPS masked for that device), and
  2. a udev rule grants the server process write access to the node (e.g. a group on /dev/usb/lp*), since the appliance server does not run as root.

This is a appliance-provisioning concern, recorded as open-questions #14 until the on-site printer is confirmed USB and the rule is baked into the image and verified on hardware.

Status

Built 2026-06-24 behind the existing render layer. sendRawUsb/probeUsb/transportFromConfig/ sendTo/probeTo in printer-escpos.ts; cashino + rongta resolve a Transport and dispatch. The setup UI offers a Connection select (Network / USB) + a USB device path field (default /dev/usb/lp0); host/port are not-required so a USB printer needs neither. Covered by printer-escpos.test.ts (USB writes the exact rendered bytes; probe present/absent; transportFromConfig TCP back-compat) and printer-cashino.test.ts (a USB-configured driver prints to the node and reports ready/offline). The on-hardware confirmation + the udev/usblp provisioning are pending (open-questions #14).

Field bug — the NONBLOCK partial-write truncation (found + fixed 2026-07-06)

First on-hardware USB test (ICS XP-K200L, an ESC/POS clone): over TCP it printed + cut fine; over USB it printed the ticket's TEXT but no barcode and no cut. Root cause was in OUR transport, not the printer: sendRawUsb opened the node with O_NONBLOCK and issued ONE write() for the whole job. On a non-blocking usblp fd the kernel accepts only what fits the printer's USB buffer (~8 KB) and returns a short write; the old code never checked bytesWritten, closed the handle, and silently dropped the tail — which is exactly where the barcode (mid-payload) and the CUT (last bytes) live. Small jobs fit one buffer, hence "text prints fine". The regular-file test stand-in can't short-write, so tests never caught it.

Fix: writeAllUsb — chunked loop (4 KB, safely under the usblp buffer) that continues after partial writes, retries EAGAIN/zero-byte writes with a short pause, and fails at the caller's deadline with a (N/M bytes accepted) diagnostic. Driven by fake-handle tests (short writes, EAGAIN interleave, wedged-printer timeout, non-EAGAIN passthrough) since a real file can't reproduce the char device's behaviour.

Second truncation mode — close() cancels the in-flight transfer (lab bench, 2026-07-07). The chunked loop alone STILL truncated on hardware (test slip stopped mid-sentence, no feed, no cut — "press the feed button to see the text"). Verified against drivers/usb/class/usblp.c: write() returns at URB submission (not completion), only ONE write URB is in flight at a time (the next write EAGAINs until it completes), and usblp_release() — i.e. our close() — kills in-flight URBs. The printer drains bulk data at PRINT speed (tiny internal buffer on these clones), so closing right after the last accepted write cancels the still-transferring tail — exactly where the feed + GS V cut bytes live. Kernel-accepted ≠ printer-received.

Fix: the one-URB rule makes acceptance of write N a completion certificate for write N−1. So writeAllUsb now writes the payload's FINAL BYTE alone: when that 1-byte write is accepted, every byte before it is physically in the printer; a short drain pause (USB_DRAIN_MS 300 ms) covers the lone final-byte packet, then close is safe. (usblp also implements poll(POLLOUT) as the true completion signal, but Node cannot poll an arbitrary char-device fd without a native dep — the hold-back + drain gets the same guarantee for all but the final byte, whose packet the printer ACKs immediately after having just freed its buffer.)

✅ HARDWARE-VERIFIED (lab bench, 2026-07-07): with both fixes, the ICS XP-K200L over USB prints the complete slip, feeds, and CUTS — parity with TCP. The USB transport is done.

Device discovery (2026-07-07). The kernel numbers usblp nodes by plug/boot order — park-buzi's printer is lp1, and the admin had to shell in and ls /dev/usb to learn that. The wizard now lists REAL printers: GET /api/setup/usb-printers enumerates /dev/usb/lpN (visible via the compose bind-mount) and enriches each with the printer's self-reported make/model from sysfs (/sys/class/usbmisc/lpN/device/ieee1284_id — readable through Docker's default ro /sys). The devicePath field becomes a SELECT ("/dev/usb/lp1 — Xprinter XP-K200L") with a fresh form preselecting the first present device; a saved-but-unplugged path stays selectable, flagged "saved — not present now"; zero devices found falls back to the free-text path + a check-the-cable hint. The transport option label no longer hardcodes lp0.

Driver-choice note for this clone: the ICS XP-K200L does NOT serve the Rongta /prn_stat.htm status page (checked on hardware at 10.0.10.11 — print socket 9100 open, status page absent), so on NETWORK the honest driver is cashino (reachability-only monitoring); under rongta the monitor would mark a perfectly working printer offline/degraded. Over USB the two drivers behave identically (reachability floor), so either works post-fix. See rongta-printer.

Field bug — cover-open re-enumeration wedges the container's /dev/usb view; only docker restart, not a host reboot, clears it (investigated 2026-08-30, unconfirmed root cause)

Symptom (park-buzi, unknown/"Generic" USB printer, model not yet identified — see below): every time the booth operator opens the printer's paper-roll cover to reload paper, the printer's status goes offline/faulty in the app and never self-recovers — not after the cover closes, not after a full appliance reboot. The only fix found so far is SSH in and docker restart server.

Ruled out at the application layer. Traced sendRawUsb/probeUsb in printer-escpos.ts: every print AND every poll tick (device-monitor.ts 8s / printer-monitor.ts 5s) does a fresh open() → write/probe → close() against the configured devicePath. No fd, socket, or driver instance is held across calls — driver.create(config) is a throwaway object with no persistent handle. So a naive "stale Node file descriptor" explanation does not fit this codebase; the app-layer retry-by-fresh-open-every-poll should self-heal within one poll cycle if the kernel's view of the device node is current.

Leading hypothesis: the container's bind-mount of /dev/usb, not the Node process, holds the stale state. Docker Compose wires the printer in as a directory bind-mount (docker-compose.prod.yml, volumes: - /dev/usb:/dev/usb), chosen deliberately (per its own comment) so the app survives the printer renumbering to a different lpN. But many USB thermal printers cut power to their own USB interface board when the cover-open microswitch trips (a hardware safety/power feature, not just a status flag) — the printer drops off the bus and re-enumerates, potentially as a new device node, when the cover closes. The host kernel picks this up fine; the container's mount namespace, once established, is a known Docker/OverlayFS sharp edge for /dev subtree bind-mounts — it can keep resolving the old node until the mount itself is redone.

  • docker restart server recreates the container's mount namespace → the /dev/usb bind-mount is redone against current host state → the new node is picked up → fixed.
  • A full host reboot restarts the container too (restart: always), but as a boot-time race: if the container starts before the USB subsystem finishes settling, or the printer re-enumerated some time before the reboot and Docker doesn't necessarily redo an already-satisfied bind-mount target on a policy-driven restart, the container can come back up still bound to the pre-incident view. This matches the exact reported asymmetry (reboot doesn't fix it; explicit restart does).

Not yet confirmed on hardware — this is the leading theory, not a verified root cause. To confirm at the next occurrence, BEFORE restarting anything:

# host:
ls -la /dev/usb/ && stat /dev/usb/lp1
# container:
docker exec server ls -la /dev/usb/ && docker exec server stat /dev/usb/lp1

A major:minor or inode mismatch between host and container is the smoking gun. Also worth capturing on the lab RONGTA (different printer, but same cover-open mechanism is plausible): watch -n1 lsusb + sudo dmesg -w | grep -i -E 'usb|disconnect' while cycling the cover, to see whether the Bus/Device number changes.

Candidate fixes, not yet implemented (ranked cheapest-to-most-invasive):

  1. A host-side watchdog/udev rule that detects re-enumeration of this printer (match vendor:product ID) and runs docker restart server automatically — turns the manual SSH fix into a self-healing one without touching app code.
  2. Same idea but event-driven via a udev rule or systemd path unit watching /dev/usb, rather than polling.
  3. Switch the compose device wiring from the directory bind-mount to a specific --device= cgroup passthrough + a udev rule pinning a stable symlink name — reintroduces the renumbering fragility the directory bind-mount was chosen to avoid, so only worth doing alongside (1)/(2), not instead.

Open sub-question — printer identity. The park-buzi unit shows as "Generic (unknown)" in the app; not yet identified by vendor/product ID. Lab reproduction uses a RONGTA unit instead (not the same hardware), so the lab cannot currently reproduce the park-buzi symptom directly — only validate the general re-enumeration mechanism. Commands to identify the real park-buzi printer next time it's reachable via SSH: lsusb, udevadm info -q property -n /dev/usb/lp1, udevadm info -a -n /dev/usb/lp1. This mirrors the same discovery gap already noted above under "Device discovery" (sysfs ieee1284_id enrichment) — once identified, fold the model into that mechanism's coverage.

Related: rongta-printer, printer-status-monitoring, printer-roles-failover, appliance-provisioning, network-isolation, technology-stack.