Cameras watch a robot workcell. A vision-language model on the device reports who is in the frame and what they are wearing, safety rules turn that into structured events, and an operator reviews them. It measures what it costs to run, and it records what the model gets wrong.
reports/thor/The operator draws restricted zones on a snapshot from each camera and picks which rules run on it. Every rule works on what the model reports for that frame: labels, boxes, and for each person whether a hard hat and a high-visibility vest are worn.
A person without a hard hat or without a high-visibility vest. The blue circle is the ISO 7010 shape for a mandatory action.
A person's box overlaps a zone the operator drew as restricted, such as the robot cell.
The centre of a person's box within 90 pixels of the centre of a robot's box.
A pallet, cart or box the model reports as blocking an emergency path.
Anything else the model itself calls unsafe, passed through for review.
One thread per camera decodes the stream and keeps the newest frame. The console shows that video at the camera's own rate.
A worker samples the newest frame, asks the model for people, objects and PPE, and runs the rules bound to that camera.
Events go to a bounded queue drained by its own thread, with retries, a drop counter, and latency measured from frame capture to the backend's acknowledgement.
If the model server stops, the video keeps playing and the console shows "inference unavailable". Events only come from frames the model actually answered.
Read from rules/policies.py and rules/severity.py. The rule functions are pure: no I/O, the same detections always give the same findings.
Hard hat and vest only. Gloves, glasses and boots are not checked.
A degraded runtime lowers confidence by 0.08, a critical one by 0.18, and each dropped frame by 0.01 up to 0.10.
Every number below is read from a committed run artifact. The model runs used 60 frames of an Isaac Sim workcell recording with a robot arm and no people in it.
Hover a dot for its value. A sixth run is left off on purpose: a grammar with a mismatched prompt made the model return nothing on every frame, so its 0.31 s p50 is the cost of generating nothing. The README keeps it as a labelled failure case.
With a JSON-schema grammar, out-of-vocabulary labels went from 59 to 0 on the 2B. On footage with no people, it then reported six persons and fired seven false PPE and proximity events.
Under the same grammar and prompt, the 8B reported zero persons and zero events on the same frames, at 3.20 s p50. Sixty frames of one video is a signal, not a rate.
With vLLM 0.14's reasoning parser on, a JSON schema is accepted but not enforced. That is now blocked in code: before any worker uses constrained mode, a canary request asks for one word while the schema allows only another, and the server has to answer with the schema's value. The probe is committed as reports/thor/constrained_guard_probe.json.
Constrained mode refused when starting the reasoning-parser container.
Constrained mode refused when saving settings with the reasoning parser on.
The reasoning-parser server's answer to the canary. The guard refused.
The container without the parser honoured the grammar and was allowed. Ready 55.6 s after launch.
Screens from the running app, captured with its built-in synthetic workcell feed and the mock detector. The console labels both on screen so they are never mistaken for a real camera or a real model.
Each run writes one artifact with the device, date, inputs, per-frame timings, token use, event counts, peak memory and a 1 Hz tegrastats record.
A timed-out run, an empty-output run and a model echoing its own schema are all committed and labelled, not edited out.
Ten vendor profiles, a connection test that decodes one frame or names the problem, and zones drawn on a live snapshot.
Passwords are encrypted at rest, never returned by the API, only reused for the host they were saved for, and masked in logs, events and reports.
The model container keeps running across restarts. Stopping it needs an explicit, logged confirmation.
The system flags and records. It never stops a robot or a line; a person decides.
Out of scope by design: stopping machines or acting without a person. This is observability with a human in the loop.