EmbodiedEdge Labs · Evidence
Runtime safety · Robot workcells · Jetson AGX Thor Functional · live-camera validation in progress

Physical AI Safety Observability

Cameras watch a robot workcell. A vision-language model on the device reports who is in the frame and what they are wearing, safety rules turn that into structured events, and an operator reviews them. It measures what it costs to run, and it records what the model gets wrong.

DeviceJetson AGX ThorL4T R38.4.0, 120 W mode
ModelCosmos-Reason2 2B and 8Bserved on device by vLLM 0.14
Rules5 safety rulesPPE, zone, proximity, path, summary
Evidence11 committed artifactsunder reports/thor/
Cell plan

What the rules look for in a workcell

The operator draws restricted zones on a snapshot from each camera and picks which rules run on it. Every rule works on what the model reports for that frame: labels, boxes, and for each person whether a hard hat and a high-visibility vest are worn.

camera robot hat ✗ vest ✓ hat ✓ vest ✓ EXIT pallet 1 PPE MISSING 2 RESTRICTED ZONE ENTRY 3 ROBOT PROXIMITY 4 BLOCKED EMERGENCY PATH Schematic. Not from a run.
The worker in the cell trips three rules at onceProximity is measured in camera pixels
  • PPE missing HIGH

    A person without a hard hat or without a high-visibility vest. The blue circle is the ISO 7010 shape for a mandatory action.

  • Restricted zone entry CRITICAL

    A person's box overlaps a zone the operator drew as restricted, such as the robot cell.

  • Human and robot proximity HIGH

    The centre of a person's box within 90 pixels of the centre of a robot's box.

  • Blocked emergency path MEDIUM

    A pallet, cart or box the model reports as blocking an emergency path.

  • Unsafe event summary LOW

    Anything else the model itself calls unsafe, passed through for review.

Signal path, per camera

Three loops that never wait on each other

CaptureNative frame rate

One thread per camera decodes the stream and keeps the newest frame. The console shows that video at the camera's own rate.

InferenceOn its own interval

A worker samples the newest frame, asks the model for people, objects and PPE, and runs the rules bound to that camera.

PostingAsynchronous

Events go to a bounded queue drained by its own thread, with retries, a drop counter, and latency measured from frame capture to the backend's acknowledgement.

A model outage never invents an event

If the model server stops, the video keeps playing and the console shows "inference unavailable". Events only come from frames the model actually answered.

Logic as shipped

When a person is flagged, and when a human must look

Read from rules/policies.py and rules/severity.py. The rule functions are pure: no I/O, the same detections always give the same findings.

PPE check

for each person: if not hard_hat or not vest: PPE_MISSING, HIGH confidence = highest among those people

Hard hat and vest only. Gloves, glasses and boots are not checked.

Review gate

review required if severity is HIGH or CRITICAL or confidence < 0.85 or runtime is not nominal

A degraded runtime lowers confidence by 0.08, a critical one by 0.18, and each dropped frame by 0.01 up to 0.10.

Measured on Jetson AGX Thor, 2026-09-09

The pipeline keeps camera rate. The model sets the pace.

Every number below is read from a committed run artifact. The model runs used 60 frames of an Isaac Sim workcell recording with a robot arm and no people in it.

28.588 fpsPipeline at 30 fps pacing with the mock model, 1,800 framesmock_30fps_no_post.json
0.9877 msRule evaluation per frame, p95mock_30fps_no_post.json
2.69 sCosmos-Reason2-2B per frame, p50, with the JSON grammarcosmos2b_video_60_schema.json
62.5 WBoard power p50 during that run, against about 24 W idlecosmos2b_video_60_schema_tegrastats.jsonl

Inference latency per frame, five model runs

p50p95
0 s51015202530 s 2B, think block2B, no think2B, JSON grammar8B, JSON grammar8B, no think 2B, think block · p50 4.35 s2B, think block · p95 12.85 s 2B, no think · p50 3.57 s2B, no think · p95 6.58 s 2B, JSON grammar · p50 2.69 s2B, JSON grammar · p95 3.36 s 8B, JSON grammar · p50 3.20 s8B, JSON grammar · p95 8.26 s 8B, no think · p50 8.29 s8B, no think · p95 27.78 s 12.856.583.368.2627.78

Hover a dot for its value. A sixth run is left off on purpose: a grammar with a mismatched prompt made the model return nothing on every frame, so its 0.31 s p50 is the cost of generating nothing. The README keeps it as a labelled failure case.

The grammar fixed the labels and exposed phantom people

With a JSON-schema grammar, out-of-vocabulary labels went from 59 to 0 on the 2B. On footage with no people, it then reported six persons and fired seven false PPE and proximity events.

The 8B did not see anyone who was not there

Under the same grammar and prompt, the 8B reported zero persons and zero events on the same frames, at 3.20 s p50. Sixty frames of one video is a signal, not a rate.

Measured on the Thor, 2026-10-02

The config refuses a schema the server would ignore

With vLLM 0.14's reasoning parser on, a JSON schema is accepted but not enforced. That is now blocked in code: before any worker uses constrained mode, a canary request asks for one word while the schema allows only another, and the server has to answer with the schema's value. The probe is committed as reports/thor/constrained_guard_probe.json.

HTTP 422

Constrained mode refused when starting the reasoning-parser container.

HTTP 422

Constrained mode refused when saving settings with the reasoning parser on.

"banana"

The reasoning-parser server's answer to the canary. The guard refused.

schema value

The container without the parser honoured the grammar and was allowed. Ready 55.6 s after launch.

Operator console

Watch, filter and review in one place

Screens from the running app, captured with its built-in synthetic workcell feed and the mock detector. The console labels both on screen so they are never mistaken for a real camera or a real model.

Live page: synthetic workcell feed with boxes for a person marked hat missing and vest present, a robot and a pallet, a dashed restricted zone, a red No PPE banner and a column of high severity events marked review required.
Live. Each person's box shows what the model reported for the hard hat and the vest. Events land in the side column as they post, marked when a human must review them. The strip below shows frames, drops, the model and the poster's queue.
Events page: a table of safety events with time, severity, rule, camera, confidence and review status, with filters for camera, rule, severity and review required.
Events. Filter by camera, rule, severity or review required. Repeated events are grouped on the Incidents tab, and each event keeps the frame the model saw.
Model page: host memory and Docker status, the model server state, and a catalog of three Cosmos-Reason2 containers, one marked reasoning parser with no constrained mode.
Model. The catalog checks memory before anything starts and marks the reasoning-parser container as think mode only. Captured on a machine without the weights, so each shows weights missing.
Built for a real deployment

Details that matter on a factory floor

Every number has a file

Each run writes one artifact with the device, date, inputs, per-frame timings, token use, event counts, peak memory and a 1 Hz tegrastats record.

Failures are kept

A timed-out run, an empty-output run and a model echoing its own schema are all committed and labelled, not edited out.

Cameras set up in the UI

Ten vendor profiles, a connection test that decodes one frame or names the problem, and zones drawn on a live snapshot.

Credentials stay put

Passwords are encrypted at rest, never returned by the API, only reused for the host they were saved for, and masked in logs, events and reports.

Shared model server guarded

The model container keeps running across restarts. Stopping it needs an explicit, logged confirmation.

Nothing acts on its own

The system flags and records. It never stops a robot or a line; a person decides.

Status, October 2026

What is proven, and what is next

Working now
  • Runs end to end from camera setup to events in the console, and is deployed on the Thor as a service.
  • Runtime cost measured: pipeline overhead, model latency, power and temperature for 2B and 8B.
  • Constrained-mode guard measured against the real containers.
  • Tests and CI on every push to main.
Not yet measured
  • No live-camera run is committed. Capture rate, dropped frames, packet-to-event latency and power with a camera attached are unknown.
  • No rule accuracy. With no labelled ground truth, false events are counted, not rated.
  • The asynchronous poster has not been measured on the Thor.
  • Long-run stability. Committed runs last 60 to 390 seconds.
Planned upgrades
Live-camera evidence run on the Thor PPE check against known poses: vest, hat, both, neither Cosmos-Reason2-8B on live video Simulated scenes with ground truth for per-rule precision and recall Zones from camera calibration

Out of scope by design: stopping machines or acting without a person. This is observability with a human in the loop.

An EmbodiedEdge Labs project. Measured on real hardware, published as found.