Physical AI Safety Observability
Every number on this page is read from a run artifact under reports/thor/; each row links to its JSON. Device: jetsonthor, tegra264, # R38 (release), REVISION: 4.0, GCID: 43443517, BOARD: generic, EABI: aarch64, DATE: Wed Dec 31 00:15:19 UTC 2025, NV Power Mode: 120W 1. Runs on 2026-09-09, repository commit 254ef27. These are runtime-overhead and inference-cost measurements on simulation footage; none is a detection-quality number.
Runtime overhead, mock adapter
The mock adapter returns fixed detections at microsecond cost, so this is the price of frame sampling, the policy engine, event construction and the backend path. Events are four per frame by construction.
| Run | Frames | Frames/s | Rule eval p95, ms | Backend POST p50, ms | VIN p50, mW | Peak RSS, GB | Events |
|---|---|---|---|---|---|---|---|
| mock, 30 fps pacing, no posting mock_30fps_no_post.json | 1,800 | 28.588 | 0.9877 | n/a | 24,604 | 0.061 | 7,200 |
| mock, 30 fps pacing, per-event POST to FastAPI + SQLite mock_30fps_backend.json | 400 | 4.357 | 0.7123 | 160.9 | 25,570 | 0.051 | 1,600 |
| mock, 30 fps pacing, one batch POST per frame mock_30fps_backend_batched.json | 400 | 7.075 | 0.8565 | 68.2 | 25,158 | 0.057 | 1,600 |
| mock, unpaced, no posting mock_no_post.json | 500 | 3,077.556 | 0.2257 | n/a | n/a | 0.041 | 2,000 |
Real-model inference cost, Cosmos-Reason2 via vLLM on the same device
Same 60 frames (one every 60 from an Isaac Sim recording with no people in it), same server image, same device. Inference in seconds per frame. "Phantom persons" are detections labelled person on people-free footage; "events" are the safety events those detections fired, all false by construction.
| Run | Frames | p50, s | p95, s | max, s | Frames/s | Completion tokens p50 | VIN p50, mW | VIN peak, mW | Detections | Out-of-vocab labels | Phantom persons | Events |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Cosmos-Reason2-2B, think block, 4096 cap cosmos2b_video_60.json | 60 | 4.35 | 12.85 | 15.41 | 0.154 | 260 | 66,568 | 107,102 | 97 | 59 | 0 | 0 |
| Cosmos-Reason2-2B, no think, 512 cap cosmos2b_video_60_nothink.json | 60 | 3.57 | 6.58 | 9.35 | 0.241 | 157 | 60,338 | 105,206 | 106 | 57 | 0 | 0 |
| Cosmos-Reason2-2B, grammar, original prompt cosmos2b_video_60_schema_v1prompt.json | 60 | 0.31 | 0.33 | 4.19 | 2.303 | 7 | 37,052 | 101,448 | 0 | 0 | 0 | 0 |
| Cosmos-Reason2-2B, grammar, JSON-only prompt cosmos2b_video_60_schema.json | 60 | 2.69 | 3.36 | 8.03 | 0.367 | 135 | 62,522 | 83,120 | 193 | 0 | 6 | 7 |
| Cosmos-Reason2-8B, no think, 512 cap cosmos8b_video_60_nothink.json | 60 | 8.29 | 27.78 | 33.67 | 0.098 | 126 | 69,136 | 108,376 | 81 | 20 | 0 | 0 |
| Cosmos-Reason2-8B, grammar, JSON-only prompt cosmos8b_video_60_schema.json | 60 | 3.2 | 8.26 | 14.96 | 0.234 | 47 | 65,976 | 131,582 | 74 | 0 | 0 | 0 |
What is not established
- No precision or recall for any rule: there is no labelled ground truth. The grammar fixes the vocabulary; only ground truth can score the detections.
- No real camera input in these runs; the footage is simulation.
- Runs are 60 to 390 seconds, not sustained operation.
- Sixty frames of one video is a signal about the 8B, not a rate.