Bench2Field · Case study
Case study 01: YOLOX-s on the ROSMASTER rover
01_perception_detector/runs/Model alone versus the full frame
The model-alone bar is what b2f run times: inference plus the input and output copies, response p95 at 26 Hz, median of the baseline sweep's repeats. The full-frame bar is the pipeline profile on the same 300 frames of 640×480, set b39dce6618b8…: decode, preprocess, copies, inference and postprocess, one frame at a time.
Camera rate
Configured 30 Hz; delivered as measured with ros2 topic hz on the rover with its stack running. Every comparison on this page is made at the swept tier nearest the delivered rate, 26 Hz.
Latency against request rate
Response p95 (median of repeats, bar = range) for the baseline sweeps. Solid lines are trt_fp32, dashed cuda_fp32. The same model gets slower as frames get sparser, because the GPU drops into lower power states between them.
Power state by tier
trt_fp32, p50 over each tier. Where the board reports a GPU rail or an SM clock, it is shown; the Jetsons report board power, the 5090 GPU power only, and the two are never compared.
RTX 5090
| tier | p95, ms | GPU power, W | SM clock, MHz |
|---|---|---|---|
| 10 Hz | 2.29 | 68.3 | 2,452 |
| 26 Hz | 2.04 | 73.2 | 2,460 |
| 30 Hz | 2.02 | 74.5 | 2,467 |
| 100 Hz | 1.96 | 97.8 | 2,527 |
Jetson AGX Thor
| tier | p95, ms | board power, W | GPU rail, W |
|---|---|---|---|
| 10 Hz | 11.30 | 20.9 | 2.8 |
| 26 Hz | 11.04 | 21.1 | 3.5 |
| 30 Hz | 10.98 | 21.2 | 3.9 |
| 100 Hz | 4.25 | 34.7 | 11.0 |
Jetson Orin NX (rover)
| tier | p95, ms | board power, W |
|---|---|---|
| 10 Hz | 33.96 | 7.1 |
| 26 Hz | 26.85 | 9.2 |
| 30 Hz | 23.02 | 9.7 |
| 100 Hz | 11,185 | 23.8 |
Repeat-to-repeat spread on the field board
Jetson Orin NX (rover), trt_fp32 at 26 Hz, one row per repeat of the baseline sweep. Spread between repeats: 16% of the fastest. Temperature and GPU load are shown for each repeat as recorded; a difference between repeats is not attributed to either here, and the findings say what is and is not established about its cause.
| repeat | position in sweep | response p95, ms | p50 | junction °C p50 | peak | GPU load % |
|---|---|---|---|---|---|---|
| 1 | 1 | 23.10 | 22.85 | 59.7 | 60.6 | 53 |
| 2 | 3 | 26.88 | 26.64 | 63.3 | 64.5 | 69 |
| 3 | 5 | 26.85 | 26.59 | 63.5 | 64.6 | 64 |
The same frames on every board
Stage cost per frame, p50, with the p95 total marked. 300 frames of 640×480, set b39dce6618b8….
| stage, p50 / p95 ms | RTX 5090 | Jetson AGX Thor | Jetson Orin NX (rover) |
|---|---|---|---|
| decode | 0.562 / 0.623 | 3.415 / 3.443 | 4.320 / 5.189 |
| preprocess | 1.508 / 1.998 | 1.689 / 1.722 | 3.001 / 3.486 |
| copy to GPU | 0.265 / 0.333 | 0.629 / 0.663 | 1.227 / 1.366 |
| inference | 1.118 / 1.310 | 5.469 / 5.681 | 24.828 / 25.104 |
| copy back | 0.229 / 0.302 | 0.252 / 0.316 | 1.195 / 1.322 |
| postprocess | 0.992 / 1.297 | 4.465 / 4.505 | 9.533 / 11.676 |
| total | 4.686 / 5.667 | 15.899 / 16.217 | 43.851 / 47.826 |
| host stages (decode + preprocess + postprocess) | 3.1 ms, 65% | 9.6 ms, 60% | 16.9 ms, 38% |
| inference inside the frame / back to back | 1.1 / 1.5 | 5.5 / 3.2 | 24.8 / 16.8 |
Stage breakdown on the 5090 and the spin effect
onnxruntime's intra-op threads spin-wait between runs by default. In an inference-only sweep that costs almost nothing; in a frame with host-side stages they compete for the CPU.
| stage, p50 / p95 ms | rover 720p, spin-on | rover 720p, no spin |
|---|---|---|
| decode | 1.579 / 1.771 | 1.565 / 1.634 |
| preprocess | 3.286 / 39.985 | 1.974 / 2.230 |
| copy to GPU | 0.295 / 0.419 | 0.258 / 0.297 |
| inference | 1.122 / 1.393 | 1.110 / 1.224 |
| copy back | 0.248 / 0.297 | 0.229 / 0.268 |
| postprocess | 0.970 / 1.199 | 0.962 / 1.031 |
| total | 7.536 / 44.748 | 6.088 / 6.584 |
| host stages (decode + preprocess + postprocess) | 5.8 ms, 77% | 4.5 ms, 74% |
| inference inside the frame / back to back | 1.1 / 1.6 | 1.1 / 1.6 |
Provenance
Every run set this report draws on. Baselines are the sweeps marked as such in report.yaml; references are kept as evidence and shown in the tables above only where named.
| board | files | status | tiers × repeats | ORT spinning | runtime | power mode | commit | containers up | stopped for the run |
|---|---|---|---|---|---|---|---|---|---|
| RTX 5090 | runs/bench_5090_nospin/ | baseline | 10, 26, 30, 100 Hz × 3 | off | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | ac95271b (recorded afterwards) | telecom-pg | nothing |
| RTX 5090 | runs/bench_5090_rerun/ | reference – spin-on, four tiers | 10, 26, 30, 100 Hz × 3 | on | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | nothing |
| RTX 5090 | runs/bench_5090/ | reference – first sweep, three tiers, spin-on | 10, 30, 100 Hz × 3 | on | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | nothing |
| Jetson AGX Thor | runs/bench_thor_nospin/ | baseline | 10, 26, 30, 100 Hz × 3 | off | ORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2 | NV Power Mode: 120W | 6dad16b9 | none | systemd user unit physical-ai-safety.service (Safety Observability API on :8081) stopped with systemctl --user stop, then docker stop physical-ai-vllm (--rm, removed); confirmed no container and no GPU client after 3 min; both restored after the sweep |
| Jetson AGX Thor | runs/bench_thor_4tier_spin/ | reference – clean of vLLM but spin-on (stale checkout) | 10, 26, 30, 100 Hz × 3 | on | ORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2 | NV Power Mode: 120W | not recorded | none | systemd user unit physical-ai-safety.service (Safety Observability API on :8081) stopped with systemctl --user stop; its physical-ai-vllm container (--rm) went with it; both restored after the sweep |
| Jetson AGX Thor | runs/bench_thor/ | reference – first sweep, three tiers, spin-on, vLLM container up | 10, 30, 100 Hz × 3 | on | ORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2 | NV Power Mode: 120W | not recorded | physical-ai-vllm | docker container physical-ai-vllm (stopped before the sweep, started again after it) |
| Jetson AGX Thor | runs/bench_thor_rerun_stopped/ | stopped, not a baseline – stopped after one run, not a baseline | 10, 26, 30, 100 Hz × 3 (incomplete) | on | ORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2 | NV Power Mode: 120W | not recorded | physical-ai-vllm | nothing |
| Jetson Orin NX (rover) | runs/field_orin/ | baseline | 10, 26, 30, 100 Hz × 3 | off | ORT 1.24.0, TRT 10.7.0, cuDNN 9.3.0, CUDA 12.6 | NV Power Mode: MAXN_SUPER | ac95271b (recorded afterwards) | none | nothing |
| RTX 5090 | runs/profile_5090_rover_frames_720p.json | pipeline profile – rover 720p, spin-on | 300 frames, data/rover_frames_720p | on | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | – |
| RTX 5090 | runs/profile_5090_rover_frames_720p_nospin.json | pipeline profile – rover 720p, no spin | 300 frames, data/rover_frames_720p | off | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | – |
| RTX 5090 | runs/profile_5090_rover_frames_480p.json | pipeline profile – rover 480p, spin-on | 300 frames, data/rover_frames_480p | on | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | – |
| RTX 5090 | runs/profile_5090_scene_frames_640.json | pipeline profile – busy scene 640, spin-on | 300 frames, data/scene_frames_640 | on | ORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0 | – | not recorded | telecom-pg | – |
| Jetson AGX Thor | runs/profile_thor_rover_frames_480p.json | pipeline profile – rover 480p, no spin | 300 frames, data/rover_frames_480p | off | ORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2 | NV Power Mode: 120W | dadbabef | physical-ai-vllm | – |
| Jetson Orin NX (rover) | runs/profile_orin_rover_frames_480p.json | pipeline profile – rover 480p, no spin | 300 frames, data/rover_frames_480p | off | ORT 1.24.0, TRT 10.7.0, cuDNN 9.3.0, CUDA 12.6 | NV Power Mode: MAXN_SUPER | noted afterwards | none | – |