EmbodiedEdge Labs · Evidence

Bench2Field · Case study

Case study 01: YOLOX-s on the ROSMASTER rover

YOLOX-s, FP32 ONNX, TensorRT provider · Bench2Field 1.0.1 report, generated from 01_perception_detector/runs/
47.8 msfull frame on the Jetson Orin NX (rover), p95, against the 33.3 ms deadline: does not fit
26.9 msthe model alone on the same board at 26 Hz (inference plus copies, median of repeats)
25.9 Hzdelivered by the rover camera, configured 30 Hz; the headline tier is 26 Hz

Model alone versus the full frame

The model-alone bar is what b2f run times: inference plus the input and output copies, response p95 at 26 Hz, median of the baseline sweep's repeats. The full-frame bar is the pipeline profile on the same 300 frames of 640×480, set b39dce6618b8…: decode, preprocess, copies, inference and postprocess, one frame at a time.

Model alone versus the full frame, per board02040milliseconds, p95RTX 5090model alone 2.0 msfull frame 5.7 msJetson AGX Thormodel alone 11.0 msfull frame 16.2 msJetson Orin NX (rover)model alone 26.9 msfull frame 47.8 msdeadline 33.3 ms

Camera rate

Configured 30 Hz; delivered as measured with ros2 topic hz on the rover with its stack running. Every comparison on this page is made at the swept tier nearest the delivered rate, 26 Hz.

Delivered camera rate over the capture222426283032configured 30 Hzdelivered, mean 25.9 Hz59 window averages over the 60 s capture

Latency against request rate

Response p95 (median of repeats, bar = range) for the baseline sweeps. Solid lines are trt_fp32, dashed cuda_fp32. The same model gets slower as frames get sparser, because the GPU drops into lower power states between them.

Response p95 against request rate125102050100200500100020005000100002000050000100000102630100request rate, Hz (log)response p95, ms (log)deadline 33.3 mscamera 26 Hz✕✕
RTX 5090, trt_fp32RTX 5090, cuda_fp32Jetson AGX Thor, trt_fp32Jetson AGX Thor, cuda_fp32Jetson Orin NX (rover), trt_fp32Jetson Orin NX (rover), cuda_fp32✕ saturated: the tier could not be served

Power state by tier

trt_fp32, p50 over each tier. Where the board reports a GPU rail or an SM clock, it is shown; the Jetsons report board power, the 5090 GPU power only, and the two are never compared.

RTX 5090

tierp95, msGPU power, WSM clock, MHz
10 Hz2.2968.32,452
26 Hz2.0473.22,460
30 Hz2.0274.52,467
100 Hz1.9697.82,527

Jetson AGX Thor

tierp95, msboard power, WGPU rail, W
10 Hz11.3020.92.8
26 Hz11.0421.13.5
30 Hz10.9821.23.9
100 Hz4.2534.711.0

Jetson Orin NX (rover)

tierp95, msboard power, W
10 Hz33.967.1
26 Hz26.859.2
30 Hz23.029.7
100 Hz11,18523.8

Repeat-to-repeat spread on the field board

Jetson Orin NX (rover), trt_fp32 at 26 Hz, one row per repeat of the baseline sweep. Spread between repeats: 16% of the fastest. Temperature and GPU load are shown for each repeat as recorded; a difference between repeats is not attributed to either here, and the findings say what is and is not established about its cause.

repeatposition in sweepresponse p95, msp50junction °C p50peakGPU load %
1123.1022.8559.760.653
2326.8826.6463.364.569
3526.8526.5963.564.664

The same frames on every board

Stage cost per frame, p50, with the p95 total marked. 300 frames of 640×480, set b39dce6618b8….

Stage cost per frame on each board02040milliseconds per frame, p50 by stage; p95 total at rightRTX 5090p95 5.7Jetson AGX Thorp95 16.2Jetson Orin NX (rover)p95 47.8deadline 33.3 ms
decodepreprocesscopy to GPUinferencecopy backpostprocess
stage, p50 / p95 msRTX 5090Jetson AGX ThorJetson Orin NX (rover)
decode0.562 / 0.6233.415 / 3.4434.320 / 5.189
preprocess1.508 / 1.9981.689 / 1.7223.001 / 3.486
copy to GPU0.265 / 0.3330.629 / 0.6631.227 / 1.366
inference1.118 / 1.3105.469 / 5.68124.828 / 25.104
copy back0.229 / 0.3020.252 / 0.3161.195 / 1.322
postprocess0.992 / 1.2974.465 / 4.5059.533 / 11.676
total4.686 / 5.66715.899 / 16.21743.851 / 47.826
host stages (decode + preprocess + postprocess)3.1 ms, 65%9.6 ms, 60%16.9 ms, 38%
inference inside the frame / back to back1.1 / 1.55.5 / 3.224.8 / 16.8

Stage breakdown on the 5090 and the spin effect

onnxruntime's intra-op threads spin-wait between runs by default. In an inference-only sweep that costs almost nothing; in a frame with host-side stages they compete for the CPU.

Stage cost per frame on the 5090, spinning on and off02040milliseconds per frame, p50 by stage; p95 total at rightrover 720p, spin-onp95 44.7rover 720p, no spinp95 6.6
decodepreprocesscopy to GPUinferencecopy backpostprocess
stage, p50 / p95 msrover 720p, spin-onrover 720p, no spin
decode1.579 / 1.7711.565 / 1.634
preprocess3.286 / 39.9851.974 / 2.230
copy to GPU0.295 / 0.4190.258 / 0.297
inference1.122 / 1.3931.110 / 1.224
copy back0.248 / 0.2970.229 / 0.268
postprocess0.970 / 1.1990.962 / 1.031
total7.536 / 44.7486.088 / 6.584
host stages (decode + preprocess + postprocess)5.8 ms, 77%4.5 ms, 74%
inference inside the frame / back to back1.1 / 1.61.1 / 1.6

Provenance

Every run set this report draws on. Baselines are the sweeps marked as such in report.yaml; references are kept as evidence and shown in the tables above only where named.

boardfilesstatustiers × repeatsORT spinningruntimepower modecommitcontainers upstopped for the run
RTX 5090runs/bench_5090_nospin/baseline10, 26, 30, 100 Hz × 3offORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–ac95271b (recorded afterwards)telecom-pgnothing
RTX 5090runs/bench_5090_rerun/reference – spin-on, four tiers10, 26, 30, 100 Hz × 3onORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pgnothing
RTX 5090runs/bench_5090/reference – first sweep, three tiers, spin-on10, 30, 100 Hz × 3onORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pgnothing
Jetson AGX Thorruns/bench_thor_nospin/baseline10, 26, 30, 100 Hz × 3offORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2NV Power Mode: 120W6dad16b9nonesystemd user unit physical-ai-safety.service (Safety Observability API on :8081) stopped with systemctl --user stop, then docker stop physical-ai-vllm (--rm, removed); confirmed no container and no GPU client after 3 min; both restored after the sweep
Jetson AGX Thorruns/bench_thor_4tier_spin/reference – clean of vLLM but spin-on (stale checkout)10, 26, 30, 100 Hz × 3onORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2NV Power Mode: 120Wnot recordednonesystemd user unit physical-ai-safety.service (Safety Observability API on :8081) stopped with systemctl --user stop; its physical-ai-vllm container (--rm) went with it; both restored after the sweep
Jetson AGX Thorruns/bench_thor/reference – first sweep, three tiers, spin-on, vLLM container up10, 30, 100 Hz × 3onORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2NV Power Mode: 120Wnot recordedphysical-ai-vllmdocker container physical-ai-vllm (stopped before the sweep, started again after it)
Jetson AGX Thorruns/bench_thor_rerun_stopped/stopped, not a baseline – stopped after one run, not a baseline10, 26, 30, 100 Hz × 3 (incomplete)onORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2NV Power Mode: 120Wnot recordedphysical-ai-vllmnothing
Jetson Orin NX (rover)runs/field_orin/baseline10, 26, 30, 100 Hz × 3offORT 1.24.0, TRT 10.7.0, cuDNN 9.3.0, CUDA 12.6NV Power Mode: MAXN_SUPERac95271b (recorded afterwards)nonenothing
RTX 5090runs/profile_5090_rover_frames_720p.jsonpipeline profile – rover 720p, spin-on300 frames, data/rover_frames_720ponORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pg–
RTX 5090runs/profile_5090_rover_frames_720p_nospin.jsonpipeline profile – rover 720p, no spin300 frames, data/rover_frames_720poffORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pg–
RTX 5090runs/profile_5090_rover_frames_480p.jsonpipeline profile – rover 480p, spin-on300 frames, data/rover_frames_480ponORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pg–
RTX 5090runs/profile_5090_scene_frames_640.jsonpipeline profile – busy scene 640, spin-on300 frames, data/scene_frames_640onORT 1.30.0, TRT 10.16.1, cuDNN 9.19.0, CUDA 13.0–not recordedtelecom-pg–
Jetson AGX Thorruns/profile_thor_rover_frames_480p.jsonpipeline profile – rover 480p, no spin300 frames, data/rover_frames_480poffORT 1.24.0, TRT 10.13.3, cuDNN 9.12.0, CUDA 13.2NV Power Mode: 120Wdadbabefphysical-ai-vllm–
Jetson Orin NX (rover)runs/profile_orin_rover_frames_480p.jsonpipeline profile – rover 480p, no spin300 frames, data/rover_frames_480poffORT 1.24.0, TRT 10.7.0, cuDNN 9.3.0, CUDA 12.6NV Power Mode: MAXN_SUPERnoted afterwardsnone–