EmbodiedEdge Labs · Evidence
Bench2Field · Model benchmarking measured on the robot v1.0.1 released · phase 2 under way

The benchmark said GO. The robot would drop frames.

A benchmark tells you how fast a model runs on a bench. Bench2Field measures how much of that survives on the robot, and where the rest goes. It runs the model at the rate the robot's camera really delivers, times the whole frame and not just the model, and refuses any number that doesn't trace to a committed run.

Field boardJetson Orin NXinside a ROSMASTER M3 Pro rover
Bench boardsJetson AGX Thor, RTX 5090one report schema across all three
Case study 01YOLOX-s detectorFP32 ONNX through TensorRT
Toolb2f command linePython, 145 tests, Apache 2.0
Phase 1 result · field board, Orin NX · 26 Hz camera

Full frame misses the 33.3 ms deadline
Rover camera, Orbbec DaBai DCW2
0 Hz25.9 delivered30 configured

The launch file asks for 640×480 at 30 fps. With the rover's stack running, the colour topic arrived at a mean of 25.9 Hz over 60 s (windows 23.7 to 27.0), because the driver decodes MJPEG on the CPU. The benchmark follows the robot, so 26 Hz is the headline tier.

Bench vs field, whole frame

Same frames, three boards: only the rover misses

Median time per stage, 300 frames from the rover's camera at 640×480, one frame at a time, TensorRT FP32, threads not spinning. Same frames and code on every board, drawn to one scale. The black figure at the end of each bar is the frame's p95.

Bench vs field

Latency depends on how often frames arrive

Response time from frame arrival to result, median of three repeats; whiskers show the range across repeats. Log scale, because the boards differ by more than ten times. Sparse frames let the GPU drop into a lower power state, and each frame pays to wake it.

Why

The power state follows the frame rate

Each board's own telemetry, TensorRT FP32, median across repeats. The measures differ by board, so each gets its own axis.

RTX 5090 · SM clock (MHz)

Thor · GPU rail (W)

Orin NX · CPU+GPU rail (W)

The Jetsons report no GPU clock through tegrastats, so their power-state explanation is inferred from the rails rather than measured directly. At 100 Hz the Orin is saturated, so its rail shows full load.

End-to-end profile, RTX 5090

Inference is the smallest part of the frame

Median time per stage over 300 frames, one frame at a time, TensorRT FP32. Rover frames come from the DaBai camera; the busy scene comes from recorded video.

ONNX Runtime thread spinning, rover 720p

Spin-waiting threads steal CPU from the host stages. Turning spinning off leaves inference unchanged and collapses the tail, which is why every baseline above runs with it off.

Repeats, Orin NX

Why every number is a median of alternated repeats

TensorRT FP32 at 26 Hz, the three phase 1 repeats in the order they ran. The first came in 3.8 ms faster than the other two.

Corrected in v1.0.1

v1.0 put the fast repeat down to the board being cooler. A later sweep on the same board measured 26.84 to 26.90 ms in all three repeats at similar temperatures, so heat doesn't explain it and the claim was withdrawn. The cause isn't established. The median, 26.9 ms, is the board's figure, and the repeats are what exposed the outlier.

Method

Timed the way the robot sees it

01 · arriveFrames at the camera's rate

Requests come at a fixed rate whether or not the last frame is done, like a camera.

02 · timeFrom arrival, not from start

A frame is late relative to when it arrived, so queueing shows up in the percentiles.

03 · profileEvery stage of the frame

Decode, preprocess, copies, inference and NMS, per stage, on the same frames on every board.

04 · repeatAlternated repeats

Variants run A, B, A, B in separate processes, so board state can't favour one of them.

05 · decideVerdict and report

GO or NO-GO against a budget written before testing, and this page built from the runs.

Every number traces to a committed run

Each run file records the code commit, library versions, power mode, thread settings and everything else running on the board. The report is built from those files and nothing else, so it can't drift from the data.

The tool

What b2f does

One schema for every run, so a workstation GPU and two Jetsons compare directly. Telemetry comes from NVML on discrete GPUs and tegrastats on Jetson; a rocm-smi reader for AMD is written but not yet run on AMD hardware.

b2f run

Benchmark an ONNX model at fixed request rates through ONNX Runtime on CPU, CUDA or TensorRT, with MIGraphX and ROCm wired in for the AMD phase. Records latency from arrival, deadline misses, power, clocks and temperature.

b2f sweep

Repeats of several variants, alternated, one process per run, with a manifest of the order. Refuses a checkout that doesn't match the expected commit.

b2f record-load

Record the robot's own background load so it can be replayed on the bench as CPU, memory-bandwidth and GPU stressors.

b2f retention

Of the speedup an optimization shows on the bench, how much is left on the robot. Warns when power modes or library versions differ between the two sides.

b2f attribute

Split the bench-to-field loss into causes, one stressor at a time, with the unexplained remainder reported rather than hidden.

b2f verdict

GO, NO-GO or INCOMPLETE against a deployment budget: deadline, miss rate, power, temperature, accuracy.

b2f report

One self-contained HTML page from a case study's committed runs. Works offline, light and dark.

Controls

Built so the numbers hold up

Provenance on every run

Commit and dirty flag, ONNX Runtime, TensorRT, cuDNN and CUDA versions as loaded, power mode, thread settings and TensorRT build options.

What else was running

Each report snapshots containers, the busiest processes and load average, so a quiet bench has to prove it was quiet.

Nothing changed silently

Power modes, fans and governors are recorded, never changed during a run. Any board change is the owner's decision.

Reproduced from a clean clone

A fresh clone following only the README reproduced a baseline and the report. The first pass found two install gaps, and the README was fixed.

Model provenance

The exported model's hash, export options and a check against the original framework's output are committed.

Speed is not accuracy

No optimized variant is claimed until its accuracy is measured. Phase 2 speed results stay labelled until then.

Status, October 2026

Measured, and still to measure

v1: measure and diagnose
  • Bring-up on three boards with real-hardware fixtures for every telemetry parser.
  • FP32 baselines on all three boards at 10, 26, 30 and 100 Hz, three alternated repeats each.
  • Full-frame profiles on all three boards on the rover's own frames, plus an Nsight Systems capture on the 5090.
  • 145 tests on Python 3.10 and 3.12, CI on every push, tagged releases with the report attached.
Not yet measured
  • Accuracy of any optimized variant. The harness comes first in the next phase.
  • The rover under load: field runs with its SLAM stack up, load replay on the bench, and gap attribution.
  • Busy scenes. The profiled frames show a plain wall, so decode and NMS were at their cheapest.
  • AMD hardware for the ROCm port.
v2: optimize and close the gap
Accuracy harnessFP16 and INT8 with calibration2:4 structured sparsityDistillation, YOLOX-l to YOLOX-sFused CUDA preprocessing kernelZero-copy on Jetson, pinned memory on the 5090HIP port and MIGraphX on MI300XField load replay and attributionTriton serving on the 5090
Provenance

Every run behind this page

Versions, power mode, thread spinning, code commit and what else was running, as recorded in each report.

Rows marked reference are kept as evidence: spin-on runs, the first Thor sweep with an idle vLLM container resident, and a stopped re-run. A dash in the commit column means the run predates the provenance guard. The Thor and Orin differ in TensorRT and cuDNN because they run JetPack 7 and 6.

An EmbodiedEdge Labs project. Measured on real hardware, published as found.