A benchmark tells you how fast a model runs on a bench. Bench2Field measures how much of that survives on the robot, and where the rest goes. It runs the model at the rate the robot's camera really delivers, times the whole frame and not just the model, and refuses any number that doesn't trace to a committed run.
The launch file asks for 640×480 at 30 fps. With the rover's stack running, the colour topic arrived at a mean of 25.9 Hz over 60 s (windows 23.7 to 27.0), because the driver decodes MJPEG on the CPU. The benchmark follows the robot, so 26 Hz is the headline tier.
Median time per stage, 300 frames from the rover's camera at 640×480, one frame at a time, TensorRT FP32, threads not spinning. Same frames and code on every board, drawn to one scale. The black figure at the end of each bar is the frame's p95.
Response time from frame arrival to result, median of three repeats; whiskers show the range across repeats. Log scale, because the boards differ by more than ten times. Sparse frames let the GPU drop into a lower power state, and each frame pays to wake it.
Each board's own telemetry, TensorRT FP32, median across repeats. The measures differ by board, so each gets its own axis.
The Jetsons report no GPU clock through tegrastats, so their power-state explanation is inferred from the rails rather than measured directly. At 100 Hz the Orin is saturated, so its rail shows full load.
Median time per stage over 300 frames, one frame at a time, TensorRT FP32. Rover frames come from the DaBai camera; the busy scene comes from recorded video.
Spin-waiting threads steal CPU from the host stages. Turning spinning off leaves inference unchanged and collapses the tail, which is why every baseline above runs with it off.
TensorRT FP32 at 26 Hz, the three phase 1 repeats in the order they ran. The first came in 3.8 ms faster than the other two.
v1.0 put the fast repeat down to the board being cooler. A later sweep on the same board measured 26.84 to 26.90 ms in all three repeats at similar temperatures, so heat doesn't explain it and the claim was withdrawn. The cause isn't established. The median, 26.9 ms, is the board's figure, and the repeats are what exposed the outlier.
Requests come at a fixed rate whether or not the last frame is done, like a camera.
A frame is late relative to when it arrived, so queueing shows up in the percentiles.
Decode, preprocess, copies, inference and NMS, per stage, on the same frames on every board.
Variants run A, B, A, B in separate processes, so board state can't favour one of them.
GO or NO-GO against a budget written before testing, and this page built from the runs.
Each run file records the code commit, library versions, power mode, thread settings and everything else running on the board. The report is built from those files and nothing else, so it can't drift from the data.
One schema for every run, so a workstation GPU and two Jetsons compare directly. Telemetry comes from NVML on discrete GPUs and tegrastats on Jetson; a rocm-smi reader for AMD is written but not yet run on AMD hardware.
b2f runBenchmark an ONNX model at fixed request rates through ONNX Runtime on CPU, CUDA or TensorRT, with MIGraphX and ROCm wired in for the AMD phase. Records latency from arrival, deadline misses, power, clocks and temperature.
b2f sweepRepeats of several variants, alternated, one process per run, with a manifest of the order. Refuses a checkout that doesn't match the expected commit.
b2f record-loadRecord the robot's own background load so it can be replayed on the bench as CPU, memory-bandwidth and GPU stressors.
b2f retentionOf the speedup an optimization shows on the bench, how much is left on the robot. Warns when power modes or library versions differ between the two sides.
b2f attributeSplit the bench-to-field loss into causes, one stressor at a time, with the unexplained remainder reported rather than hidden.
b2f verdictGO, NO-GO or INCOMPLETE against a deployment budget: deadline, miss rate, power, temperature, accuracy.
b2f reportOne self-contained HTML page from a case study's committed runs. Works offline, light and dark.
Commit and dirty flag, ONNX Runtime, TensorRT, cuDNN and CUDA versions as loaded, power mode, thread settings and TensorRT build options.
Each report snapshots containers, the busiest processes and load average, so a quiet bench has to prove it was quiet.
Power modes, fans and governors are recorded, never changed during a run. Any board change is the owner's decision.
A fresh clone following only the README reproduced a baseline and the report. The first pass found two install gaps, and the README was fixed.
The exported model's hash, export options and a check against the original framework's output are committed.
No optimized variant is claimed until its accuracy is measured. Phase 2 speed results stay labelled until then.
Versions, power mode, thread spinning, code commit and what else was running, as recorded in each report.
Rows marked reference are kept as evidence: spin-on runs, the first Thor sweep with an idle vLLM container resident, and a stopped re-run. A dash in the commit column means the run predates the provenance guard. The Thor and Orin differ in TensorRT and cuDNN because they run JetPack 7 and 6.