Edge vision analytics · Streets and intersectionsFunctional · field validation in progress
Urban EdgeVision
Network cameras are set up in a web console. A vision-language model running on the device looks at each frame, use-case packs turn what it sees into traffic events, and an operator reviews them. Processing stays on the device, and nothing is enforced automatically.
Field deviceJetson AGX ThorCosmos-Reason2-2B on vLLM
WorkstationRTX 5090 class GPUGemma 4 on Ollama or vLLM
Cameras10 vendor profilesRTSP, HTTP MJPEG, webcam
Packs4 use casesper camera, combinable
Site plan
Four packs, drawn on the street
Each pack works on geometry the operator draws in the Studio page over a live snapshot from the camera. Coordinates are stored normalized to the frame, so they survive a change of stream resolution.
Geometry stored 0 to 1, relative to the frameOne camera can carry several packs
1
Stop sign
Watches each vehicle track through the stop zone and calls it a full stop, a rolling stop or no stop from its dwell time and slowest speed in the zone.
2
Speed between two gates
Times a track from gate A to gate B and converts the distance the operator entered into km/h or mph. Only measured crossings produce a reading.
3
Vehicle count
Counts each tracked vehicle once as it crosses the A to B line, by direction and by class: car, truck, bus, motorcycle. Hourly totals per camera.
4
Moving object
Tracks people or other chosen classes with direction and a short descriptor, across the whole frame or inside a drawn zone.
Speed and stop sign are refused on the same camera. They need different sight lines, so the console blocks the combination instead of producing weak results for both.
Signal path, per camera
From stream to reviewed event
01 · captureDecode at native rate
One capture thread per camera writes the newest frame to a single slot.
02 · watchLive video
The browser reads that slot as MJPEG at the camera's own frame rate.
03 · lookModel on its own cadence
Inference samples the slot every interval. vLLM, Ollama or NVIDIA NIM, switchable live.
04 · decideTracker and packs
Detections get stable track ids, then each pack bound to the camera applies its rule.
05 · reviewOperator queue
Events persist in SQLite. Flagged ones wait in Review with the frame the model saw.
Video never waits on the model
The model server is not in the video path. If it stops or falls behind, the Live page keeps playing and shows "inference unavailable" instead of inventing events. An end-to-end runtime test checks exactly that.
Decision rules, as shipped
What each pack actually decides
The defaults below are read from the pack code. Every one can be changed per camera in Studio.
Stop sign
full stop slowest ≤ 0.02 frame widths/s
and dwell ≥ 1,000 ms
rolling stop slowest ≤ 0.08
no stop anything faster
Decided once, when the track leaves the zone
Rolling and no stop go to the Review queue
Speed between two gates
km/h = distance_m ÷ (t_B − t_A) × 3.6
discard if transit > 30 s
compare with posted limit (default 50 km/h)
Distance comes from a tape measure or a map, entered once
Readings shown in km/h or mph
Vehicle count
count once per crossing of A→B or B→A
track seen ≥ 2 times before counting
jitter on the line = one record
Direction labels are the operator's, such as Eastbound
Counts never enter the Review queue
Moving object
classes pedestrian by default
zone whole frame, or a drawn polygon
report direction, speed band, descriptor
Useful near crosswalks and sidewalks
Works on any camera, street view or not
Operator console
Set up, watch and review in one place
Screens from the running app. These captures use the built-in synthetic test feed and the mock detector, and the console labels both on screen so they are never mistaken for a real camera.
Live. Video at the camera's frame rate with detections and the count line drawn over it. Vehicle counts by class and direction, recent events, and per-camera health: frame rate, dropped frames, reconnects, packs running.Cameras. Pick a vendor profile and the stream path fills itself in for main or sub stream. Test connection decodes one frame or names the problem: unreachable, wrong password, wrong path, codec.Studio. Choose packs per camera and draw their zones, gates and count line on a snapshot. Here Speed is blocked because Stop sign is selected on the same camera.Models. The catalog is checked against the host before anything loads: free VRAM on a discrete GPU, available memory on Jetson. Models that will not fit are shown blocked with the reason. On Jetson the console can launch NVIDIA's vLLM container itself.
Built for a real deployment
Details that matter once it is on a pole
Credentials stay put
Camera passwords are encrypted at rest with a key file outside the repository. The API never returns one, and a saved password is only reused for the host and port it was saved for.
Nothing leaks into logs
Stream URLs are masked in logs, errors, status payloads and run artifacts.
Right model for the box
On first start the console picks a default for the hardware: Cosmos-Reason2-2B on Thor, Gemma 4 on a workstation. The operator can switch at any time.
Shared model server guarded
A model container another app depends on cannot be stopped by accident. Stopping it needs an explicit, logged confirmation.
Runs as a service
A systemd unit starts the API and console at boot and keeps a persistent log.
People decide
Packs flag, operators confirm or dismiss. Each item can carry a ground-truth note, such as "my car, full stop", so pack output can be checked against known passes.
Status, October 2026
What is proven, and what is next
Working now
Runs end to end from camera setup to the Review queue, on the Thor and on a workstation.
271 automated tests pass covering camera setup, connection probing, credential handling, model switching and memory preflight, all four packs and the review flow.
CI on every push: lint, type checks, tests and the console build.
Not yet measured
Live-camera latency and power on the Thor. No live-camera run is committed yet.
Pack accuracy against known passes: a car at a known speed, a full stop, a rolling stop.
Speed calibration error, which depends on gate placement, the entered distance and the inference interval.
Long-run stability over hours and days, and behaviour at night or in rain.
Planned upgrades
Live-camera evidence runs on the ThorGround-truth check of speed and stop signRecorded-video summarization (NVIDIA VSS)AWS edge-to-cloud syncWrong-way detectionIllegal parkingPedestrian counts and yieldQueue lengthNear-miss detectionRed-light running
Out of scope by design: license plate reading and automated ticketing. This is operational analysis with a person in the loop.
An EmbodiedEdge Labs project. Measured on real hardware, published as found.