EmbodiedEdge Labs · Evidence
Edge vision analytics · Streets and intersections Functional · field validation in progress

Urban Edge Vision

Network cameras are set up in a web console. A vision-language model running on the device looks at each frame, use-case packs turn what it sees into traffic events, and an operator reviews them. Processing stays on the device, and nothing is enforced automatically.

Field deviceJetson AGX ThorCosmos-Reason2-2B on vLLM
WorkstationRTX 5090 class GPUGemma 4 on Ollama or vLLM
Cameras10 vendor profilesRTSP, HTTP MJPEG, webcam
Packs4 use casesper camera, combinable
Site plan

Four packs, drawn on the street

Each pack works on geometry the operator draws in the Studio page over a live snapshot from the camera. Coordinates are stored normalized to the frame, so they survive a change of stream resolution.

STOP t12 1 STOP ZONE A B distance entered in metres t31 → 2 SPEED BETWEEN GATES A B Eastbound Westbound ← t44 bus t45 3 VEHICLE COUNT p7 → 4 MOVING OBJECT ZONE N ↑ Schematic. Track ids are illustrative, not from a run.
Geometry stored 0 to 1, relative to the frameOne camera can carry several packs
  1. 1
    Stop sign

    Watches each vehicle track through the stop zone and calls it a full stop, a rolling stop or no stop from its dwell time and slowest speed in the zone.

  2. 2
    Speed between two gates

    Times a track from gate A to gate B and converts the distance the operator entered into km/h or mph. Only measured crossings produce a reading.

  3. 3
    Vehicle count

    Counts each tracked vehicle once as it crosses the A to B line, by direction and by class: car, truck, bus, motorcycle. Hourly totals per camera.

  4. 4
    Moving object

    Tracks people or other chosen classes with direction and a short descriptor, across the whole frame or inside a drawn zone.

Speed and stop sign are refused on the same camera. They need different sight lines, so the console blocks the combination instead of producing weak results for both.

Signal path, per camera

From stream to reviewed event

01 · captureDecode at native rate

One capture thread per camera writes the newest frame to a single slot.

02 · watchLive video

The browser reads that slot as MJPEG at the camera's own frame rate.

03 · lookModel on its own cadence

Inference samples the slot every interval. vLLM, Ollama or NVIDIA NIM, switchable live.

04 · decideTracker and packs

Detections get stable track ids, then each pack bound to the camera applies its rule.

05 · reviewOperator queue

Events persist in SQLite. Flagged ones wait in Review with the frame the model saw.

Video never waits on the model

The model server is not in the video path. If it stops or falls behind, the Live page keeps playing and shows "inference unavailable" instead of inventing events. An end-to-end runtime test checks exactly that.

Decision rules, as shipped

What each pack actually decides

The defaults below are read from the pack code. Every one can be changed per camera in Studio.

Stop sign

full stop slowest ≤ 0.02 frame widths/s and dwell ≥ 1,000 ms rolling stop slowest ≤ 0.08 no stop anything faster
  • Decided once, when the track leaves the zone
  • Rolling and no stop go to the Review queue

Speed between two gates

km/h = distance_m ÷ (t_B − t_A) × 3.6 discard if transit > 30 s compare with posted limit (default 50 km/h)
  • Distance comes from a tape measure or a map, entered once
  • Readings shown in km/h or mph

Vehicle count

count once per crossing of A→B or B→A track seen ≥ 2 times before counting jitter on the line = one record
  • Direction labels are the operator's, such as Eastbound
  • Counts never enter the Review queue

Moving object

classes pedestrian by default zone whole frame, or a drawn polygon report direction, speed band, descriptor
  • Useful near crosswalks and sidewalks
  • Works on any camera, street view or not
Operator console

Set up, watch and review in one place

Screens from the running app. These captures use the built-in synthetic test feed and the mock detector, and the console labels both on screen so they are never mistaken for a real camera.

Live page: synthetic camera with detections, the dashed count line from A to B, vehicle counts by direction and a feed of recent events.
Live. Video at the camera's frame rate with detections and the count line drawn over it. Vehicle counts by class and direction, recent events, and per-camera health: frame rate, dropped frames, reconnects, packs running.
Add camera form with the Hikvision profile selected, a documentation-range host address, port 554 and the stream path filled in from the profile.
Cameras. Pick a vendor profile and the stream path fills itself in for main or sub stream. Test connection decodes one frame or names the problem: unreachable, wrong password, wrong path, codec.
Use-case Studio with Moving Object, Stop Sign and Vehicle Count selected and Speed Violation marked blocked.
Studio. Choose packs per camera and draw their zones, gates and count line on a snapshot. Here Speed is blocked because Stop sign is selected on the same camera.
Models page showing the host's GPU and memory, the current model's health, and a catalog where models that will not fit are marked with the reason.
Models. The catalog is checked against the host before anything loads: free VRAM on a discrete GPU, available memory on Jetson. Models that will not fit are shown blocked with the reason. On Jetson the console can launch NVIDIA's vLLM container itself.
Built for a real deployment

Details that matter once it is on a pole

Credentials stay put

Camera passwords are encrypted at rest with a key file outside the repository. The API never returns one, and a saved password is only reused for the host and port it was saved for.

Nothing leaks into logs

Stream URLs are masked in logs, errors, status payloads and run artifacts.

Right model for the box

On first start the console picks a default for the hardware: Cosmos-Reason2-2B on Thor, Gemma 4 on a workstation. The operator can switch at any time.

Shared model server guarded

A model container another app depends on cannot be stopped by accident. Stopping it needs an explicit, logged confirmation.

Runs as a service

A systemd unit starts the API and console at boot and keeps a persistent log.

People decide

Packs flag, operators confirm or dismiss. Each item can carry a ground-truth note, such as "my car, full stop", so pack output can be checked against known passes.

Status, October 2026

What is proven, and what is next

Working now
  • Runs end to end from camera setup to the Review queue, on the Thor and on a workstation.
  • 271 automated tests pass covering camera setup, connection probing, credential handling, model switching and memory preflight, all four packs and the review flow.
  • CI on every push: lint, type checks, tests and the console build.
Not yet measured
  • Live-camera latency and power on the Thor. No live-camera run is committed yet.
  • Pack accuracy against known passes: a car at a known speed, a full stop, a rolling stop.
  • Speed calibration error, which depends on gate placement, the entered distance and the inference interval.
  • Long-run stability over hours and days, and behaviour at night or in rain.
Planned upgrades
Live-camera evidence runs on the Thor Ground-truth check of speed and stop sign Recorded-video summarization (NVIDIA VSS) AWS edge-to-cloud sync Wrong-way detection Illegal parking Pedestrian counts and yield Queue length Near-miss detection Red-light running

Out of scope by design: license plate reading and automated ticketing. This is operational analysis with a person in the loop.

An EmbodiedEdge Labs project. Measured on real hardware, published as found.