Pre-launch · Technical demonstration

Warden Edge

A technical demonstration of governed, human-in-the-loop edge-AI perimeter patrol — the patrol unit and the person authorizing it on one real-time link, on site or remote.

“Every autonomous security robot acts on its own. Ours holds and asks.”

Pre-launch — a concept and working technical demonstration, not a shipping product.

Governed patrol, not autonomous patrol

Warden Edge is a concept for continuous perimeter coverage at commercial sites, where the machine never takes an action on its own.

The coverage gap

Commercial sites have areas fixed cameras cannot reach — back lots, far fence lines, unwired outdoor storage, loading docks. Coverage is bounded by where cable runs. What cannot be wired is not watched, and what is not watched leaves no record.

The approach

Intelligence runs on a fixed edge-AI box on site. A low-cost patrol dog acts purely as an actuator, carrying a camera into the gap. Every event runs through the same loop, and the loop stops at a human before anything happens.

Why governed matters

An unsupervised machine making decisions on a property is an exposure, not a safeguard. Requiring human authorization before any response is the point of the design — it is what makes the resulting record meaningful, and what the value actually rests on.

01DetectThe patrol dog flags an event on its patrol
02Cross-verifyONVIF events from the fixed cameras answer whether the detection is real
03HoldThe reasoning tier answers whether it warrants waking someone; the system stops here and takes no action
04AuthorizeThe notification carries what was seen, why it warrants a response, and a recommended action with its reasoning. The operator approves, declines, or chooses a different authorized action
05Execute + confirmThe authorized deterrent is dispatched — the patrol dog plays an audio message through the audio module on its expansion port and reports back that the track played. It holds position; it does not approach
06RecordA timestamped authorization record captures the whole chain, including the execution confirmation

Hold and ask

The system never acts autonomously. Detection escalates to a person; nothing executes without an explicit human decision.

From detection to authorization, the patroller does not move. No command reaches the locomotion controller while an event is open.

🏠

Runs on site

Inference happens on the local edge box. No cloud dependency for the perception or decision path, and video need not leave the premises.

👤

Role-based authorization

Authorization reaches the operator wherever they are, as a notification on their phone. Who approved, in what role, and when are treated as part of the event itself — not as logging bolted on afterwards.

📋

The authorization record is the output

Every event produces a timestamped authorization record capturing what was detected, how it was corroborated, who authorized the response, and confirmation that the authorized action completed. The record shows the action happened, not just that it was ordered.

Cameras watch. Guards visit. Autonomous robots act on their own.

Warden Edge holds and asks — and proves it.

Each of those alternatives does something well. A fixed camera is the cheapest way to see one point continuously; a guard brings judgement a machine does not have; a robot covers ground without staffing it. The difference is what happens in the moment something is found.

Versus fixed cameras

Coverage that is not bounded by cable

  • Reaches the gaps cameras cannot — back lots, far fence lines, unwired zones.
  • Resolves an event while it is happening, rather than leaving footage to review afterwards.
  • Can respond during the event with an authorized audio deterrent, not only record it.
Versus guards and mobile patrol

Continuous, not periodic

  • Continuous coverage instead of drive-throughs, so there is no gap between rounds.
  • A fraction of ongoing patrol cost, displacing a recurring contract rather than adding to it — intended to pay back in months.
  • No staffing or turnover, and no person placed into a confrontation.
Versus autonomous security robots

It never acts on its own

  • Every response is authorized by a human first — no unsupervised machine making decisions on your property.
  • Produces an authorization record: who approved, when, in what role, with confirmation it executed. A liability instrument, not an activity log.
  • Contained response only — an audio deterrent from a held position, no approach or pursuit, so none of the escalation risk.
And beyond all three

Governed from anywhere

  • Authorization and the authorized command both travel a remote-capable link, so the person approving — and the command that follows — can be anywhere.
  • One operator, or a central monitoring centre, can govern patrols across multiple sites without being on any of them.
  • Approval works after hours, when nobody is on site — which is when a perimeter matters most.

These are the concept’s intended advantages, not results from deployments. Warden Edge is pre-launch: nothing here is measured from a running system, and no customer outcomes are claimed.

The governance loop, running

A single event, end to end — from the patrol noticing something to the authorization record being written.

The scenario: Mid-afternoon at a busy distribution yard. The patrol dog spots someone loading goods from a restricted area into a private car. By day that could be anyone — but four seconds of video tells the story: items carried out one at a time, glances around, no uniform, no work order. So the dog doesn’t act on its own. It shows a person what it saw, recommends a response, and waits for approval — then plays an audio warning to leave, never approaching. The whole chain becomes one timestamped record.

One thing worth noting about the link itself: the patroller and the edge-AI box no longer have to share a local network. The unit joins as a WebRTC peer over the same STUN/TURN path that lets the operator authorize from anywhere, so it can attach over a separate segment or a cellular link instead of the box’s LAN. The dog is still walking this site — only its network path is flexible — but that is what would let one box, or one monitoring centre, command units that are not on its LAN. Intended target-platform behaviour for the demonstration.

  1. 01 Detect The patroller, stopped in dwell, flags a person at the restricted storage area — confidence 0.91
  2. 02 Cross-verify The yard-side camera corroborates via ONVIF at 14:32:09 — the detection is real. This answers is it real, and nothing more
  3. 03 Hold The resident model reasons over the buffered four-second segment: the subject carries goods from the storage area to a private car over several trips, pausing to look around between them. No hi-vis, no yard vehicle, and no work order logged for this zone. This answers does it warrant waking someone — and nothing happens next without a person
  4. 04 Authorize A notification reaches the operator's phone — and the on-site console, where one is installed — carrying the behaviour summary and a recommended response, audio deterrent from a held position, with the reasoning for it. The operator approves, declines, or chooses a different authorized action
  5. 05 Execute + confirm Only on approval does the patrol dog play an audio deterrent through the audio module on its expansion port — asking the person to leave the premises — and confirms the track played. It does not approach
  6. 06 Record A timestamped authorization record is written for the whole chain

Where the time goes

ElapsedStepWhat happens
T−3.6→T+0.4Pre-rollRolling buffer already holding the approach
T+0.001 DetectDetector fires — one person, 0.91
T+0.401 DetectFrames agree. Event raised. Buffered 4s segment sent
T+0.502 Cross-verifyONVIF event at 14:32:09 inside window — corroborated
T+0.802 Cross-verifyFull-resolution confirm
T+0.902 Cross-verifyTrack summary: dwell 3.8s, net translation ~0.4m
T+1.0–4.0ReasoningResident model reasons over the 4s segment
T+4.103 Hold / 04 AuthorizeEscalate. Gate opens; notification fires
T+14.504 AuthorizeApproved — role, operator, timestamp, surface
T+14.7–20.205 ExecuteDeterrent plays from held position
T+20.205 ConfirmUnit reports the track played
T+24.006 RecordRecord signed and closed

The four-second segment costs about 1.3 seconds — all of it absorbed by the ten-plus seconds a human takes to reach for a phone. Illustrative timings, consistent with the example authorization record below; representative of the intended design, not measurements from a deployed system.

Local network · on-site LAN
Perimeter Camera Fixed viewpoint · perception in
Edge-AI Box Central intelligence · perception, reasoning, governance
Entry written when the event completes
Authorization record saved
2026-08-03 14:32:35
Detected → corroborated → held
Authorized by operator, on phone
Audio deterrent played · no approach
Timestamped record
Outside the local network
Warden Edge
No pending authorization
Authorization required Person at restricted storage Recommended: audio deterrent, held position
Decline Approve
✓ Approved 14:32:26
Operator's phone Primary path · remote, out-of-network authorization the console is an option this example does not use — the phone answered

Detector: person at restricted storage — confidence 0.91

Yard-side camera agrees — corroborated. Real, but does it warrant a person?

Four-second segment: loading goods into a private car, no work order. Escalating — held for a person

Approved on phone — Site Supervisor · role: supervisor

Deterrent played. Authorization record written.

Illustrative perimeter scenario — detect → cross-verify → hold → human authorization → execute → timestamped record. Scenario content is representative, not a recording of a deployed system.

Step 06 produces an artifact, and the artifact is the point. This is an example of the authorization record the loop above generates — the full chain of causation behind a single event, from what was perceived through to confirmation that the authorized action actually completed.

Example authorization record

Person loading goods · restricted storage area

Zone
Restricted storage — no work order logged
Trace
RYNN-tr-4471
Opened
2026-08-03 14:32:11
Closed
2026-08-03 14:32:35
Authorized · execution confirmed
Loop sequence Segment widths show relative duration only — illustrative, not measured
Detect
Cross-verify
Reason
Classify
Gate
Authorization request
Detection — fast path at the sensor Correlate, reason, classify Governance — the loop pauses here
Perceived
A person at the restricted storage area — ground no fixed camera covers; detected by the patrol-dog camera while stopped in dwell, confirmed over consecutive frames at confidence 0.91
Corroboration
Corroborated — the yard-side camera emitted an ONVIF person event at 14:32:09, 2 seconds before the patroller's detection at 14:32:11 and inside the correlation window. The storage area itself is uncovered; the corroborating source saw the same person nearby, not the same ground. Two independent sources agree that the detection is real — which is all corroboration can establish
Assessment
Behaviour: goods carried from the restricted storage area to a private vehicle across repeated trips, with pauses to scan the surroundings between them. Against site expectation: no work order or scheduled pickup logged for this zone; no site uniform, no yard vehicle present. Escalation reason: removal of goods from a restricted zone with no authorized activity recorded for it. Operator summary, as sent: “Person at the restricted storage area loading goods into a private car, checking around between trips. No site uniform and no work order for this zone.”
Classified
Unexpected presence, not on the pre-registered list; removal of goods from a restricted zone
Governance
Explicit operator authorization required before any action
Recommended action
System recommended: audio deterrent from a held position, drawn from the bounded, pre-authorized action set. The system selects from that fixed set and explains the choice; it does not invent an action, and nothing in the set involves approach or pursuit. Rationale presented: goods being removed from a restricted zone with no work order logged for it; a contained audio response is proportionate to that. Operator adjudication: approved as recommended. The operator could equally have declined, or authorized a different action from the same set — the recommendation is presented for a decision, not applied for one.
Action
held authorized audio deterrent played (expansion-port audio module) execution confirmed No autonomous approach. The authorized response is a fixed, non-confrontational deterrent: the patrol dog held its position and played a recorded message asking the person to leave the premises. It did not navigate toward the subject and has no authority to. The unit reported the track played, and that confirmation is part of this record.
Authorized by
Site Supervisor role: supervisor 2026-08-03 14:32:26 Authorization is attributed to a role, not just a person. A decline is recorded here exactly as an approval is.

Why the confirmation matters: the record shows the action was carried out, not merely ordered. A log bolted onto a robot after the fact can say a command was sent; it cannot show that a human authorized it first and that the authorized action then completed. That gap is the difference between an activity log and a defensible record.

Intended target platform

These are the two components the demonstration is being built toward. They are a target platform for the concept — not a finished, integrated product, and not a system available to buy.

Edge-AI box
Fixed · The intelligence

Edge-AI box

An edge-AI box that enables real-time vision-language model execution on-site. It carries the perception and governance stack: on-device inference, camera ingestion, and the authorization workflow. This is where the reasoning happens and where the authorization record is produced.

12-DOF patrol dog
Mobile · The actuator

12-DOF patrol dog

A deliberately low-cost commodity actuator. Its job is to patrol, to carry a camera into the places fixed cameras cannot reach, and to play an authorized audio deterrent through an audio module on its expansion port. It holds no decision authority of its own — the intelligence stays on the fixed box, and the authorized response is a contained deterrent rather than pursuit.

What the demonstration runs on

The stack behind the demo above: on-device vision-language inference paired with a real-time operator surface, communicating over a local WebRTC link.

Edge-AI box

The on-site compute for perception and reasoning — running locally, with no cloud service in the decision path.

Tiered detection architecture

Cheap perception at the sensor gates expensive reasoning. The system watches continuously at low power and reasons hard only when something happens.

Tier 0 · at the sensor Detection at the camera node

A small detection model runs on the camera node carried by the patroller, so detection happens before any video crosses the network. It does one job: notice that something is worth looking at. Running this rung all night is cheap enough to be practical.

Tier 1 · on the box Correlator and confirmer

The ~2.3 TOPS integrated NPU confirms the detection on a full-resolution frame and merges the patroller's detection events with ONVIF person events from the fixed cameras, correlating across a lookback window to reach a corroborated or uncorroborated outcome before anything heavier runs.

Optional: Tier 1 can also match a detected face against a facility-maintained roster held on the box — off by default, opt-in per site, and a known/not-known signal only, never an action.

The vision-language model on the discrete high-performance NPU is Tier 2. For the demonstration that model is RynnBrain-4B — an open embodied vision-language model from Alibaba DAMO Academy, built on Qwen3-VL-4B. It is purpose-built for egocentric video understanding, meaning spatial and fine-grained reasoning from a moving camera’s viewpoint, which is the situation a patrolling unit’s camera is in. It reasons over a short video segment — roughly four seconds — not a single frame. That matters because a still frame cannot distinguish walking past from loitering, approaching from leaving (someone walking away needs no deterrent at all), or a hand resting on a gate from one testing it. “Perimeter breach” is inherently temporal: it means crossed, which one frame cannot show. Static false positives — a person printed on a truck side, a mannequin, a poster — also disappear under motion. The model is resident and warm at all times: idle means zero requests in flight, not powered down and not unloaded. It is resident precisely because a cold model would stall inside the window where a person is waiting to be told something — resident is about never paying a first-event penalty, not about throughput. The saving is not avoiding a model load. It is avoiding spending inference on nothing: issuing a request for every frame of an empty yard would cost power and silicon budget all night for no information. Gating requests behind detection and correlation is what makes continuous coverage viable on a box this size. This describes the intended target platform architecture.

Neural axis (reference model)
Warden Edge
no equivalent layer
Governance human authorizes · slowest on purpose
Cerebrum reasoning · ~10–20 W continuous
Edge box · VLM tier resident and warm; a request per event, not per frame
Cerebellum coordination · ~1–2 W continuous
Detector tier · patrol route at the sensor, on the patroller — gates the tier above
Spinal cord reflex · event-driven, ~1 µJ
Patrol dog MCU · gait, IMU, servos never gated by the layers above
reference model Warden Edge equivalent the layer we add The three reference layers and their power figures are as presented in NXP CEO Rafael Sotomayor’s COMPUTEX 2026 keynote transcript; they describe that model, not measurements of anything here. The Warden Edge column describes intended architecture for the demonstration — no measured figures are claimed for it. The governance row is the layer the reference model has no equivalent for, and it is the one this project exists to add.
🧠
On-device inference

Detection runs at the sensor; correlation and reasoning run on the box. The reasoning tier is exposed over a standard chat-completions REST API, with RynnBrain-4B — the third-party open embodied model named above, built on Qwen3-VL-4B — running locally and no internet connection in the decision path. Open weights are what make the on-premises argument hold: the model is a file on the box, not a service call.

📹
Multi-camera ingestion

Several camera streams processed together — fixed perimeter cameras alongside the camera carried by the patrol dog. Corroboration comes from ONVIF person events emitted by the fixed cameras, correlated against the patroller’s detection across a lookback window.

📡
Local WebRTC

Real-time video, audio, and data channels between the operator console, cameras, and patrol dog. Sub-second latency, carried on the local network.

🤖
Agent orchestration

Multi-agent workflows built on an agent framework, with the authorization step modelled as an explicit gate that the workflow cannot proceed past on its own.

🔧
MCP tool integration

Capabilities extended through Model Context Protocol tools, so site systems and data stores can be connected without changes to the core reasoning loop.

🎯
Domain fine-tuning

Open-source vision-language models can be fine-tuned on domain-specific data, so detection and description reflect what actually matters at a given kind of site.

🐳
Containerized deployment

Components ship as containers behind a gateway service, which keeps the demonstration environment reproducible.

🛡️
Model guardrails

Constraints on model output, kept separate from the authorization gate — guardrails shape what the system says, human authorization governs what it does. Qwen3Guard-Gen-0.6B runs on the edge box alongside the reasoning tier for this.

What it guards is narrower than chat safety. It constrains what goes into the operator notification and the authorization record, because a description of a person carries real consequences once it is written into a record an insurer or a legal team may read. Its job is keeping protected-attribute inference — apparent race, ethnicity, gender, age — out of that record. “Person carrying goods from the restricted storage area” is operational; a demographic description is a different kind of document entirely.

This shapes description, not action. It is not a substitute for the gate: the loop still stops and waits for a human either way.

Edge-AI box

Patrol dog

The mobile actuator that carries a camera into the gap and plays an authorized audio deterrent from a held position. The specifications below describe the intended actuator platform for the demonstration — a deliberately low-cost commodity unit that holds no decision authority of its own.

🦿
12-DOF articulation

Three joints per leg across four legs, driven through multi-link connecting rods with inverse kinematics for greater effective torque. Enough articulation to cross uneven ground and steps, and to keep the camera steady as a moving viewpoint.

⚙️
Actuation

Twelve serial-bus servos with real-time position, speed and voltage feedback — roughly 2.3 kg·cm nominal torque, up to about 5.2 kg·cm locked-rotor. Aluminium-alloy and cast-nylon structure running on 40 bearing joints.

🧭
Onboard sensing

A 9-axis IMU — accelerometer, gyroscope and magnetometer — drives self-balancing and gait stability. An onboard voltage and current monitor tracks power draw in real time.

📷
Camera

A 5-megapixel ultra-wide camera, roughly 160° field of view. This is the mobile viewpoint in the loop — the one whose report gets cross-verified against a fixed perimeter camera before anything escalates. Visual detection and reasoning run on this daylight RGB camera, so the perception loop as demonstrated is daylight-dependent. Night operation is a supported extension by adding infrared or thermal imaging to the camera payload, and is not part of the current demonstration.

🧮
Onboard compute

A dual-core microcontroller runs the real-time loop for inverse kinematics and gait generation. An optional single-board Linux host can be added on top for on-unit vision work.

📡
Connectivity

The unit joins as a full WebRTC peer: one connection carries its camera video and, on that connection’s DataChannel, the authorized actuation commands. Because the DataChannel uses the same STUN/TURN traversal the operator’s phone already relies on, an authorized command can be dispatched from outside the local network as well as from the on-site box — which is what makes remote-monitoring-centre and multi-site operation possible. The DataChannel gives reliable, ordered delivery with connection state built in, so dispatch and the execution confirmation return on the same synchronized channel rather than being bridged back from a separate protocol. Control rides the same WebRTC DataChannel transport this project’s architecture is built around. In this architecture the decision path stays on the edge-AI box — the unit receives authorized commands, it does not originate them.

🔋
Power and runtime

Two lithium-ion cells in series, around 7.4 V nominal and 5200 mAh, with over-charge, over-discharge, over-current and short-circuit protection. Roughly one to two hours of continuous operation, or twenty to thirty minutes under sustained high load. Operates while charging.

🐾
Mobility and gait

Gait generation runs on the controller with IMU-driven self-balancing. Movement sequences can be recorded and played back as task files, which is what makes a repeatable patrol route possible.

🔆
Pre-roll buffer

The camera node keeps a rolling buffer of recent low-resolution frames, so when a detection fires the four-second segment is already complete and includes the approach — the most informative part, and the part a post-trigger-only system throws away. Without pre-roll, a four-second segment means waiting four seconds after detection; with it, the segment costs no added wait.

📍
Route recording

The patrol route is established by walking the unit through it once. Waypoints are captured as RTK coordinates as it goes, and every subsequent pass replays the route against those coordinates. This is deliberately not a mapping session: no SLAM, no path planner, no engineer on site. When the yard layout changes, re-recording the route is walking it again.

Two consequences follow, and they are the point. Install stays simple enough for a security dealer — a deployment model that needs a robotics engineer per site does not scale. And a pre-registered route is a governance property rather than a limitation: a site manager can be told exactly where the unit walks, and a record can name the waypoint an event occurred at. A robot that plans its own path goes wherever the planner decides, which is the unsupervised machine making decisions on your property that this concept exists to avoid.

🛰️
Positioning

A GNSS receiver on the unit takes corrections from a base station on the site itself, giving centimetre-class position. Corrections travel over the local network — no internet dependency and no subscription service, consistent with the rest of the decision path running on-premises. The receiver is a small addition on the unit’s expansion port, not part of the base commodity platform.

RTK needs sky view. A perimeter running beside a building wall or between stacked containers will drop from a fixed solution to a degraded one, and that is normal rather than exceptional. Position confidence is recorded with the event, so a degraded fix is visible in the record rather than silent.

Precise position is what lets the record state which zone was breached, and how far the unit was from the subject, as measurements rather than assertions.

Obstacle handling

Roughly five ranging units are fanned across the front and daisy-chained on a single CAN bus. Each carries an embedded ranging MCU, reports a distance, and is given its own ID. Unlike the visual camera, the ranging sensors work in complete darkness — obstacle sensing does not depend on light.

The outdoor-capable units have a narrow beam, and rather than treat that as a limitation the design turns on it: each narrow beam is a distinct angular bin, so the fan produces a multi-point horizontal scan. Presence, range, and which side — from one bus, with no moving parts and nothing spinning to wear out. One unit in the chain can be angled down and forward, where a sudden increase in range indicates a drop-off or a dock edge. Same bus, one more ID. Like the GNSS receiver, these are additions on the expansion bus, not part of the base commodity platform.

The response is a fixed maneuver, not a planner:

  1. 1Stop~0.6 m short of the obstruction
  2. 2Reverse~0.4 m
  3. 3Offset~0.75 m laterally, away from the fence line
  4. 4Run parallel~2 m alongside the route
  5. 5Rejoinback to the route line, and continue

Each leg is verified against RTK rather than dead-reckoned. If the route is still blocked, the offset widens once to ~1.5 m and retries. That corridor is a hard limit. All distances here are design parameters for the demonstration, not measured results.

This is not navigation. No map, no costmap, no path search, no SLAM. One canned maneuver with a single variable, bounded by a corridor fixed at install, and every leg provable from the position log. It is not a navigation stack and is not intended to become one.

When it fails. Two attempts is the whole budget. The unit then skips to the next waypoint and records the gap — segment 6→7 blocked, waypoint 7 not scanned this pass. That is deliberate. An obstacle is a coverage gap, and coverage gaps are what this concept exists to close. A robot that silently routes around a problem and never mentions it has given the operator nothing; one that reports which ground went uncovered has given them something a patrol contractor does not.

Entanglement is separate and does not use the maneuver. Sustained servo current with no translation means netting, cable or shrink wrap: stop immediately, no retry, and report the unit immobilised with its position. Continuing to walk makes entanglement worse.

A blocked route can also be handed to the reasoning tier, since the segment buffer and the model are already there — “route blocked by what appears to be a stacked pallet; this position was clear on the previous pass.”

Route deviation happens only while patrolling. Once a detection is raised the unit is already stopped and stays stopped — no route logic runs while an event is open.

Payload interconnect A payload node MCU at the centre connects to five blocks. A CAN transceiver to its left carries CTX to D and R to CRX plus ground, and drives a cascade of five time-of-flight ranging units over a CANH and CANL differential pair with no defined direction. An RTK GNSS receiver above right is wired in both directions: transmit to GNSS receive carrying corrections in, and GNSS transmit to receive carrying position out, plus ground. An I2S amplifier and speaker to the right take bit clock, word clock and data in, sharing the ground net. A locomotion controller below is wired transmit to receive and receive to transmit, with a heavier common ground because the two boards sit on separate regulators. A two-cell battery pack feeds a five volt buck converter directly over VBAT and ground. Five volt and ground are drawn as net label stubs on each block rather than as distribution wiring. RTK GNSS receiver 5 V · GND ToF cascade ×5 5 V · GND CAN transceiver 5 V · GND Payload node MCU 5 V · GND I2S amplifier + speaker 5 V · GND Locomotion controller own regulator 2S battery pack VBAT 5 V buck converter CANH CANL CTX → D R → CRX GND TX → GNSS RX · corrections in GNSS TX → RX · position out GND BCLK LRCLK DIN shares the GND net TX → RX RX ← TX GND · common reference VBAT GND 5 V and GND are shown as net labels on each block, not drawn as distribution. The buck is fed from the pack directly, not through the expansion header.
Intended payload interconnect for the demonstration. Power distribution is shown as net labels rather than wired out, so the signal paths stay legible. This describes design intent — it is not a built, wired or validated board.
🔆
Status output and expansion

An onboard OLED status display, RGB indicators and an audible buzzer for status tones. Spoken deterrent audio is not this buzzer: it comes from an audio module fitted to the multi-function expansion port, which also exposes spare I/O, serial lines and power for additional sensing.

12-DOF patrol dog
Illustrative console mockup — intended target hardware, not a shipped product.

Operator console (optional)

Optional. An on-site surface where an event can be presented, held, and either authorized or declined. It is not required for the loop to function — authorization reaches the operator's phone regardless, which is what makes approval possible when nobody is on site. The console is a second surface for sites that want one, not a dependency.

One of two authorization surfaces

An escalated event is pushed to the operator's phone and, where a console is installed, presented here too. Either surface can approve or decline; the loop waits for whichever answers first. Declining is as much a recorded outcome as approving.

👤
On-device operator verification

Face recognition running on the console itself establishes who is at the panel, so an authorization can be attributed to a person and a role. Processed on the device.

⏱️
RTOS foundation

A real-time operating system gives consistent, predictable response times — useful when the interface is the thing standing between a detection and an action.

🎥
Video capabilities

Hardware video encoding for WebRTC streaming, so camera feeds and the patrol dog's view can be reviewed live before a decision is made.

Dual RISC-V MCU

Dual-core RISC-V running at up to 400 MHz, built for HMI workloads with rich I/O rather than general-purpose compute.

🔌
Expansion I/O

Available I/O for additional sensing — 60/77 GHz radar for 3D person sensing, additional cameras, or bus protocols such as RS-485.

Demonstration architecture The fixed perimeter camera and the mobile patrol dog both feed perception into a central edge-AI box, which runs vision-language reasoning and the governance logic. The edge-AI box and the patrol dog share one WebRTC peer connection: the dog's video arrives on it and authorized actuation commands are dispatched back on that connection's DataChannel. Because that DataChannel uses the same STUN/TURN traversal as the operator's phone, an authorized command can be dispatched from outside the local network as well as from the on-site box. For any response it first requests authorization from a human — primarily as a notification to the operator's phone, which works remotely and out of network, and optionally through an on-site operator console. The approved or declined decision returns to the edge-AI box, which writes a timestamped authorization record. Perimeter Camera Fixed viewpoint Perception input Patrol Dog Mobile viewpoint Perception input Commanded actuator Edge-AI Box Central intelligence · on site Perception Vision-language reasoning Governance & authorization Operator's Phone Primary authorization path Remote · out of network Approve or decline Operator Console (optional) On-site surface Authorization Record Timestamped record of the whole chain of causation video in actuation · DataChannel in-band, one link remote via STUN/TURN authorization request approve / decline writes on completion
Demonstration architecture — the fixed perimeter camera and the mobile patrol dog patroller are both perception inputs to the edge-AI box, which is where reasoning and governance run. Actuation commands reach the patrol dog in-band, on the DataChannel of the same peer connection that carries its video, and can be dispatched remotely over STUN/TURN rather than only from the on-site box. Authorization goes to a human before any response, primarily as a notification to the operator's phone, with the on-site console as an optional surface. This is the target architecture for the demonstration, not a deployed system.

Where this concept could apply

Site types with large outdoor footprints and coverage gaps fixed cameras cannot close. These are candidate environments for the concept — none of them are deployments.

🚚

Distribution & logistics yards

Large outdoor footprints with trailer parking, gates, and fence lines well beyond the reach of wired camera coverage.

🔐

Self-storage sites

Long drive aisles and outdoor unit rows where continuous coverage is difficult and after-hours activity typically goes unwitnessed.

🏗️

Equipment rental lots

High-value assets parked outdoors across an open lot, with perimeters longer than the camera runs that serve them.

🚗

Vehicle lots

Open lots where inventory sits outside overnight and the back rows are the least observed part of the site.

📦

Warehousing

Loading docks, side yards, and outdoor storage — the areas around a building rather than inside it.

Frequently asked questions

Common questions about the concept and the demonstration.

No. Warden Edge is pre-launch. What this page shows is a concept and a technical demonstration of the underlying loop. There is no shipping product, no deployment, no pilot programme, and no results to report. Anything on this page describing site types or hardware describes intent, not something already delivered.

No — that is the whole point of the design. The patrol dog is an actuator. It patrols, carries a camera, and plays a response that a person has already authorized. Detection and reasoning happen on the fixed edge-AI box, and the loop deliberately stops and waits while an authorization notification goes to the operator's phone.

From detection to authorization, the patroller does not move. No command reaches the locomotion controller while an event is open. The patroller detects while already stopped in dwell, so there is no motion of any kind to interrupt or reverse.

The authorized response is deliberately contained: an audio deterrent played through the audio module on the unit's expansion port, asking the person to leave the premises, from wherever the unit already is. It does not navigate toward the subject, and it has no authority to. Speak, do not chase — the restraint is part of the design, not a limitation of it.

The distinction matters because an unsupervised machine taking action on a commercial property creates exposure rather than removing it.

Yes, and that needs stating plainly. If the model can escalate, it can also decline to, and that is a machine making a decision. The distinction that matters: the model decides what reaches a human’s attention. It never decides what happens. Triage, not authorization. No action is ever taken without a person — that claim is unchanged.

Three constraints bound the triage:

1. Suppression is recorded. A non-escalation writes a record entry with the segment reference and the reasoning. This is not a loss, it is due-diligence evidence. An operator who can produce “at 01:40 a person was observed in zone C, assessed as site staff on the scheduled dock round, not escalated” is in a stronger position than one whose system silently saw nothing.

2. Suppression is bounded. Above a detector-confidence threshold, or with corroboration present in a restricted zone, escalation is unconditional and the model’s assessment does not apply. Judgment is exercised only inside a band the operator defines.

3. Suppression fails open. Timeout, malformed output, or a hedged answer all escalate. A slightly under-informed notification costs ten seconds; a suppressed real event costs the thing this product is for.

It recommends. The notification carries what was observed, why it warrants a response, and a recommended action with the reasoning behind it — so the operator is making an informed decision rather than interpreting a raw detection at two in the morning.

The system recommends; the person decides. The operator adjudicates the recommendation: approve it, decline it, or authorize a different action. Nothing executes without that decision, exactly as before — presenting a recommendation does not move the gate, it only means the human is better informed when they reach it.

Recommendations are bounded. They are drawn only from the same pre-authorized, non-confrontational action set the system could ever execute — an audio deterrent from a held position, no approach, no pursuit. The system selects from that fixed set and explains why; it does not compose new actions. That keeps this consistent with triage: the model shapes what reaches a person and what it suggests, never what happens.

A bounded sidestep, tried twice. The unit stops short of the obstruction, reverses, offsets laterally away from the fence line, runs parallel to the route and rejoins it. If it is still blocked the offset widens once and retries. That is the whole budget.

After the second attempt the unit skips to the next waypoint and records the gap — which segment was blocked, and which waypoint went unscanned on that pass. An obstacle is a coverage gap, and coverage gaps are the thing this concept exists to close, so reporting the uncovered ground is the point rather than an admission.

This is not a navigation stack. There is no map, no path search and no SLAM — one fixed maneuver with a single variable, inside a corridor set at install. And it never applies while an event is open: once a detection is raised the unit is already stopped and stays stopped.

Corroboration comes from ONVIF person events emitted by the fixed cameras the site already runs. When the patroller raises a detection, the box correlates it against ONVIF events inside a lookback window and reaches one of two named outcomes.

Corroborated — the event falls inside fixed-camera coverage and an ONVIF person event agrees within the window. Two independent sources.

Uncorroborated, high confidence — the event is in a blind spot, which is the case the patroller exists for, so there is no second source to correlate against. A higher detection-confidence bar carries it alone. This is the honest limit of the approach: where the patrol is most valuable is exactly where corroboration is unavailable.

Both outcomes escalate to a human. Which one applies is stated in the notification and written into the record, and an uncorroborated event is never described as verified.

Corroboration and reasoning answer different questions. Corroboration answers is this real — it kills phantom detections. It does nothing about a person who is genuinely there and genuinely should be: a contractor, or scheduled staff on a legitimate task. There the detector is right, both sources agree, and the event still should not wake anyone. Deciding that is the video model’s job, not corroboration’s.

Corroboration uses the cameras already on site — closing the gap needs no additional fixed hardware. The coverage audit establishes which of them qualify: a camera has to emit classified person events to corroborate. Cameras that emit motion events only are non-corroborating, and their zones join the uncorroborated class. Motion on a windy night is noise, and counting it as agreement would make false positives worse rather than better — so the audit maps person-event coverage, not merely camera coverage.

The intent either way is to reduce the number of times an authorization notification lands on someone's phone for something that isn't real, so that an escalation carries weight.

The perception and decision path is designed to run on-premises. Inference happens on the local edge box, and the console, cameras, and box communicate over the local network — so video need not leave the site for the loop to function.

Connectivity would be optional and relevant to things like remote access or software updates, not to detection or authorization.

Three reasons. Continuous video against a cloud API means continuous per-query cost, which scales badly for something meant to watch a perimeter all night. Round-trip latency sits in a loop that may need to respond quickly. And continuous outdoor video from a commercial site is exactly the kind of data an operator has good reason to keep on their own premises.

Running the model on the edge box addresses all three at once.

They are optimized for different jobs. The console handles display, touch, camera, audio, and real-time interaction on an RTOS, where predictable sub-millisecond response matters. The edge box handles compute-intensive work — running vision-language models, reasoning over multiple camera streams — on ARM64 Linux with NPU acceleration.

Keeping heavy compute off the interface is what lets the interface stay responsive.

Yes, in principle. Open-source vision-language models can be fine-tuned on domain-specific data, and retrieval over site-specific context can shape how events are described. A vehicle lot and a self-storage site care about different things, and the descriptions attached to a record should reflect that.

Questions about the concept

If you work with commercial sites, build in this space, or just want to understand the approach, get in touch.

Pre-launch — this is a concept and a demonstration
Technical questions welcome

Your information is kept private and never shared.