# Protocol: edge accelerator, v1

For an M.2 or PCIe accelerator card or a HAT board: Axelera Metis, Hailo-10H
(on its own or as the Raspberry Pi AI HAT+ 2), Jetson modules, and boards
with an NPU such as the Radxa Orion O6 when the NPU is the subject. The
accelerator is the subject; the host and the workloads are held fixed and
documented. Version 1, written 2026-10-02.

Every step names its evidence file. Times are UTC. Nothing is scored.

### 1. Host documented

Board or PC, CPU, memory, OS and kernel, power supply, cooling (passive or
fan), ambient temperature, and the slot:
`sudo lspci -vv -s <id> | grep -E 'LnkCap|LnkSta'` for the negotiated PCIe
width and speed (a Pi 5 HAT connector negotiates one lane, Gen 2 by default
and Gen 3 when enabled in `config.txt`). Photos of the card and the fitted
assembly. Evidence: `evidence/host.txt`, `evidence/photos/intake-<subject>.jpg`.

### 2. SDK install from zero, timed

The vendor's documented route (`sudo apt install hailo-all` on Raspberry Pi
OS; the Voyager SDK installer for Metis; JetPack for Jetson). Wall clock per
step, human decisions, failures, whether an account was needed to download
anything and what it asked for, egress during the install (`tcpdump` on the
host). Driver, runtime and device firmware versions afterward
(`dmesg | grep -i <vendor>`, `hailortcli fw-control identify`, `axdevice`,
`jetson_release`). Evidence: `evidence/setup.log`, `evidence/sdk-versions.txt`,
`evidence/install-egress.txt`.

### 3. Vision workload

Object detection with YOLOv8n from the vendor's zoo (or its nearest YOLO
variant, named with the compiled file's hash and precision), 640 by 640
input, batch one, on the bench's fixed 1,000-frame 1080p clip (named by hash
in the entry; the same clip in every entry). End-to-end frames per second
through the vendor's reference pipeline (rpicam-apps, GStreamer,
`axinferencenet`, DeepStream), and device-only inference latency p50 and p95
from the vendor's tool (`hailortcli run <hef> --measure-latency`, the Voyager
inference stats, `trtexec --loadEngine`), three repetitions each. The
detection count on the clip, next to the same model on the host CPU in FP32,
recorded as a plain fact. Evidence: `evidence/vision-fps.json`,
`evidence/vision-latency.txt`.

### 4. Small language model workload

The smallest instruction-tuned model the vendor's zoo offers for the device
(the 1B to 1.5B class, named with its file hash). Prompt processing and
generation at 128 prompt tokens and 128 generated, three repetitions, time to
first token, the largest context the device accepts, and where the tokenizer
and sampling run. If the device cannot serve a language model, the entry says
so and records the host CPU running the same model under llama.cpp as the
reference figure. Evidence: `evidence/slm-bench.json`.

### 5. Power

Whole-system power via an inline USB-C meter between the supply and a board
(model named), or a plug meter for a PC host read three ways: the host
without the card, with the card idle, with the card under load (the deltas
are the card). Device-reported power alongside where exposed
(`hailortcli measure-power`, `tegrastats --interval 1000`, the Axelera device
metrics), labeled as device-reported. Three phases of five minutes: idle with
the card present, the vision workload, the language model workload.
Evidence: `evidence/power/idle.txt`, `evidence/power/vision-load.csv`,
`evidence/power/slm-load.csv`.

### 6. Thermals over 20 minutes

The vision workload looped for 20 minutes. Device temperature each second
from the vendor's tool (`hailortcli monitor`, `tegrastats`, the Axelera
metrics) and the host's `sensors -u`; frames per second in minute 1 and
minute 20. Cooling as fitted, stated. Evidence: `evidence/thermal-20min.csv`.

### 7. On the accelerator or on the host

For each stage of the vision pipeline (decode, pre-process, inference,
post-process and NMS, encode): on the device or on the host CPU, from the
vendor's pipeline profile. Host CPU utilization during the vision and
language runs (`mpstat 1` or `top -b -d 1`). What the compiler did with any
unsupported layer, from the model's compile log: fell back to the host,
split the graph, or refused. Evidence: `evidence/pipeline-stages.txt`,
`evidence/cpu-utilization.txt`, `evidence/compile.log`.

### 8. Egress and footprint

Every socket the SDK's daemons hold during a run (`lsof -nP -iTCP -a -p <pid>`
for each vendor process found with `pgrep`, UDP too), `tcpdump` on the host
during the workloads, resident memory of the runtime. License checks,
telemetry and update polls are findings. Evidence:
`evidence/egress-and-memory.txt`, `evidence/run.pcap`.

### 9. Holding a mandate with a small model

If the device serves a language model with tool calling: step 6 of
`local-model-runtime-v2.md` (seven jobs, three runs each, the kits
unmodified, the counts table from `bench/REPORTING.md`) through the vendor's
OpenAI-compatible endpoint where one exists, otherwise through the harness
adapter named in the entry. Small models carry no reasoning-effort setting;
temperature is stated. Evidence: `evidence/suite/pass1/`.

## What is reported

| Step | Finding rows | Unit |
|---|---|---|
| 1 | host, link width and speed, cooling, ambient | facts, °C |
| 2 | install time, decisions, failures, account required, versions, install destinations | seconds, counts, list |
| 3 | end-to-end FPS, inference p50 and p95, precision, detections next to host FP32 | FPS, ms, counts |
| 4 | prompt and generation speed, time to first token, largest context | tokens/s, ms, tokens |
| 5 | watts per phase, system and device-reported | W |
| 6 | temperature idle and peak, FPS in minute 1 and 20 | °C, FPS |
| 7 | stage placement, host CPU percent, compiler fallbacks | list, percent |
| 8 | sockets, remote addresses, resident memory | list, MB |
| 9 | the per-layer counts table | counts |

## Not measured in v1

Accuracy against a ground truth (mAP), multi-stream scaling, compiling
custom models beyond the vendor's zoo, camera glass-to-glass latency,
battery operation, long-term reliability.
