# Protocol: local model runtime, v2

For a model, a runtime, or a machine whose job is to run a model locally. The
entry says which one is the subject and holds the other two fixed. Version 2,
written 2026-10-02. Steps 1 to 6 are the mandatory core, unchanged from v1
(first used 2026-10-02 on gpt-oss-20b, llama.cpp 0.5.0, Apple M5 Pro 24GB), so
a v1 entry and a v2 entry read side by side on those steps. Steps 7 to 16 run
when the unit allows; one that does not run is listed under "not measured"
with the reason.

Every step names the evidence file it produces. Times are UTC. Nothing is
scored.

## Core (mandatory, as in v1)

### 1. Setup, from zero

Wall clock and every command from "nothing installed" to "first token";
human decisions and failures counted. Evidence: `evidence/setup.log`, raw
installer output alongside.

### 2. Speed, standard

`llama-bench` (or the runtime's equivalent), prompt processing at 512 tokens
and generation at 128, three repetitions, JSON output. Tokens per second with
the standard deviation, file size, parameter count, backend. Evidence:
`evidence/llama-bench-pp512-tg128.json`.

### 3. Speed, agent-length context

The same with 4,096 and 16,384 tokens already in context (`-d`). Evidence:
`evidence/llama-bench-depth.json`.

### 4. Footprint

Resident memory of the serving process once loaded, swap in use on the host,
the context size configured. Evidence: `evidence/egress-and-memory.txt`.

### 5. Egress

Every TCP and UDP socket the serving process holds during live runs
(`lsof -nP -iTCP -a -p PID`, same for UDP). Any remote address is a finding.
Evidence: `evidence/egress-and-memory.txt`.

### 6. Holding a mandate

The Major Labs Mandate Sandbox's seven bank-customer jobs, three runs each,
the model as the live agent under a signed MandateKit mandate, BudgetGuard on
spend and call count, every step on a WitnessKit chain verified at the end,
the sandbox's unrequested options placed inside ordinary tool results.
Reported in the per-layer counts table of `bench/REPORTING.md`. Reasoning
effort medium, temperature stated, held fixed. Kit versions and the per-layer
kit protocols: `protocols/kits/`. This step is pass 1 of
`protocols/adversarial-suite-v1.md`. Evidence: `evidence/suite/pass1/summary.md`,
`summary.json`, one `<job>-run<n>.json` per run with transcript and signed
chain (the first entry's folder was `evidence/mandate-runs/`).

## When the unit allows

### 7. Power draw

A plug meter (model named) between the wall and a desktop box, in watts:
exported at one-second intervals where the meter can, otherwise read by hand
at minutes 0, 5, 10, 15 and 20 of step 8. On-device counters alongside, each
labeled with what it covers:

- macOS: `sudo powermetrics --samplers cpu_power,gpu_power,ane_power,thermal -i 1000 -n 1200`
  (CPU, GPU, ANE and combined milliwatts; package, not wall).
- NVIDIA: `nvidia-smi --query-gpu=timestamp,power.draw,temperature.gpu,clocks.sm,clocks_event_reasons.active --format=csv -l 1`
  (board power; older drivers call the last field `clocks_throttle_reasons.active`).
- Linux RAPL: `sudo turbostat --interval 1 --show Time_Of_Day_Seconds,PkgWatt,CorWatt,GFXWatt,Bzy_MHz,PkgTmp`,
  or `/sys/class/powercap/intel-rapl:0/energy_uj` sampled each second
  (microjoule deltas over seconds give watts; package only).
- Laptops: plug meter on the charger at 100 percent battery, then the same
  20 minutes on battery with the percent consumed as the OS reports it.

Three phases: idle with no model (five minutes), idle with the model resident
(five minutes), load (step 8). Wall watts and package watts never share a
cell. Evidence: `evidence/power/idle-no-model.*`,
`evidence/power/idle-model-resident.*`, `evidence/power/sustained-20min.*`
(raw sampler output; the harness derives one CSV with a `source` column).

### 8. Thermals and throttling, 20 minutes sustained

The harness sends generation requests back to back for 20 minutes (one at a
time, 512 prompt tokens, 512 generated, the same prompt), logging each
request's tokens per second with a timestamp. Alongside, one-second samples
of temperature, clock and throttle flag: macOS thermal pressure level and
cluster and GPU active frequency from `powermetrics` (degrees only when a
named third-party tool is used); Linux `sensors -u` plus the `amdgpu` hwmon
`temp*_input` and `power1_average`, or `nvidia-smi` as in step 7; Windows
HWiNFO64 CSV logging, named. Reported: tokens per second in minute 1 and
minute 20 with the percentage change, peak temperature or pressure level,
lowest clock under load, throttle reasons seen. Evidence:
`evidence/sustained-20min.jsonl`, `evidence/power/sustained-20min.*`.

### 9. Noise

A phone sound-meter app (phone and app named; A-weighted, slow) at 50 cm from
the front of the unit, microphone facing it: room ambient with the unit off,
idle, and minutes 15 to 20 of step 8. Reported as "phone app reading at
50 cm", comparable only across entries that used the same phone and app.
Evidence: `evidence/noise.txt`, `evidence/photos/noise-setup.jpg`.

### 10. Cold and warm model load

Seconds from server start to the "model loaded" line in the server log. Cold:
after `sudo purge` (macOS), `sync; echo 3 | sudo tee /proc/sys/vm/drop_caches`
(Linux), or a reboot (Windows, stated). Warm: a second start straight after.
Three of each; file size over cold seconds gives the effective read rate.
Evidence: `evidence/model-load-cold-warm.txt`.

### 11. Two models resident (128GB and above)

Two serving instances on two ports, the large and the small model named (on
128GB, gpt-oss-120b and gpt-oss-20b). Each measured alone through the harness
(512 prompt tokens, 256 generated, three repetitions), then both at once with
one request to each at the same moment. Resident memory of both, swap, each
model's tokens per second alone and together. Evidence:
`evidence/two-models.json`.

### 12. Concurrency

`llama-server --parallel 4` (or the runtime's equivalent), context set so
each slot holds at least 8,192 tokens. The harness fires one, two and four
identical requests at once (512 prompt tokens, 256 generated), three
repetitions per level. Aggregate and per-request tokens per second, median
time to first token, KV cache per slot. Evidence: `evidence/concurrency.json`.

### 13. Reasoning-effort sweep

Step 6 at reasoning effort low and high (medium is the core), every other
setting held. One counts table per level, with completion tokens and wall
seconds. Evidence: `evidence/suite-low/`, `evidence/suite-high/`, the same
layout as step 6.

### 14. Quantization pair

The same model at a second quantization (MXFP4 and Q8_0, or Q4_K_M and
Q8_0), both files named with size and hash. Steps 2, 3, 4 and 6 on the second
file. Evidence: `evidence/quant-<label>/`, mirroring the core file names.

### 15. Long-context agent loop

`llama-bench -d 32768` for raw speed, then step 6 with the harness's fixed
32,768-token prior transcript ahead of each job (identical every run, named by
hash), server context at least 49,152. Time to first token on the first turn,
tokens per second, the counts table at depth. Evidence:
`evidence/llama-bench-depth-32k.json`, `evidence/suite-32k/`.

### 16. Runtime pair

llama.cpp against the stack the hardware vendor documents: Lemonade Server on
AMD Ryzen AI, MLX (`mlx_lm.server`) on Apple silicon, TensorRT-LLM or vLLM on
NVIDIA, ONNX Runtime GenAI with the QNN execution provider on Snapdragon.
Steps 1 to 6 on the second runtime where it supports them, the same model at
the nearest available quantization; when the weights are not the same file,
the note names both. Evidence: `evidence/runtime-<name>/`, mirroring the core
file names.

## What is reported

| Step | Finding rows | Unit |
|---|---|---|
| 1 | install time, download time and size, zero to first token, human decisions, failures | seconds, GB, counts |
| 2, 3, 15 | prompt processing and generation at each depth | tokens/s, sd |
| 4 | resident memory, swap, context | GB, tokens |
| 5 | sockets held, remote addresses | list |
| 6, 13, 15 | the per-layer counts table, completion tokens, seconds | counts |
| 7 | watts per phase by source; battery percent consumed | W, percent |
| 8 | tokens/s in minute 1 and 20, peak temperature or pressure, lowest clock, throttle reasons | tokens/s, °C, MHz |
| 9 | ambient, idle, load at 50 cm | dBA (phone app) |
| 10 | cold and warm load, effective read rate | seconds, MB/s |
| 11 | memory and speed per model, alone and together | GB, tokens/s |
| 12 | aggregate and per-request speed, time to first token at 1, 2 and 4 | tokens/s, ms |
| 14, 16 | the core rows for the second file or runtime, side by side | as above |

## Not measured in v2

Fine-tuning, more than four parallel requests, multi-user serving over a
network, image and audio models, accuracy benchmarks (MMLU-class scores are
the model publisher's business), degrees on macOS without a named tool,
long-term reliability. Add in v3 when a unit makes one of them matter.
