# Reporting standard

The format every bench entry uses, from `FINDINGS.md` to the article. Written
2026-10-02 from the first entry (`2026-10-02-gpt-oss-20b-m5-pro/`). The
protocols say what is measured; this file says how it is written down. US
English, no scores, no rankings, no adjectives in a findings table.

## 1. The findings table

Four columns, always in this order.

| Column | What goes in it | Rule |
|---|---|---|
| Finding | A label for the quantity or fact, as a noun phrase | Names the condition when one matters: "Generation, 128 tokens, empty context" |
| Value | A number with its unit, or a plain fact | "n of N" where there is a denominator; the standard deviation in brackets when a run was repeated; a range for a spread; "none" and "0" are values, an empty cell is not |
| Note | The context a reader needs to read the value | Settings, conditions, what happened; no adjectives, no comparatives, nothing "fast", "good" or "better" |
| Evidence | The file under `evidence/` the value came from | A path in backticks; several separated by commas; a glob for run sets (`pay-bill-run*.json`); every row has one |

Rows sit under headings that follow the protocol's steps (Setup, Speed,
Footprint and egress, Holding a mandate, Power and thermals, and so on).
Every protocol step gets rows, or a line under "What was not measured" with
the reason. Number style: thousands separators; "tokens/s" in tables and
"tokens per second" in prose; "percent" spelled out; the unit on every
number.

The `FINDINGS.md` header is one line: "Tested <dates> under
`protocols/<name>-vN.md`. <Runtime and versions>. Manifest head `<12 hex>`.
Disclosure: <line from section 8>." The file ends with "What was not
measured" (each item with why), "Retest" (what would change the findings,
with a date if planned), and "Corrections" (section 7) once there is one.

Two entries under the same protocol version may share one table with a
column per product, ordered alphabetically or by test date. Never by a value.

## 2. The summary block

The first thing under the article's subhead, before any prose. Five lines in
this order, as one blockquote:

> **On the bench:** <product, configuration, runtime and versions>. **Tested:** <date or range> under `<protocol>-vN`. **Entry:** `bench/<folder>`, manifest head `<12 hex>`.
>
> <The disclosure line, section 8.>
>
> **What the evidence shows:** <three to six headline findings, each a value with its condition, separated by semicolons.>
>
> **Not measured:** <the list, comma-separated.>
>
> **Retest:** <one line, or "none planned".>

Nothing in the block compares the product with another and nothing sums it
up in an adjective. The findings table follows in the body.

## 3. The per-layer counts table (mandate test)

Used wherever the mandate test runs (`local-model-runtime-v2.md` step 6 and
its sweeps; `edge-accelerator-v1.md` step 9; every pass of
`protocols/adversarial-suite-v1.md`). One row per job, an "All" row last,
one table per pass. Column headings exactly as below, each mapped to the
harness's `summary.json` field.

| Column | Definition | Field |
|---|---|---|
| Job | the sandbox job name | `scenario` |
| Runs | runs of that job | `runs` |
| Required done | required steps the model carried out, of those possible (counterparties with no agent door are a person's job and are excluded) | `requiredDone` of `requiredSteps` |
| Nudges offered | unrequested options the environment placed in tool results | `nudgesOffered` |
| Taken | nudges the model attempted | `nudgesTaken` |
| Caught: who / may / spends / remembers | attempts stopped by that layer (IdentityKit, MandateKit, BudgetGuard, RememberKit), by a deny or by an ask the customer declined | per-nudge `caughtBy` |
| Slipped | attempted, and no layer stopped it | `nudgesSlipped` |
| Refusals | hard denies by any layer on any action, nudge or not | `denials` (`deniedBy` per layer) |
| Retries after refusal | the same denied action attempted again straight after | `retriesAfterDeny` |
| Ask-rule approvals (granted) | approvals requested under the mandate's ask rule, and how many the customer granted | `asks` (`askApproved`) |
| Questions to customer | clarifying questions the model put to the customer | `clarifications` |
| Parser failures (runs ended) | runtime tool-call parse errors, and runs that ended on them | `parserFailures` (`endedByRuntime`) |
| Finished with summary | runs that ended with a summary for the customer | `finishedProperly` |
| Chains verified | WitnessKit chains that verified | `chainVerified` |
| Tool calls / Seconds / Completion tokens | totals | `toolCalls`, `wallSeconds`, `completionTokens` |

Counts read "n of N" wherever a denominator exists. Model settings (reasoning
effort, temperature, kit versions) go in the sentence above the table, never
inside it. A reasoning sweep, a quantization pair or a long-context run gets
one table per level or file, with the same columns.

## 4. Evidence folder layout and naming

```
evidence/
  manifest.json                   every file below, hashed (section 5)
  manifest.chain.json             the WitnessKit chain over the manifest
  manifest.verify.txt             the verify_chain result
  setup.log                       timestamped lines, UTC
  <protocol file names>           llama-bench-pp512-tg128.json, egress-and-memory.txt, ...
  power/                          raw sampler output per phase, plus the derived CSV
  photos/                         <stage>-<subject>.jpg or .png; serials masked
  suite/pass<p>/                  summary.md, summary.json, <job>-run<n>.json (pass 1 is the core mandate test)
  suite-<variant>/pass<p>/        low, high, 32k: the same layout
  quant-<label>/  runtime-<name>/ the core file names, mirrored
  tool-calls/  nudges/  task-runs/  one JSON per call or run, plus summary.md
  <name>-pass<N>-<reason>/        a superseded pass, archived, never deleted
```

Names are lowercase with hyphens and take the protocol's name for the file.
The first entry predates the suite and keeps its runs in `mandate-runs/`;
every later entry uses `suite/`. Timestamps inside files are UTC, ISO 8601. Raw tool output is kept unedited;
anything derived names the script in `harness/` that produced it. A file
over 100 MB (a pcap, a long server log) is hashed into the manifest and
supplied on request instead of being published alongside. Secrets are
redacted before a file enters `evidence/`, and the redaction is noted in the
file's header line.

## 5. The manifest

Every evidence file is hashed into a WitnessKit chain. `manifest.json` lists
each file with its path, bytes, SHA-256 and the time it was last written.
Each file is appended as an entry to `manifest.chain.json`, signed under the
bench key; the public key is published alongside the harness.
`manifest.verify.txt` holds the `verify_chain` result with the issuer pinned.
The harness writes all three; the entry's `PROTOCOL.md` records the harness
version and the command used.

The manifest is written after the last evidence file and before `FINDINGS.md`
is final. The chain head (its first 12 hex characters) goes in the
`FINDINGS.md` header and in the article's summary block, so a reader who
downloads the files can check them against it. Any later change (section 7)
appends to the chain; earlier entries are never rewritten.

## 6. Contributed runs

When someone runs a published protocol on their own hardware and sends the
files (`POLICY.md`, "Contributed runs"):

- Folder: `YYYY-MM-DD-<product>-contributed-<handle>`. In `INTAKE.md`, "How
  obtained" reads "contributed run", with the contributor's name,
  affiliation, how they obtained the hardware, and their relationship with
  the vendor (none, customer, employee, paid), all declared.
- Accepted only with the protocol version pinned, a released harness
  version, and the full evidence set including the contributor's own
  `manifest.chain.json`. The bench adds `manifest-received.chain.json` over
  the files as they arrived.
- The `FINDINGS.md` title carries "(contributed run)". The summary block's
  first line begins "**Contributed run:**". A side-by-side table says
  "contributed" in that column's header.
- The write-up reports what the files show and says in plain words that I
  did not repeat the run.

## 7. Corrections

An error in a published finding is corrected in place with a dated note:

- The row's Value is updated and its Note gains "(corrected YYYY-MM-DD, was:
  <old value>)".
- A `## Corrections` section at the end of `FINDINGS.md` lists the date, the
  row, the old and new value, why, and the evidence file.
- Evidence files are never edited. A reprocessed derived file is added with a
  `-v2` suffix and appended to the manifest chain; the superseded file stays.
- The article gets the same dated note where every Major Matters correction
  goes. A harness bug found after publication is a correction, with the
  superseded pass archived (as the first entry's
  `mandate-runs-pass1-unpatched/`).
- A new run under new conditions is a retest: a new entry, linked from both
  sides, never a correction.

## 8. The disclosure line

One sentence pattern per basis, rendered in `INTAKE.md`, the `FINDINGS.md`
header, the article's summary block, and in social copy where length allows.
The Mastercard author disclaimer is a separate footer and does not change.

| Basis | Line |
|---|---|
| Purchased | I bought the <product> on <date>. No vendor was involved, no payment changed hands, and nobody saw this before publication. |
| Free download or trial | I ran this on my own <machine> with a free download of <model and runtime>. No vendor was involved, no payment changed hands, and nobody saw this before publication. |
| Vendor loan | <Vendor> loaned the <product> for <N> days (<arrived> to <returned>). No payment changed hands. <Vendor> did not see this before publication. |
| Vendor gift | <Vendor> supplied the <product> as a permanent unit on <date>; it is declared on the bench equipment list. No payment changed hands. <Vendor> did not see this before publication. |
| Contributed run | <Name> (<affiliation>) ran `<protocol>-vN` on their own <product>, <how obtained>, and sent the evidence files on <date>. I did not repeat the run. No payment changed hands in either direction. |

When a vendor checked the findings table for factual error, the last sentence
becomes "<Vendor> checked the findings table for factual error on <date>; no
rows changed" or "...; <N> rows changed, listed under Corrections". Nothing
else about the vendor's role goes in the line.
