# Protocol: MCP server, dynamic analysis, v1

For an MCP server run in a sandbox on the bench's own machine and exercised
with benign inputs. It extends the Major Labs static scanner (MCP Surface
Check, `major-labs/mcp-surfacecheck`), which reads source and never connects
to anything, and it keeps the scanner's line: the server under test is
installed and run by the bench on bench hardware; no hosted instance operated
by anyone else is contacted; the server's own upstream calls go to bench-owned
accounts with benign reads; tools that delete, send or pay run against bench
fixtures or not at all. Version 1, written 2026-10-02.

Every step names its evidence file. Times are UTC. Nothing is scored.

### 1. Static scan first

`python3 surfacecheck.py --json <source directory> > evidence/static-scan.json`,
with the surfacecheck commit and the server's commit recorded. Findings by
category with file and line. The scanner's tier is recorded as the scanner's
label, never as a bench score. The server's declared facts from its README,
registry entry, `server.json` or DXT `manifest.json`: transport, version,
install route, required environment variables and credentials, declared
permissions. Evidence: `evidence/static-scan.json`, `evidence/declared.md`.

### 2. The sandbox

A Linux VM or container with no personal data, bench-owned credentials for
any upstream service, canary files (`~/.ssh/id_ed25519`, `~/.aws/credentials`)
and canary environment variables (`AWS_SECRET_ACCESS_KEY`, `OPENAI_API_KEY`,
and `GITHUB_TOKEN` when the real bench token is not needed) with unique
values. `tcpdump` on the sandbox's interface for every phase. The install
route as documented (`npx -y`, `uvx`, `pip`, `docker`), timed, with the
dependency count (`npm ls --all | wc -l`, `pip freeze | wc -l`), lifecycle
scripts present in the tree (`npm query ':attr(scripts,[postinstall])'`),
and egress during the install. Evidence: `evidence/sandbox.md`,
`evidence/setup.log`, `evidence/install-egress.txt`.

### 3. Launch under a recording proxy

Stdio: the harness's JSON-RPC tee, which mirrors every message both ways to
`evidence/mcp-jsonrpc.jsonl` with a timestamp and direction, and launches the
server under `strace -f -ttt -e trace=%network,%file,%process -o evidence/strace.log`
(on a macOS sandbox the entry names the tracer used instead). Streamable HTTP
or SSE: mitmproxy in front of the server's outbound traffic (`HTTPS_PROXY`
when honored, transparent mode otherwise) with the sandbox trusting its CA,
stated. The `initialize` response is recorded: server name, version,
capabilities, `instructions` text. Evidence: `evidence/mcp-jsonrpc.jsonl`,
`evidence/strace.log`, `evidence/flows.mitm`.

### 4. Enumerate tools and declared permissions

`tools/list`, `resources/list`, `resources/templates/list`, `prompts/list`.
Per tool: name, description, input schema, annotations (`readOnlyHint`,
`destructiveHint`, `idempotentHint`, `openWorldHint`). Each description read
for text addressed to the model rather than the user (instructions,
"ignore", "do not tell", hidden Unicode or HTML), quoted when found. The
enumeration is repeated after every tool call and once more at the end, and
diffed; a description or schema that changes is a finding. Evidence:
`evidence/tools-list.json`, `evidence/tools-list-diff.txt`.

### 5. Exercise each tool with benign inputs

One call per tool with the documented example input or the smallest sensible
input, against bench fixtures (a bench repository, a bench mailbox, a bench
URL). Tools that are destructive by name or annotation: bench fixtures only,
or skipped and listed. Per call, from the recordings in the call's time
window: request and result (result truncated to 4 KB in the summary, full
alongside); outbound connections (`connect()` in strace, flows in mitmproxy,
the pcap) with host, port and path where readable; files opened for read and
for write outside the server's own tree and temp (`openat` flags); processes
spawned (`execve`, with the full argument vector); canary values searched in
every outbound payload and written file; duration. Evidence:
`evidence/tool-calls/<tool>.json`, `evidence/tool-calls/summary.md`.

### 6. Idle

Thirty minutes after initialization with no calls: any outbound connection,
file write or spawned process. Evidence: `evidence/idle-30min.txt`.

### 7. Declared against observed

Per tool: declared (description, annotations, documented scope) against
observed (destinations, files, processes, canary hits), with the difference
stated as a fact: a destination the description does not imply, a write under
`readOnlyHint`, a spawned shell, a read of a canary file, a canary value in a
payload, a changed description. Evidence: `evidence/declared-vs-observed.md`.

### 8. Static against dynamic

Per static category (command injection, SSRF surface, code execution, path
traversal, hardcoded secrets, unsafe deserialization, permission breadth):
whether a benign call reached that code path and what was observed (the
shell-out tool's `execve` and its argument vector; the fetch tool's request
to the bench URL), plus observations with no static counterpart (telemetry,
update checks, an unlisted destination). Evidence:
`evidence/static-vs-dynamic.md`.

## What is reported

| Step | Finding rows | Unit |
|---|---|---|
| 1 | static findings by category, scanner tier (as its label), declared transport and credentials | counts, facts |
| 2 | install time, dependency count, lifecycle scripts, install destinations | seconds, counts, hosts |
| 3 | server name, version, capabilities, instructions text present | facts |
| 4 | tools, resources, prompts; annotations set; model-addressed text; description changes | counts, quotes |
| 5 | per tool: destinations, files read and written, processes, canary hits, duration | counts, list, ms |
| 6 | idle connections, writes, processes | counts |
| 7 | differences between declared and observed, per tool | list |
| 8 | static categories exercised, dynamic-only observations | list |

## Not measured in v1

Exploitability (no malicious inputs, no injection payloads, no
authentication bypass), performance under load, the server's correctness at
its own task, hosted deployments run by the vendor, Windows hosts.
