This week
Run of 2026-09-20, GitHub-hosted runner, US data center, 92 storefronts.
The full picture
| robots.txt withheld (refused or no answer) | 29 / 92 | 32% |
| robots.txt readable | 63 / 92 | 68% |
| of readable: names any AI crawler or agent | 16 / 63 | 25% |
| of readable: blocks a training crawler outright | 5 / 63 | 8% |
| of readable: blocks an AI search crawler outright | 3 / 63 | 5% |
| of readable: blocks a user-triggered agent outright | 2 / 63 | 3% |
| of readable: carries Content-Signal lines | 1 / 63 | 2% |
| homepage served, readable by a non-browser client | 50 / 92 | 54% |
| homepage refused or challenged | 28 / 92 | 30% |
| homepage an empty shell (script challenge or JavaScript-only) | 7 / 92 | 8% |
| homepage no answer | 5 / 92 | 5% |
| homepage other | 2 / 92 | 2% |
| of served homepages: JSON-LD structured data | 31 / 50 | 62% |
| llms.txt published | 16 / 92 | 17% |
| Universal Commerce Protocol profile at /.well-known/ucp | 3 / 92 | 3% |
| A2A agent card | 0 / 92 | 0% |
Where the agent stands matters
The primary run comes from a US cloud address, which is where most AI agents run. The same scan from other vantage points on 2026-09-20:
| Vantage point | robots.txt withheld | llms.txt found | UCP |
|---|---|---|---|
| GitHub-hosted runner, US data center | 29 | 16 | 3 |
| Dedicated VPN address, US data center | 30 | 16 | 3 |
| UK residential broadband | 39 | 12 | 3 |
A refused site cannot show its llms.txt, so file counts from a refused vantage point are undercounts. The gap between the US and UK figures is consistent with US retailers turning away visitors from outside the US, a geography effect rather than an agent one.
Method
Sample: NRF Top 100 Retailers 2026, one consumer storefront each; 92 of 100 have one. Where a company runs several banners, its own storefront is used if it has one, otherwise its largest US brand. Six plain requests per site, 1.5 seconds apart, under a user agent that names the research project and links to it: robots.txt, /llms.txt, /.well-known/ucp, two A2A agent-card paths, and the homepage. No login, no forms, no cart, no browser impersonation, no impersonating another company's bot, no retry after a refusal. A homepage counts as served only if it returns a title and at least ten links; a 200 response carrying a script challenge or an empty JavaScript shell is not a page an agent can read. Timeouts are reported as no answer, not as refusals, because a slow site and a deliberate stall look the same from outside.
The scan runs every Saturday from a GitHub-hosted runner in a US data center. Bot-management decisions are not fully deterministic, so a handful of sites answer differently between runs; read the trend, not a single week. Only aggregates are published here. Per-merchant results stay private; nothing on this page rates or ranks any retailer.
The launch piece: Amazon Blocked Meta's Shopping Agent in Public. A Third of Big Retailers Do It Quietly. For the identity layer that would change these numbers, see the Major Labs Agent Identity Tracker.
Machine-readable: /trackers/merchant-readiness/feed. Every figure is reproducible from the method above. Spot an error? Reply to any edition. See also all trackers.