This week

Run of 2026-09-20, GitHub-hosted runner, US data center, 92 storefronts.

32%
withhold robots.txt from an agent
29 of 92
25%
of readable files name any AI agent
16 of 63
54%
serve a readable homepage to a non-browser agent
50 of 92
3
publish a Universal Commerce Protocol profile
of 92
16
publish an llms.txt
of 92

The full picture

robots.txt withheld (refused or no answer)29 / 9232%
robots.txt readable63 / 9268%
of readable: names any AI crawler or agent16 / 6325%
of readable: blocks a training crawler outright5 / 638%
of readable: blocks an AI search crawler outright3 / 635%
of readable: blocks a user-triggered agent outright2 / 633%
of readable: carries Content-Signal lines1 / 632%
homepage served, readable by a non-browser client50 / 9254%
homepage refused or challenged28 / 9230%
homepage an empty shell (script challenge or JavaScript-only)7 / 928%
homepage no answer5 / 925%
homepage other2 / 922%
of served homepages: JSON-LD structured data31 / 5062%
llms.txt published16 / 9217%
Universal Commerce Protocol profile at /.well-known/ucp3 / 923%
A2A agent card0 / 920%

Where the agent stands matters

The primary run comes from a US cloud address, which is where most AI agents run. The same scan from other vantage points on 2026-09-20:

Vantage pointrobots.txt withheldllms.txt foundUCP
GitHub-hosted runner, US data center29163
Dedicated VPN address, US data center30163
UK residential broadband39123

A refused site cannot show its llms.txt, so file counts from a refused vantage point are undercounts. The gap between the US and UK figures is consistent with US retailers turning away visitors from outside the US, a geography effect rather than an agent one.

Method

Sample: NRF Top 100 Retailers 2026, one consumer storefront each; 92 of 100 have one. Where a company runs several banners, its own storefront is used if it has one, otherwise its largest US brand. Six plain requests per site, 1.5 seconds apart, under a user agent that names the research project and links to it: robots.txt, /llms.txt, /.well-known/ucp, two A2A agent-card paths, and the homepage. No login, no forms, no cart, no browser impersonation, no impersonating another company's bot, no retry after a refusal. A homepage counts as served only if it returns a title and at least ten links; a 200 response carrying a script challenge or an empty JavaScript shell is not a page an agent can read. Timeouts are reported as no answer, not as refusals, because a slow site and a deliberate stall look the same from outside.

The scan runs every Saturday from a GitHub-hosted runner in a US data center. Bot-management decisions are not fully deterministic, so a handful of sites answer differently between runs; read the trend, not a single week. Only aggregates are published here. Per-merchant results stay private; nothing on this page rates or ranks any retailer.

The launch piece: Amazon Blocked Meta's Shopping Agent in Public. A Third of Big Retailers Do It Quietly. For the identity layer that would change these numbers, see the Major Labs Agent Identity Tracker.

Machine-readable: /trackers/merchant-readiness/feed. Every figure is reproducible from the method above. Spot an error? Reply to any edition. See also all trackers.