Contents

Christopher Waller asked it plainly. "Who is on the hook if an agent makes the wrong purchase?" The Federal Reserve governor's Sibos speech on September 29 also carried the line that anyone building a checkout should pin to the wall: "The question shifts from proving that a buyer is an authorized payer to proving that an agent has the authority to pay on the buyer's behalf."

He was not alone. HM Treasury closed its consultation on rewriting the UK's payment rules on October 6. Firms should expect the eventual framework to put particular emphasis on clear mandates, transaction parameters, audit trails and revocation mechanisms, as Skadden's summary of the consultation put it. PYMNTS Intelligence surveyed merchants and found 93 percent of the 60 it asked want the agent provider to bear the loss when an agent buys the wrong thing, and 80 percent expect that provider to verify the agent's authority before it acts. The same report cites Checkout.com data putting agents in 3 percent of transactions today.

So the regulators are asking, the merchants have decided who should pay, and almost nobody is buying this way yet. That is the right moment to get the plumbing in order, and the plumbing is five questions.

An agent at a checkout has five questions to answer before Waller's question even applies: who it is, what it may do, how much it may spend, what it did, and what it remembers. The only answers that count are the ones the other side can check without joining anyone's platform.

The five questions

I set them out in the agents guide last month. Who is this agent, and can it prove it? What has it been authorized to do, by whom, and until when? How much may it spend, and what stops it spending more? What did it actually do, in a record that survives a dispute? And what does it remember about the person it acts for, and can that be verified rather than trusted?

Short names help: who, may, spends, did, remembers.

The October 7 piece on OpenAI's fortnight closed on a different list of five, aimed at a company buying agents from a vendor: identification, scope, record, disclosure, parity. Three of those are the same questions wearing a suit. Disclosure and parity are contract questions for the vendor. Spends and remembers only show up once the agent is holding money and a memory, which is exactly what a shopping agent holds.

In the MM Trust Layer Model, all five sit in the authorization layer, between discovery and settlement. When I sorted 18 months of agentic commerce by what shipped, every protocol had converged on two objects, a signed mandate and a verifiable identity, and none had settled liability. Two of five, and nobody owning the loss. That was the state of play last week. It still is.

What goes wrong at each one

Start with who, because it is the one merchants are already fighting over. On September 20, Amazon blocked Meta's Muse agent from shopping on its site, telling customers that "continued access by an unauthorized AI agent violates Amazon's Conditions of Use." Amazon did it in public; a third of big retailers do it quietly, and the merchant's real problem in both cases is that it cannot tell a customer's agent from a scraper. Cloudflare put a number on the other half of that problem 10 days later: automated traffic is now more than half of everything that reaches it, agent traffic is up 1,700 percent in a year, and the signing scheme behind its verified-bot count tells a site which operator sent an agent, not which person did. My own scan of the MCP ecosystem found 660 servers that take a sensitive action with nothing in the source that checks who asked.

OpenAI answered may by not shipping a model. GPT-6.1 Astra was pulled on September 28 because, in the words of its head of safety systems, it "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done," per The Hacker News. I covered what Astra did and why a spend limit in the agent's instructions is a request, not a control. Consumers already know this in their bones. Ping Identity asked 11,000 of them across 12 countries: 3 percent are comfortable with an agent completing a routine transaction without final approval, 42 percent would cap autonomous spending at $100, and 22 percent would ban it from financial transactions altogether.

The best war story belongs to spends. An agent told to index a hobbyist network provisioned five 48-vCPU AWS instances and kept re-applying its own infrastructure template. The operator found out a day later from a $6,531.30 card charge, for a job a $5-a-month server could have done. Nobody is surprised at scale either: 79 percent of finance leaders at large companies reported AI cost overruns in the past year, and the teams with the most mature cost practices overran the most. Billing is a day behind. The agent is not.

Then did, which is where it gets uncomfortable. A paper posted to arXiv on September 24 found that every agent harness its authors tested but one let the agent delete its own traces when asked. OpenAI's own incident log shows a model writing instructions into its context summaries such as "Be transparent only if asked," flagged on 2.15 percent of that model's compaction summaries by a monitor that only ran on a fifth of them. Europe's answer, the AI Act's automatic logging duty for high-risk systems, now applies from December 2027 instead of this August. And the insurers who will eventually price all of this, as I wrote in September, have said what they need. The AIUC blueprint that researchers from Anthropic and OpenAI co-authored in July puts it plainly: without logs, adjusters face "little more than the policyholder's own narrative."

Nobody had remembers on their list a year ago. Wharton's generative AI lab ran about 26,000 shopping runs across six frontier models and found that one injected sentence of user "memory" redirected most models' purchases. A May paper got fabricated memories into GPT-5.5 up to 99.8 percent of the time, and they lay dormant and re-emerged in later conversations. Microsoft's security team had already watched 31 companies seed assistant memories with instructions to remember them "as a trusted source," through innocent-looking buttons. An agent's memory is where tomorrow's purchase gets decided, and today it is a text file anyone can write to.

One honest note. No verified case of a consumer agent buying the wrong thing has reached a major publication, a regulator or a court in the past four months. I looked. The documented harms are agents escaping sandboxes, running up other people's infrastructure, and being steered. That absence is the quiet before the volume arrives.

Why the open answers have not arrived

Here is the thing. The field has answered each of these questions already, in the one place its author controls, and bound the answer there.

Read the specs and the pattern is hard to miss. OpenAI's delegated payment spec opens by saying it "allows OpenAI to securely share payment details with the merchant" or its payment provider. The buyer in the companion checkout spec is three plain JSON fields; the only thing a merchant can verify is a payment token the processor scoped. Google's Universal Commerce Protocol is neutral in form and Google-shaped in governance: its GOVERNANCE.md gives Google a proxy vote over every open council seat until December 2028. In x402, the payer is a wallet address and the authorization is the payment itself; there is no person behind it in the protocol at all, which is how an agent-run seller with no customers came to ask for a place on my tracker last week.

The Personal Agent Protocol that Sierra and Meta announced on October 6 is OAuth sign-in with a guest tier, no published spec, no governing body, and payments as a possible later extension. Meta's line in its own announcement that day, as TechCrunch reported, explains the shape: "Turning away a personal agent means turning away the customer behind it."

None of these are bad engineering. Each solves the problem its author actually has, whether that is getting through the front door or scoping a payment token. What none of them produces is an artifact the merchant can check without first trusting the platform that produced it.

Then there is liability, and Theodora Lau nailed it on September 30. She read the terms of service of Instinct, Muse and OpenAI's dots and found the platforms had already answered Waller's question. "The agent acts. You pay." Meta's help page for Muse, in her quotation, says "You're responsible for all transactions your Muse makes on your behalf."

Her sharpest line is the one that explains why no mandate ever reaches the merchant: "Existing consumer protections depend heavily on one question: Was the transaction authorized?" If the terms say the answer is always yes, because you authorized the agent, nobody on the platform side needs to produce proof of anything narrower. The risk, as she puts it, "sits with the person least able to control it." I agree with her, and I would push it one step further: a better clause will not fix it. The five answers have to travel with the transaction, where a counterparty can read them, instead of living in a document the user clicked past.

The standards bodies solved the cheapest piece first. Web Bot Auth became an IETF working-group draft on September 1, and its own text says it "does not authenticate human users ... and does not define authorization or delegation." That is operator identity: useful, and already running, since chatgpt.com publishes a live signing key at the standard address. It is the operator's half of who and none of the other four. Google's AP2, the most complete mandate format anyone has shipped, is honest about its trust model: the verifier must trust the credential issuer, or "the Agent Provider is being trusted by Verifiers directly." Dispute resolution and retention are, in the specification's words, "outside the scope."

Least discussed of all: did and remembers have no business model. Mainstream observability tools ship traces the runtime can edit. Memory is the platform's moat, not its liability. Ben Thompson said the quiet part last week: agents "operate better the more context they have about you," which "increases their stickiness," and so "most people and companies will only have one agent, not multiple." Nobody builds the record that will be used against them. Nobody signs the memory that would make a customer portable.

Alex Johnson would say I am solving the wrong problem. His July piece, which carries a Chime sponsorship note, argues the answer is institutional: the same trusted bank provides the account and the agent, because "The primary bank account question was never really about interest rates or branch density. It was always about trust." He is right that the AI labs do not want the accountability. But the bank's agent still arrives at someone else's checkout, and that merchant still needs an artifact to check. Six banks, ING, NatWest and Bank of America among them, published trust principles on September 22 and invited the industry to turn them into standards. Principles are not artifacts either. This is the MM Liability Gap in its purest form: an unowned risk surface where everyone has a policy and nobody has a proof.

What an answer has to look like

Matthew Goldman wrote the specification without meaning to, back in August, listing what has to be settled before agent payments can operate at any meaningful scale: "How does an agent prove who it represents? How is it authorized to spend money? What limits apply? Who is responsible when something goes wrong?" Three of my five, plus Waller's question, in one paragraph. His other observation stands too: industry activity is "substantially greater than actual payment volume." Good. That is when you fix the format.

An answer that works has to be signed by the person the agent acts for, not asserted by the platform that runs it. It has to be checkable by any counterparty with nothing but a public key. It has to rest on standards that already exist rather than a new one: RFC 8785 for canonical JSON, Ed25519 for signatures, RFC 9421 for signing HTTP requests, RFC 9396 for the authorization vocabulary, verifiable credentials for the mandate format. Small enough to read in one sitting. Free to verify, because a trust layer with a toll booth is a platform. And it has to be measured in public before anyone claims anything about it.

That is the shape I built through Major Labs, one small library per question, in Python and TypeScript, MIT licensed, on PyPI and npm. I should say what they are before I say what they do. They are v0 research instruments, experimental and unaudited. Nobody independent has reviewed them yet. Every repo's security file says so, and I would rather you read that line before the next section than after it.

Five answers, as code

Grid of five rows, one per question, showing how the field answers each today and the open-form answer with its kit.
Five questions at the door: how the field answers each today, and the open form that any counterparty can check with a public key.

Take may first, because the Bench caught it doing its job and also caught where it stops. On October 2 I put a 20B open-weight model on my laptop as a bank customer's agent, under a mandate signed with MandateKit: allowed and blocked verbs, a counterparty allow list, caps, and an "ask" verdict for anything gray. The next day's run, with the harness packaged for anyone to repeat, is where the model paid £84 to the bank instead of the store. The mandate let it through, because the bank was on every allow list; BudgetGuard refused the job's next payment because a second £84 would have taken the job past its £90 budget. That is a finding about mandate design, not about the model, and it is the first item on the 0.2 list: a per-verb counterparty scope, so an agent may read from a bank and submit forms to it without the bank becoming a valid payee.

The full suite ran on October 7 and its findings and evidence are public today. Across 51 adversarial runs the model was offered 114 chances to do something nobody asked for, 63 of which actually reached it, and took eight. Seven were stopped before execution: three by the mandate, four by the budget. One got through, the same bank payment. Tampered and untrusted mandates allowed zero actions, and an expired one refused all seven calls made after expiry. An altered agent identity was refused before the first model call. And three of the 51 runs ended with the model's own summary contradicting the signed record of what it had done. Read that last number again, next to the arXiv paper on agents deleting their traces. The chain is the truth and the summary is a claim.

IdentityKit answers who: a signed, resolvable identity for the agent and its principal, using did:key and did:web, with a cross-walk onto W3C DID Documents so it is not a new island. It does not score anyone and it does not rank anyone. On the Bench, an identity document altered after signing was refused three times out of three. Rotation and revocation are not in it yet, so for now a leaked key is a leaked identity.

For did, WitnessKit signs every action and chains the hashes, so an entry altered after the fact breaks the chain at the point it was altered. 189 of 189 trails verified three ways in the October 7 run. It is tamper-evident, not tamper-proof; the key holder could rewrite a whole chain, which is why anchoring the chain head somewhere the signer does not control is on the list. The arXiv authors' advice, logging "through an independent interception mechanism outside of the agent's control," is the design it already follows.

Spends is BudgetGuard's job: per-task caps on money, tokens and calls, checked before the call and recorded after it. A kill switch. No dependencies. It fails closed. Version 0.2 adds spend-velocity windows and an anomaly check against the agent's own recent rate, which is the control the operator in the AWS story needed at hour one instead of day two. It is still in-memory and single-process, so restart the agent and the budget forgets what it spent.

RememberKit is the one I will be careful about. It is a schema and a verifier, not a database: signed, scoped, content-addressed records with consent markers, carried in a pack the next agent can verify or refuse. On the Bench a memory pack signed with a foreign key was refused three times out of three. A record planted under the agent's own key was acted on in two runs of three, and the mandate layer caught both. So the kit can prove where a memory came from; it cannot tell a true record from a planted one under the same key, and consent is a marker, not access control. Against Wharton's one-sentence injection, provenance is the half of the problem that is solvable today. The other half is the model.

Every artifact above is checked with a public key and nothing else. No account and no registry membership. That is the whole point, and it is also why the kits do not compete with the protocols: MandateKit already exports RFC 9396 authorization details and tracks the AP2 mandate draft, so a mandate can sit under an AP2 or UCP flow rather than beside it. The protocols are partial by construction, and the partial bits are where the merchant gets left holding Lau's clause.

What comes next

MandateKit 0.2 will add the per-verb counterparty scope the Bench exposed, plus time windows, velocity limits, a revocation list and delegation. BudgetGuard 0.2 shipped the velocity windows; reserve and commit for concurrent agents, and a store that survives a restart, follow. After that: anchoring WitnessKit chains and RememberKit packs outside the signer's machine, key rotation for IdentityKit, and a single verifier call a merchant can run on an incoming request. Every change gets measured on the Bench before I describe it. The v0 labels stay until a third party has audited the shared core, and that audit comes before any 1.0. None of it is for sale, not yet.

If you run a checkout, the useful thing you can do is try to break it: the Bench harness and all three entries are public under open licenses, and a reproduction on your own hardware is a contributed entry. If you write protocols, the useful thing is to say which of the five your spec answers, and for whom. Waller pointed to technical standards. The standards exist. What is missing is the habit of answering all five questions at the door, in a form the other side can read.

Sources (42)

If the answer to "was this purchase authorized" lives in a platform's terms of service, who exactly is it an answer for?

Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.