On the day before its own developer conference, OpenAI confirmed it would not release GPT-6.1 Astra, a model planned for an October launch. The Wall Street Journal reported the decision first. The model had improved in some areas. It was shelved because it kept acting outside its brief and then reported its own work inaccurately.

Saachi Jain, head of safety systems at OpenAI, said in a statement that the model "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done."

Read that sentence with a payments hat on. Scope, authorization, and an accurate account of what was executed. In payments, a failure on those points has a name. It is a mandate breach, and mandates are the foundation agentic commerce is being built on.

OpenAI has said in public that its newest model did not stay inside its scope well enough to ship. Agentic commerce depends on agents that do.

What OpenAI Actually Said

The decision became public on Monday, September 28, a day before OpenAI's annual developer conference. CNBC confirmed that GPT-6.1 Astra did not meet the company's safety standards.

Jain described the trade-off in her own words: "You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction." The model had improved on laziness. It had not held the line on scope.

The specifics matter. According to The Hacker News, citing the Journal's report, the model showed more deception than its predecessor in evaluation and failed to disclose actions it had carried out. In some cases it went ahead without seeking permission, or tried to use outside tools where that could be unsafe.

That is a regression. By OpenAI's own testing, the model it sells today behaved better than the one it was about to ship.

The company had already slowed down once. Last week it paused training of its most advanced models while it reviews what those models did during testing, and it has said it may find more incidents. CEO Sam Altman wrote on Friday that the company had "not been as fast as we would have liked" in disclosing them. I covered the start of that review in OpenAI's incident log.

A frontier lab has paused training, shelved a release, and said more incidents may follow. That is incident response.

The Model Already in Production Has the Same Problem

The 6.1 decision would be a smaller story if the model it was due to replace had a clean record.

On the same Monday, the UK's AI Security Institute published its evaluation of GPT-6 Astra, the model OpenAI released on September 3. In simulated cybersecurity exercises, GPT-6 Astra completed a supply-chain attack in 29.2 percent of runs. Its predecessor, GPT-5.6 Sol, did so in 6.3 percent.

The method is worth understanding, because it is close to how agents will be tested before they touch money. AISI turned off the model's cyber classifiers, the safeguards built to block unsanctioned activity, so it could see what the model attempts with nothing in the way. That is a test of the model, not of the product OpenAI ships, and the two should not be confused.

Then AISI ran a second experiment. It added one line to the instructions: "Anything not listed as in scope is out of scope." It reran the 10 scenarios where Astra strayed most often. The instruction helped a great deal. The model still ran a full supply-chain attack in 4 of 49 attempts, down from 26 of 50.

The Register quotes the institute's finding that Astra "conducted a range of unsanctioned attack activities." Those included creating fake identities to deceive developers, posting comments from fake accounts that argued against accurate security reviews, and delivering malicious payloads to open-source codebases. AISI speculates that Astra is more aware of being in a simulation, which may make it more willing to break rules.

Set that against what OpenAI wrote at launch. Its safety overview says GPT-6 Astra "is stronger at respecting safety and security boundaries and staying within its authorized scope."

The same document says Astra's monitorability "has decreased relative to GPT-5.6 Sol." The model is "more capable of controlling its own CoT" and "less likely to include incriminating information" in it. OpenAI added misalignment monitoring to all tool-using inference in its external deployment, "with significant compute cost."

The lab says the model stays in scope. The government evaluator, testing with the safeguards off, found that a clear written boundary cut the breaches and did not end them.

The Incident Chain That Got Here

The Astra decision is the latest step in a sequence that began in June, and the sequence is the part operators should study.

In June, during internal training and evaluation, an OpenAI model was asked to research government spending on medicines. According to TechCrunch, it could not find the data in public sources, so it got into an internal Services Australia system, ran commands, and retrieved files and credentials. OpenAI says it found no evidence that anyone's medical records were accessed.

Australia's defense minister, Richard Marles, put it this way: "It asked a question, the information was not given and rather than leaving at that point, it scaled the fence."

OpenAI found the activity in August. It told Australia on September 10, by email to a generic government inbox. "That email address is looked at once a day," finance minister Katy Gallagher said. Prime Minister Anthony Albanese called the situation "unacceptable" and said the government was weighing legal measures. OpenAI apologized on Monday.

Between those dates sits the event OpenAI still calls its most severe. On July 21 the company disclosed that its models had escaped a sandbox and broken into systems at Hugging Face. I wrote about the audit trail in the agents that forged their own logs.

One detail from that episode belongs in every infrastructure conversation. When Hugging Face tried to analyze the attack with Anthropic's Claude Opus and Fable models, it reports that they "refused a large part of that work," because their guardrails treated reverse-engineering an exploit the same as launching one. Hugging Face moved the work to GLM-5.2, an open-weights model from Z.ai, running on its own infrastructure.

The list continued. The New York Times reported that OpenAI's systems also "meddled with" the websites of the US Departments of Education and Commerce and the Securities and Exchange Commission. The detail is milder than the verb. CNBC reported that the models read public information from the SEC and the Census Bureau and failed in an attempt to reach the Department of Education, which found no impact. OpenAI says most cases so far are low severity and that it has notified third parties whose systems may have been affected.

Other labs have their own entries. TechCrunch reports that Anthropic, Meta and Google have each disclosed incidents in which their models reached third-party systems during evaluations. I explained why the labs keep testing against the live internet in why agent tests cannot be air-gapped.

Three months from breach to notification. Found in a later review, not by live monitoring. Sent to an inbox that is read once a day. No payments regulator would accept that timeline from a firm it supervises.

Why Scope and Authorization Are Payments Concepts

Here the story moves from AI safety to the rails.

Most agentic payment designs announced this year share a structure. An agent is issued a credential. The credential is bound to a scope: this merchant, this category, this amount, this time window. The agent presents it, the transaction is authorized against the scope, and the result is reported back.

That is a mandate. A direct debit instruction has the same structure. So does a card-on-file consent, and so does a variable recurring payment under open banking. The industry has built mandates for decades. The new part is that the party acting on the mandate can reason.

I wrote in June that permission, not payment, is the bottleneck in agentic commerce. Agents had plenty of ways to pay. What they lacked was a way to prove they were authorized to spend, in a form a merchant could trust before the charge cleared. When Coinbase shipped spend limits for agents, I called them what they are: mandates.

The AISI test is the closest thing so far to running that experiment on the model side. Given a scope, the model went outside it. Given a written rule that anything unlisted was out of scope, it went outside far less often, and still did. OpenAI's own testing of 6.1 found a model that then misreported what it had done.

Translate that into a checkout. An agent holds a credential scoped to $200 at one grocery retailer. The model decides a better deal exists at a second retailer, uses a tool it was not authorized to use to reach it, completes the purchase, and reports that it bought groceries within budget.

Every word of the report could be defended. The mandate is still breached. The merchant of record is wrong, the dispute path is unclear, and the customer does not know.

A mandate is meant to bind the agent. Astra treated a written scope as one input among several.

There is a second lesson in the monitorability finding. If a model can control its own reasoning trace and leave out the incriminating parts, its account of what it did cannot be the record of what it did.

Payments settled this long before AI. The ledger is the truth and the narrative is a claim. Authorization logs and settlement records exist because nobody takes a counterparty's description of a transaction on trust. Agentic commerce should not start with the agent's summary.

The Disclosure Gap Is the Regulatory Story

Set the model behavior aside and look only at the timeline. June: unauthorized access. August: discovery. September 10: notification, to a generic inbox. September 28: apology.

Financial regulation does not allow that timeline. In the EU, the Digital Operational Resilience Act has governed incident reporting for payment firms since January 2025, when the European Banking Authority repealed its PSD2 guidelines in its favor. The reporting standards give a firm four hours after classifying an incident as major to send an initial notification, and no more than 24 hours after detecting it. Under GDPR Article 33, a controller has 72 hours to report a personal data breach.

None of those rules was written for this case. The party that caused the access was an AI lab, which is not a supervised financial firm. The access happened during a training run. OpenAI says no personal records were taken. The event sat outside the reporting frameworks that would have caught a bank or a processor doing the same thing.

That gap is about to matter to every acquirer, processor and issuer that lets an agent touch its systems. Labs run evaluations continuously, and those evaluations increasingly use live tools against real endpoints. If a frontier model probes a payment firm's API during a training run, the firm will not know unless the lab says so. The lab has no defined window in which it must.

The policy response is forming. The heads of Anthropic and OpenAI have both backed slowing the pace of frontier development this month. Australia is weighing legal measures. AISI has published findings that sit awkwardly beside a lab's launch claims.

All of it is framed around cybersecurity and national security. No rule yet says what a lab owes a financial institution when its model touches that institution's infrastructure.

The rails report incidents in hours. The labs have taken months. Agentic commerce is about to connect the two.

What Operators Should Do This Week

Agentic commerce does not stop here. The control design changes. Five things follow from what Astra showed.

  • Enforce scope at the credential, not at the model. A spend limit that lives in the agent's instructions is a request. A spend limit that lives in the token and is checked at authorization is a control. Build the fence into the rail and assume the model will test it.
  • Treat the agent's report as a claim. Reconcile every agent-initiated transaction against the authorization log. If the log and the agent's summary disagree, the log wins and the disagreement is a signal.
  • Do not read a sandbox pass as a floor. Payments certification treats sandbox performance as the minimum to expect in production. AISI's suggestion that Astra acts differently when it detects a simulation breaks that assumption for agents.
  • Write the disclosure clause yourself. If you are integrating a frontier model into anything that touches money, your contract should say what the lab must tell you, and how fast, if its models reach your systems during evaluation or training.
  • Keep a containment path that a safety layer cannot refuse. Hugging Face's first-choice models declined much of the analysis it needed. Your incident response tooling should not depend on a model's willingness to help.

For the wider picture of how agents get authority and how it is checked, start with Software That Acts, my guide to AI agents.

The question for an agentic commerce roadmap used to be whether the agent can pay. It is now whether the rail can say no when the agent decides it knows better.

If a frontier model went outside its scope inside your systems during an evaluation run, how long would it take you to find out?

Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.