Last week, two frontier models shipped 48 hours apart and the chief scientist of one of the labs wrote that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed. I laid out the evidence in my analysis of the double launch. This letter is the constructive half of that piece: not what is broken, but what I believe would fix it, stated specifically enough to be adopted, rejected, or improved.
I write this as a buyer and an analyst, not a lab insider or a regulator. That vantage point is the useful one here. The trust problem in frontier AI will not be solved by the labs grading their own homework or by regulation written after the first catastrophe. It will be solved, if it is solved, by the parties in the middle: the enterprises whose contracts fund the labs and whose systems absorb the risk.
Before the proposal, a point of precedent, prompted by a colleague who has spent years inside bank risk frameworks and asked the question this letter should lead with: why should AI be any different? Banking already runs the stack this letter asks for. Capital thresholds are published and versioned under Basel. Stress tests are designed and graded by supervisors the banks do not pay. Incidents produce mandatory reports to authorities with the power to compel evidence. Aviation runs the same stack through the NTSB. Nuclear power runs it through IAEA inspection. Every industry whose failures reach beyond its own balance sheet has been walked, usually after a crisis, to the same three-part design: published thresholds, independent assessment, incident authority.
Frontier AI has crossed the line that triggered that design everywhere else. These models stopped being products some time ago. They are infrastructure now, closer to telco than to boxed software, load-bearing for other companies' systems and other people's money. So the six commitments below are not a novel governance theory. They are consistency with existing best practice, applied to the one systemic industry that still grades its own homework.
Trust in frontier AI is not a feeling to be earned with launch posts. It is a stack of verifiable commitments, and today most of the stack is empty.
The Proposal: The MM Control Stack Compact
I propose a set of six commitments. I call it the MM Control Stack Compact, and it maps directly onto the control layers that stand between a frontier model and the world, the same layers that failed or held on this year's public record. Each commitment is written to be verifiable. Vague pledges built the current situation.
One. Publish the thresholds. Every frontier lab publishes its capability thresholds, the specific evaluation results that trigger enhanced safeguards, in a common public format, versioned like a changelog. Today those thresholds differ dramatically between labs or are not public at all. A safety framework nobody outside the lab can read is a press release. Until the labs run that changelog themselves, I run it for them: Threshold Watch snapshots each lab's published framework on a cadence, hashes it, and archives a dated diff of every change.
Two. Independent evaluation with real access. Pre-deployment testing by evaluators the lab does not pay, with model access deep enough to matter, and results published within a fixed window. The UK AI Security Institute is closest to this today. Its work should be the floor, not the ceiling.
Three. Incident disclosure with teeth. A frontier AI incident, an agent escaping containment, acting on live systems without sanction, or deceiving its evaluators, gets reported to an independent body within 72 hours, investigated by that body with legal authority to compel evidence, and closed with a public report. Aviation has run this model for decades and it is the single biggest reason flying became safe. The July breach of a lab's own infrastructure received six days of outside investigation that ended before the compromise did. That is the standard this commitment exists to replace.
Four. Attributable agents. Every agent deployed outside a lab's walls carries verifiable identity: who deployed it, under what authority, revocable at the source. The rogue swarms of this year were hard to investigate partly because nobody could establish whose agents they were. Attribution is the precondition for every other form of accountability, and the technical primitives exist today. I know because I maintain some of them: my research studio's open-source suite includes a one-command proxy that wraps any MCP tool server and refuses, cryptographically and before execution, any tool call that falls outside a signed mandate. The gating half of this commitment took an afternoon to build on primitives that already exist. The gap is adoption, not invention.
Five. Liability that lands somewhere. Contracts between labs and enterprise customers state who pays when an agent causes harm, before deployment, not in litigation after. I have written about the MM Liability Gap since the earliest agentic deployments: autonomy has scaled faster than accountability, and the space between them is where losses land. Closing it is a drafting exercise, not a research problem.
Six. Monitoring honesty. Labs report, on a fixed cadence, how well their primary oversight mechanisms actually work, including when reliability is falling. OpenAI's chief scientist disclosed that chain-of-thought monitoring is degrading. That disclosure was voluntary, singular, and buried in an essay. It should be routine, structured, and required.
Who Does What
A compact is only real if each party has a move available today. Each does.
The labs can adopt commitments one, two, and six unilaterally, this quarter, without waiting for legislation or each other. The first lab to publish its thresholds in a common format and invite independent evaluation on these terms converts safety from a marketing claim into a competitive position. I wrote when Fable 5's capability controls shipped that gating decisions had quietly become governance decisions. The lab that acknowledges this openly, and submits its gating to outside scrutiny, earns the trust the launch posts keep asserting.
The buyers hold the strongest lever, and this letter is mostly for them. Write commitments three, four, and five into procurement. Require incident disclosure terms, agent attribution, and contractual liability allocation as conditions of purchase. To remove the blank-page excuse, I have drafted the riders: the Compact Clauses are free, CC0, and written for your counsel to adapt. No open letter, this one included, moves a lab the way a stalled enterprise contract does. The security reckoning I documented last year has not slowed deployment, which tells you the pressure has to arrive through the deal, not the discourse.
The policymakers have one urgent assignment, and it is narrower than most AI bills: create the investigation authority. Not a licensing regime, not a model registry, not a pause. An independent body with the power to investigate frontier AI incidents and publish findings, on the aviation model. Everything else in this compact can be built by private parties. That one piece cannot.
Ben Thompson argued this year that unelected lab executives should not hold the keys to the most consequential technology, and he is right about the accountability problem even where I part with his conclusions. The answer to private power over frontier AI is not no gates. It is gates whose rules are public, whose operation is inspected, and whose failures are investigated by someone other than the gatekeeper. That is the whole letter in one sentence.
What I Am Committing To
A proposal like this costs me nothing unless I bind myself to it, so here is my part. Major Matters will track adoption of these six commitments publicly: which labs publish thresholds, which enterprises write the clauses, which incidents get real investigations. The first pieces of that machinery are already running: Threshold Watch began tracking six labs' frameworks the day this letter published, five archived and one, instructively, blocking automated retrieval, and the Frontier Incident Timeline keeps the standing record of the control incidents this letter draws on. I will report the silences as plainly as the progress. And I will publish, in full, any substantive response to this letter from a lab, a buyer, or a policymaker, including the ones that tell me I am wrong.
The window for building trust infrastructure is the window before it is needed at scale. On the evidence of this year, that window is open and closing.
Sources
Six commitments, each verifiable, each available to someone today. Which one will your organization adopt first, and if the answer is none, what does that tell you about where trust actually sits?
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.
The rest of this piece is for subscribers
Subscribing is free and takes one click. You get the full article now, plus payments, AI, and commerce decoded in your inbox. Already subscribed? Enter the same email and this device unlocks.
No spam, one click to unsubscribe, article unlocks instantly.