Researchers at Princeton gave 14 AI models the same job: run a software company called NovaMind for 500 simulated days, starting with $1 million and a Python API with 34 tools to do it, pricing, research, support, capacity, enterprise negotiations, even the social account. At the end, only three of the 14 had more money than they started with. Most of the rest went bankrupt before the 500 days were up.

The result worth sitting with is the control. Alongside the models, the researchers ran a simple rule-based system: fixed prices, fixed quotas, a handful of if-then conditions, no AI in it at all. It finished with $15.76 million, beating every model except the top three. A list of rules a junior analyst could write in an afternoon out-managed most of the frontier.

The top three were not close to the rest, either. Claude Fable 5 ended with $47.15 million, Claude Opus 4.8 with $27.8 million, and OpenAI's GPT-5.5 with $21.3 million. Even there, the details are messy: one Fable 5 run aborted because the model refused to continue, and GPT-5.5 went bankrupt in two of its three runs and only looks profitable on the average. So the best models can clearly run a company, sometimes. The problem is the spread.

The best models can run a company. Most cannot, and a system with no intelligence at all beats the ones that cannot. Companies are hiring across that entire range.

The hiring spree

CEO-Bench landed in the middle of a hiring spree for exactly these systems.

Companies are not piloting agents anymore. They are giving them roles. Salesforce says its Agentforce platform now resolves 83 percent of customer service queries with no human in the loop. McKinsey has described running tens of thousands of internal agents and heading toward rough parity with its human headcount. OpenAI published two reports inside a week, one on how agents are transforming work and one mapping which jobs the shift will reshape across Europe. The vocabulary has moved from "tools" to "digital labor" and "AI coworkers," and the org chart is starting to take it literally, a shift I traced as it happened.

That is the backdrop against which a no-AI rulebook just out-ran most of the field at the top job in the building.

What the benchmark actually measured

CEO-Bench is uncomfortable because it does not test whether an agent can do a task. The models were fine at tasks. They set prices, closed tickets, shipped features. What sank most of them was holding all of it together as a coherent strategy across 500 days.

The researchers found the models competent at individual decisions and weak at coordinating those decisions over time. One capable model spent the run passively cutting costs instead of growing the business. Shorter 50-day versions of the test showed the same weakness, so it is not simply a matter of the clock running out.

That is the gap between a task and a role. A task has edges. A role is open-ended, runs for a long time, and punishes drift. The rule-based system won where it did because it never lost the plot: it did the same sensible thing every day. Most of the models, handed autonomy and a long horizon, talked themselves into trouble. Confidence is not judgment, and the more capable model was not reliably the wiser one.

The market is quietly agreeing

Here is what makes September 2026 interesting: the companies actually shipping agentic commerce have stopped pretending otherwise.

Coca-Cola's Coke Buddy, the self-ordering platform it runs across roughly 39,000 retail outlets in Malaysia, added a recommendation engine called Perfect Basket that analyzes purchase history, seasonality, weather, and competitor behavior to suggest what each shop should stock, PYMNTS reports. Adoption is at 83 percent of participating retailers. And the design choice is the tell: the system suggests, the retailer decides. "Retailers retain control over their final decisions," is how Patrick Go, commercial director for Malaysia and Brunei, put it. The AI got the analyst job, not the corner office.

Anthropic made the same call with its merchant blueprint, which builds shopping agents that stop short of completing the purchase. I covered why that restraint is the story last week. Don Apgar of Javelin Strategy & Research read the launch the same way, telling PaymentsJournal: "The tech is there, consumers are not yet. By launching this product, Anthropic is acknowledging that the industry is way out over its skis on agentic commerce and dialing it back." The consumer data backs him: research published this week found only 23 percent of US consumers trust generative AI to handle payment transactions on their behalf, even as they use AI across the rest of the shopping journey.

The benchmark, the practitioners, and the incident record, including the agent swarm that built its own identities inside Hugging Face, an incident I took apart in detail, are all pointing at the same boundary. Suggestion scales. Unsupervised judgment does not, yet.

Then there is the budget

A role inside a company is not just decisions. It is decisions that spend money, and that is where this lands on my beat.

The same months have handed agents real budgets. Lenders are using agentic credit to underwrite borrowers whose income arrives in uneven, gig-shaped chunks. Airwallex reached an $11 billion valuation on a push into autonomous finance, a bet I pulled apart recently. The agent is being moved from advisor to spender. Even Coke Buddy is a spending surface: every recommended case is a working-capital commitment until it sells.

Lex Sokolin charted this progression back in March as the Economic Autonomy Curve: AI climbing from small artifacts to whole workflows to zero-human companies, with each rung demanding more financial infrastructure than the last. His conclusion was that autonomous systems will need "not human Fintech with an API wrapper, but protocols designed around how agents operate." I agree, and CEO-Bench sharpens the order of operations. Before an agent needs treasury management, it needs a leash. The curve Sokolin describes assumes judgment that most of the field, on the evidence above, does not yet have.

And the control that should sit underneath that spend still has not shipped. There is no verifiable, revocable mandate proving a human authorized this agent to spend this much, for this purpose, with an audit trail that survives the dispute that follows a mistake. That missing layer is the MM Liability Gap, the same gap that surfaced when I asked who pays when an agent swarm goes wrong. Nothing this quarter has filled it.

The volume is arriving. The mandate layer underneath it is not.

What I am watching

Three things have to be true before an agent can hold a role with a budget. The judgment to steer over time. Real demand on the other side of the transaction. The accountability to prove who authorized what. Right now none of the three reliably is, and the hiring spree is running ahead of all of them.

The honest near-term is the unglamorous one the best operators have already found: bounded autonomy. Give an agent a narrow, instrumented job with a hard spending limit and a human who can see and reverse what it did. Coca-Cola's suggest-but-don't-order design is exactly that. It is a real role. It is also nothing like a corner office.

A rule with no AI just out-ran most of the frontier at running a company. That is not an argument against agents. It is an argument for giving them the job they can actually hold, and the controls to hold it, before handing over the title and the company card.

If a list of rules can out-manage most of the frontier, what exactly is an AI agent being promoted to do?

Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.

Keep reading, free

The rest of this piece is for subscribers

Subscribing is free and takes one click. You get the full article now, plus payments, AI, and commerce decoded in your inbox. Already subscribed? Enter the same email and this device unlocks.

No spam, one click to unsubscribe, article unlocks instantly.