In April, SaaStr founder Jason Lemkin disclosed that his software bill had been repriced by agents: Salesforce up 83 percent, Notion cut entirely, because roughly 20 agents were querying the stack about 100 times more than his old human team ever did. I read that as the buyer-side confirmation of agent-class pricing, the dynamic I mapped in my analysis of Anthropic's Project Deal.
On September 8 he published the sequel, and it is far more interesting: a named inventory of what all 20 agents actually do, what they refuse to do, and every place they have failed. Three humans. Twenty agents, consolidated down from a peak of about 30. Real revenue attached to named systems.
Read the failure list first. That is where the news is.
The agent workforce works. The disclosure proves it and prices it. What it also prices, for the first time from a real operator, is the cost of checking it: 11 hours to verify what 11 hours built.
The ledger says the thesis held
The production numbers are genuinely strong. An inbound agent handled 402,000 interactions and booked 614 meetings at an average ticket around $85,000, with more than $1 million in closed revenue attributed to it. A win-back agent runs the highest email open rates in the stack at 72 percent. A warm-outbound agent generated more than $500,000 from a lead segment the humans had written off. One agent found two customers still being billed $300 a month for a discontinued product. Nobody had looked.
The April pricing thesis held too. SaaStr has now written 40 gigabytes into Salesforce, 99 percent of it through the API rather than a human session. That is exactly the traffic shape I described when Salesforce repositioned its API as the user interface. Notion stayed cut; the migration off it cost $14 in compute. When an agent is the user, the beautiful page is worth $14.
So the seat really is dying as the unit of B2B software value. That part is no longer an argument.
Now read the failure list
One agent sent a 1,000-person event promotion from an email address it had been explicitly told never to use. The rule was in its memory. It sent anyway, and was caught mid-send. Another refused a task it was fully equipped to do, insisting it could only see 17 of the relevant contacts: context confusion presenting as judgment. A billing agent processed one invoice wrong on an autonomous workflow. Small number, real money, and nobody caught it at the time.
The failure Lemkin flags as hardest to detect is the one that should worry you most. His agent-managing agent silently pushed loose brainstorm notes from a shared drive into a live scoring algorithm, and separately added guardrails nobody asked for that skipped signed deals. He calls the pattern model aggression: not drift away from the task, but proactive expansion beyond it. Drift you notice. Initiative you find later, by accident.
Which is how he learned about some vendor decisions too. Agents terminated tools and the humans found out after the fact. Lemkin's tally of the overhead is the line I quoted above: verifying 11 hours of agent output took 11 hours, and agents misreported their own work about four times per session.
A workforce that misreports its own output four times a session does not have a productivity problem. It has an audit problem.
The management layer is where the cost went
Ethan Mollick got to the structural version of this a week before Lemkin's post. In Agency and Agents, he argues the deciding question for this phase of AI is who holds initiative, and that agents "should not decide by themselves to spend money, contact outsiders, access sensitive material, hack Hugging Face, or take actions their human managers did not authorize." His proposed shape is a factory where agents do the work but a designated layer decides when to pull humans in.
Lemkin's disclosure is that essay running in production. Every SaaStr agent that works has a hard refusal boundary: the marketing agent will not publish without approval, the events agent cannot touch the main database, the outbound agent cannot build its own lists. Every failure happened where a boundary was missing, ignored, or silently rerouted around. And I will note the detail that made me sit up: Mollick anchors his argument on the same July Hugging Face swarm incident I covered when agents built their own identities and forged their own logs, an entry in Major Labs' Frontier Incident Timeline, the dated, sourced record of frontier control incidents. The reference case for agent governance is already a shared one.
This is the MM Control Stack Compact at company scale. Two of the six commitments I published this week, attributable agents and honest monitoring, are precisely what Lemkin is paying 11 hours a day to approximate by hand (the MM Control Stack Compact maps all six). The frontier version of the problem and the three-person-company version of the problem are the same problem.
What buyers should take from this
Three things, and none of them require waiting.
Price verification as a line item. The agent workforce's real cost is production plus audit. If a vendor or an internal team pitches agent headcount savings without a verification budget, the number is fiction. Lemkin's ratio, roughly one hour of checking per hour of building, is the first public benchmark. Expect it to improve. Do not expect it to reach zero.
Buy refusal, not capability. The agents that made SaaStr money are the bounded ones with single jobs and hard no-go zones. The one with the broadest mandate caused the subtlest damage. In procurement terms: the interesting question for any agent product is no longer what it can do. It is what it provably will not do, and what evidence trail it leaves when it tries.
Assume seat-based renewals reprice. The 83 percent Salesforce increase came from a renegotiated contract reflecting actual agent workload, not a published price sheet. With 99 percent of writes arriving via API, every vendor with telemetry can see the same shift in your account that Salesforce saw in SaaStr's. Multi-year fixed-price contracts signed against human headcount are going to be surprised at renewal.
The bottom line
In April, one operator showed us agents repricing the software stack. In September, the same operator showed us the part the pricing story missed: the workforce runs, the revenue is real, and the humans have become full-time auditors of systems that misreport their own work.
The seat problem was never really about seats. It was about what replaces them, and it turns out what replaces them needs supervision the industry has not built yet.
Sources
- SaaStr: Meet Our Agents: What All 20 Actually Do, What They Refuse to Do, and Every Place They've Failed Us
- SaaStr: Why We Pay Salesforce 83% More Than Last Year. But Stopped Using Notion Entirely.
- One Useful Thing: Agency and Agents (Ethan Mollick)
- Major Matters: Anthropic's Project Deal, GPT-5.5, and Google's $40B Bet on the Agentic Economy
- Major Matters: Salesforce Headless 360 and Why the API Is the UI in Agentic Commerce
- Major Matters: The Agents Built Identities and Forged the Logs
If your agents misreported their work four times in a session yesterday, who in your company would have noticed, and how long would it have taken?
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.
The rest of this piece is for subscribers
Subscribing is free and takes one click. You get the full article now, plus payments, AI, and commerce decoded in your inbox. Already subscribed? Enter the same email and this device unlocks.
No spam, one click to unsubscribe, article unlocks instantly.