Somewhere in the renewal you signed last quarter, there is a clause that lets your software vendor train artificial intelligence on your data. You almost certainly did not negotiate it. You may not have read it. It was not hidden so much as folded into language that has been sitting in service agreements for years, language that used to mean something far narrower than it does now.
PYMNTS named the pattern in June: enterprise software contracts have become, in effect, AI training licenses. The customer-side bar spent the summer catching up. Lawyers at Hunton Andrews Kurth published guidance in July on managing vendor data-use clauses, and their headline example is telling: HubSpot recently tried to expand its use of customer CRM data for AI-powered features, then reversed course quickly when customers noticed. The grab failed. The instinct behind it has not gone anywhere.
The clause that once let a vendor keep the lights on now lets it train a model on everything you put through the system.
The phrase that changed meaning
For most of the last decade, service agreements have contained a line giving the vendor the right to use customer data "to provide and improve the services." When that clause was written, improving the service meant fixing bugs, tuning performance, and building features. It was unremarkable, and buyers waved it through.
The same words now cover something much larger. Improving the service can mean training a model that becomes the vendor's core asset, sold back to you and to everyone else as a new capability. The text did not change. The technology underneath it did, and the permission quietly expanded to fit.
This is why the issue is easy to miss. There is rarely a dramatic new clause titled "AI training rights." There is an old clause doing new work. The Federal Trade Commission saw this coming: it warned back in February 2024 that quietly changing terms of service to enable AI training could count as an unfair or deceptive practice. Two years on, boards have noticed too. Lloyd Wilson of law firm Shumaker points out that 72 percent of S&P 500 companies now disclose AI-related risks in their SEC filings, up from 12 percent in 2023, and much of that disclosure is about third-party vendors.
Why every vendor wants this
Alex Johnson made the institutional version of this argument in a May Fintech Takes piece on AI agents coming for banks' customer data: "The differentiation will come from the data that empowers those models," not from the models themselves. He was writing advice to banks. Flip his point around and you have the SaaS vendor's playbook, because every platform vendor has already internalized it. For a company that will never train a frontier model, the most valuable thing surrounding the model is your data. Millions of real customer interactions are the one input a better-funded rival cannot buy.
So the incentive runs one way. Every platform that touches your operations has a reason to treat the exhaust from your business as raw material. Customer data stops being something the vendor stores for you. It becomes something the vendor can compound into a durable advantage, and the contract is where that right gets secured.
I wrote in the State of the Stack report that in the agentic economy the durable value accrues to whoever controls the data and the trust layer. Enterprise SaaS contracts are one of the quietest places that control is being transferred, one renewal at a time.
What the buyer is actually giving up
The exposure is larger than most teams assume, for three reasons.
The first is that training is not reversible in any practical sense. Once your data has shaped a model's weights, you cannot extract it the way you could delete a row from a database. Deletion rights that work for stored records do not map cleanly onto a trained model, which means the usual promise of "we will remove your data on request" may not mean what you think it means.
The second is competitive leakage. If a vendor trains a shared model on your data and your competitor's, the patterns learned from your business can surface, indirectly, in the help your competitor receives from the same tool. Nobody hands over your files. The model simply got better at your category because you taught it.
The third is regulatory. Data your company is obligated to protect, under privacy law or contractual duties to your own customers, does not lose those obligations because it passed through a vendor's platform. If that data ends up in a training set you did not properly authorize, the liability does not stay with the vendor.
How to read the contract
The practical defense is unglamorous, and it starts with reading the agreement as if the words mean what they now do.
Look for the "improve our services" language and ask, in writing, whether it includes training AI or machine learning models. Vendors will often answer plainly when asked directly, even when the contract is vague. Watch for terms like "aggregated" and "de-identified," which sound reassuring but can still permit training, and ask what de-identification actually means in practice. Check whether the default is opt-in or opt-out, because many platforms train by default and leave it to you to find the toggle. Trace the sub-processors, since the right to use your data often travels to partners you have never heard of.
The Hunton lawyers add two structural fixes worth stealing: write purpose limitation into the data clause, so each use of your data needs its own grant, and resist terms that let the vendor amend the agreement through a hyperlink. HubSpot's reversal shows why the second one matters. The attempted expansion did not arrive as a contract amendment. It arrived as updated terms.
None of this requires a confrontation. It requires treating the data clause as a live commercial term rather than a formality, the same way you would treat price or liability.
The governance gap underneath
Step back, and this is the same gap I keep running into across the agentic web. There is no shared, enforceable record of consent that travels with data once it leaves your hands. This is the authorization layer of the MM Trust Layer Model, and it is as unbuilt for data as it is for payments, a point I made in my coverage of the agentic identity crisis.
Until that layer exists, the contract is the only control surface a buyer has. That is why contract-layer fixes are becoming a genre. The open letter I published this week argues for contractual commitments as one of the six levers for AI accountability, and I released three open-licensed MSA riders at Major Labs covering incident disclosure, agent attribution, and liability. Notably, none of them covers training rights yet. That tells you how new this front is.
The vendors are not doing anything most of them consider underhanded. They are using the rights they were granted, in a market that now rewards data more than it ever has. The burden has shifted to the buyer to notice. The next time a renewal lands, the most valuable thing your legal team can do is read the clause that everyone has been ignoring, and ask what it permits today rather than what it meant when it was written.
Sources
- PYMNTS: Enterprise SaaS Contracts Are Secret AI Training Licenses
- National Law Review (Hunton Andrews Kurth): From "Service Improvement" to AI Exploitation
- Shumaker: When Your SaaS Vendor Becomes an AI Company Overnight
- Fintech Takes: Your Customers' AI Agents Are Coming for Your Data. Are You Ready?
- Major Labs: Compact Clauses
When the data clause in your contract has quietly changed meaning, whose job is it to notice, yours or the vendor's?
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.
The rest of this piece is for subscribers
Subscribing is free and takes one click. You get the full article now, plus payments, AI, and commerce decoded in your inbox. Already subscribed? Enter the same email and this device unlocks.
No spam, one click to unsubscribe, article unlocks instantly.