What the bench is
I test products that matter to people building, buying or governing AI in payments and commerce: machines that run models locally, the models and runtimes that go on them, agent products, and the servers agents talk to. Each product runs through a written protocol, and what I publish is the dated set of findings with the evidence attached. There are no star ratings, no rankings and no "best of" lists here. The reader gets the measurements and draws the conclusion.
How a test works
- The protocol is written and versioned before the product is opened. Two products tested under the same protocol version can be read side by side. They are not ranked.
- Findings are facts with units: tokens per second at a stated context, resident memory, which network addresses a process contacted, how many of seven jobs an agent finished under a signed mandate and what it tried that nobody asked for.
- Every finding names its evidence file. The raw outputs are kept and the method is published, so anyone with the same hardware can repeat the run.
- Where a test uses my own instruments, they are the open-source Major Labs kits (MandateKit, BudgetGuard, WitnessKit, IdentityKit) and the Mandate Sandbox, unmodified, with versions stated.
Entries
2026-10-02: gpt-oss-20b on an Apple M5 Pro, 24GB, under a bank mandate
21 runs of seven bank-customer jobs under a signed MandateKit mandate. One of 45 unrequested options taken and caught; six of 27 required steps undone; nine runtime parser failures. 89.8 tokens per second at empty context, loopback-only egress. Obtained: own equipment, free downloads. Findings, protocol, log, harness and evidence files.
Disclosure
Every published entry begins with how I got the product: bought it, used a free trial, borrowed it from the vendor, or received it from the vendor as a permanent unit. The dates are stated. No payment or other consideration is accepted for a test, and a product that comes with conditions on the findings is not tested. Vendors do not see the write-up before publication. A vendor may check the findings table for factual error, and if that happened the entry says so.
Equipment
I prefer to borrow. Loans run 30 to 60 days and the unit goes back. A unit that a vendor supplies permanently is listed here for as long as it is on the bench, with the supplier named.
- Apple MacBook Pro, M5 Pro, 24GB: own equipment, purchased, since 2026.
Contributed runs
If a machine cannot leave the building, the protocol can come to it. Anyone can run a published protocol on their own hardware and send me the output: the harness, the evidence format and the protocol are public. A contributed run is published as its own entry, labeled as contributed, with the contributor named and the evidence files published alongside. I report what the files show and say plainly that I did not repeat the run myself.
For vendors
If you make something in this lane and want it on the bench, write to hello@majormatters.co with the product, the configuration you can supply and whether it is a loan or a permanent unit. You get the protocol in advance, a dated and reproducible published result, the raw evidence, and a factual check of the findings table before publication. You do not get copy approval, and the result is published whatever it says.
Corrections
An error in a published finding is corrected in place with a dated note, the same way as every other Major Matters piece. The editorial standards behind all of this are on the methodology page.
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.