Outcome-Based Pricing Pilots in B2B SaaS
Vendors piloting outcome pricing must solve attribution and measurement before launch, not after.

Outcome-based pricing has a launch problem, not an interest problem. Roughly half of SaaS companies say they're exploring or piloting it, but the number running live, revenue-bearing deals is far smaller, and that gap is where the real work sits. The pilot phase is where the model proves itself or gets shelved quietly after a single renewal cycle, and the pitch deck doesn't decide which. Four decisions made before go-live do: what counts as an outcome, what caused it, how the invoice gets built, and whether that invoice survives a dispute.
The pressure behind this shift is structural. Enterprises have poured tens of billions into generative AI and, according to the MIT study widely cited on this topic, the overwhelming majority report no measurable return. Buyers are done paying for seats and promises. Even McKinsey, a firm that sells judgment rather than software, now ties a meaningful share of its fees to outcomes and describes the change as client-led rather than something it volunteered. Gartner's forecast that a large share of enterprise SaaS spend shifts to usage, agent, or outcome-based models by decade's end says this trend doesn't reverse. Companies piloting outcome pricing now are building institutional muscle memory that competitors will pay to acquire later, and the acquisition cost only gets steeper from here.
What "outcome" actually means in a contract, and where the definition breaks down
Most products marketed as outcome-based are still measuring output, and treating the two as interchangeable is the first mistake. A resolved ticket, an approved transaction, an automated response: these are things the product did, not necessarily things the customer wanted. A pricing spectrum for job-board services makes the distance visible: from usage-based (pay per listing) at one end, through intermediate steps like clicks and applications, all the way to a hire who stays for a defined period at the far end. Each step moves closer to what the customer actually cares about, and each step also gets harder to measure and easier to dispute. Vendors clustering at the easy end of that spectrum and calling it outcome pricing are the ones whose claims deserve scrutiny, and buyers should treat that label with real suspicion until the contract proves otherwise.
Zendesk's Automated Resolution metric is a useful, narrow example of a vendor trying to land on solid ground. An AR counts when a support ticket gets resolved by AI without a human stepping in, and it's only confirmed after 72 hours pass with no further customer activity. That waiting period serves a deliberate purpose: it's the contractual stand-in for "the customer didn't need to come back," which is the actual thing being sold. A definition that can't survive a dispute is worse than useless, because it looks precise right up until someone tests it.
Gorgias shows how fast this gets complicated inside a single customer interaction. One conversation can trigger a helpdesk ticket fee plus a separate automation fee if AI resolves it fully. If AI escalates to a human, only the ticket fee applies. If a proactive alert goes out and the customer ignores it, nothing gets billed at all. Three billing states, one interaction, and if the product can't classify which state it's in reliably, the invoice inherits that confusion.
This is why the Outcome Measurement Agreement has to exist before the contract gets signed, not after. Definitional failure comes in two flavors. Either the product genuinely can't tell success from near-success at the margins, or the definition is clean but the customer holds the measurement data, which makes independent verification difficult and disputes harder to resolve. The standard list of prerequisites for outcome pricing, a clear link between the service and a measurable benefit, tracking systems that actually work, defined timelines, stakeholder agreement on thresholds, and the technical capability to measure any of it, reads less like an audit checklist and more like a gate. Teams that treat it as cleanup work for after launch are the ones renegotiating mid-contract six months later, and by then the customer relationship has already absorbed the cost of that delay.
Attribution logic: connecting the product's action to the outcome the contract charges for
Even a well-defined outcome can be wrongly attributed, and this is the failure mode most teams underestimate. A ticket might close after the inactivity window elapses, satisfying the vendor's rule on paper, but the customer may have solved the problem some other way entirely, given up, or routed around the bot through a different channel. The billing engine doesn't know that. It bills anyway.
Fraud prevention is the cleanest version of this problem, mostly because it's already solved. Riskified charges only for transactions it approves that turn out to be fraud-free, and it guarantees those approvals against later chargebacks. Attribution here is binary, and it's verified by a record neither party controls: the chargeback itself. Identity verification services run on similar logic, charging only for successfully approved users, where the approval decision is vendor-controlled and failure is observable almost immediately. Both categories succeed because the outcome and the cause sit close together and get checked by a third party.
Customer service AI doesn't get that luxury. Intercom's Fin prices at 99 cents per successful resolution, a figure set below what a human agent interaction costs, but "resolution" is a softer signal than "fraud prevented." Did the AI answer the question, or did the customer quietly abandon the chat and solve it themselves? The product has to define, in a way that holds up, what it did versus what happened around it.
The deeper structural risk shows up when the outcome metric lives inside the customer's own systems, a CRM, an ERP, a support platform the vendor doesn't own. At that point the vendor depends on a data feed it can't verify unilaterally, which makes audit rights and data-sharing terms mandatory, not a clause buried in an appendix and forgotten. On the engineering side, attribution logic needs event capture at the moment the action happens, not reconstructed later, plus a defined lookback window and an explicit rule for what happens when an outcome gets contested. None of that is a legal afterthought. It gets built into the billing system at the same time it gets written into the contract. The same discipline applies regardless of industry: both sides need to agree on the impact metric and the measurement method before the engagement starts, not after the invoice arrives.
How to structure the pilot itself: scope, duration, and what you are actually testing
A pilot here tests trust and delivery, not price. Treating it as a price-discovery exercise is the mistake that sinks most of these programs before month three. Run it as a calibration exercise instead, one meant to prove the outcome metric behaves the way the model assumes before that metric starts generating real invoices.
Onboarding should look different depending on the segment. Low-touch, self-serve customers can get a free training window before charges kick in. Mid-market deals work better with a capped number of free outcomes and a clear rule for when paid conversion starts. Enterprise deals usually need an implementation fee, an annual minimum, or a pilot credit, something that funds the onboarding period and states, in writing, the exact date billing begins. The sequencing logic follows from this: push toward a hybrid model with a hard consumption ceiling while AI adoption is still early and unpredictable, move to pure usage-based once patterns settle, and only commit to full outcome-based billing once the instrumentation can actually support an Outcome Measurement Agreement.
Cohort selection matters more than most teams give it credit for. Run the pilot on a segment where the outcome is already happening and already measurable, rather than trying to validate the pricing model and the outcome definition at the same time. That's two experiments stacked on top of each other, and when something breaks, nobody can tell which variable broke it.
Duration needs a stopping rule decided in advance: how many outcome events make a valid sample, and what variance triggers a revision to the model. Skip that threshold and the pilot just runs forever, quietly, until someone in finance asks why it's still called a pilot after fourteen months. What the pilot actually tests comes down to three things: whether the outcome metric holds steady across different customers, whether the attribution logic produces invoices people are willing to sign off on, and whether the billing system generates records granular enough to survive an audit. A hybrid structure, a base fee plus an outcome-based variable, is the lower-risk staging ground here; industry survey data (Kyle Poyar's 2026 State of B2B Monetization survey of over 230 companies) puts more than a third of software companies already running hybrid as their primary model, which suggests the market largely arrived at the same conclusion independently.
Pricing the outcome unit: how to set rates before you have stable data
The rate has to reflect value delivered, but value is only knowable after outcomes have piled up long enough to trust the average. Early pilots are, by definition, pricing on incomplete information, and pretending otherwise is how a rate gets set too low to sustain margin or too high to survive procurement. Anyone who insists the launch rate is the right rate is guessing. The honest version of the pilot admits that up front and builds in room to move.
Zendesk's published rate architecture splits the difference: $1.50 per committed Automated Resolution for customers who commit to volume, versus $2.00 pay-as-you-go (as of August 2024). The gap rewards commitment and gives the vendor a cushion when volume is uncertain, which it always is at this stage. Intercom's 99-cents-per-resolution rate for Fin works because it's anchored below the cost of a human agent handling the same ticket, a number the customer can check against their own support costs without needing to trust the vendor's math. That's the value-anchor method in practice: price the outcome as a fraction of what the customer used to pay for it through a human or legacy process, and the ROI case makes itself.
There's a cost-side risk that rate-setting has to account for directly. AI products commonly run at gross margins in the 50 to 60 percent range, well below the 80 to 90 percent margins typical of conventional SaaS, and inference costs are falling fast, and competitive pressure from commodity models is accelerating that trend. A rate locked in today at 99 cents could look generous to the vendor next year and painfully high the year after, once cheaper commodity models put competitive pressure on price. The fix isn't picking a better number today, it's building a contractual repricing mechanism into the deal from day one, so the launch rate is understood by both sides as a starting point, not a permanent fixture.
Credit models are the practical middle ground while all this settles. The number of companies using credits has grown sharply, rising from 35 to 79 according to the PricingSaaS 500 Index, a jump of over 120 percent year on year. Credits let a vendor price something close to value without needing a fully locked outcome definition yet, which makes them a sensible bridge for teams still working out what the outcome unit even is.
The billing mechanics that outcome-based pricing actually requires
Subscription billing runs on a batch cadence: tally usage at the end of the period, generate an invoice, move on. Outcome billing can't work that way, and the underlying mistake in most of the disputes above is treating it like an extension of subscription billing anyway. It needs event capture at the moment the trigger fires, aggregation that's stateful across a customer's entire history, and invoice line items that point back to a specific event record rather than a rolled-up count.
That metering requirement isn't negotiable. A system that can't ingest and timestamp individual outcome events in real time cannot produce an invoice that holds up when a customer disputes it, because the audit trail is the invoice. Gorgias's three-state billing problem, full resolution fee, ticket-only fee, or no fee, shows why architecture matters here as much as definition does. If the classification decision gets made hours after the interaction, in a nightly batch job, the system is reconstructing intent after the fact instead of recording it as it happened, and reconstruction introduces errors that compound across thousands of daily events.
Spend visibility isn't optional polish either. A large majority of IT leaders, cited at 78 percent in industry surveys, report being hit with unexpected charges under consumption or usage-based pricing tied to intelligent-software features, and cost forecasting for such deployments reportedly ranks as the top challenge for roughly 90 percent of CIOs. A live usage dashboard needs to exist before the billing model goes live, not as a follow-up feature request after the first surprised phone call from finance.
Enterprise pilots frequently need two billing modes running at once: a prepaid credit block that funds the pilot upfront, and postpaid invoicing for anything that runs over. Both need to run on the same engine, without a manual reconciliation step stitching them together after the fact. A metering tool bolted onto a separate billing tool through an integration layer is not, functionally, a billing platform built for outcome pricing, no matter what the vendor calls it in the sales deck. Classification errors that happen at that seam stay invisible until an invoice gets disputed, and debugging them means pulling logs from two different systems at once, a task that eats a week of engineering time nobody budgeted for.
Invoice auditability: why outcome-based invoices require a different evidentiary standard
A standard SaaS invoice summarizes. An outcome-based invoice has to document each line item with a timestamp, an event reference, and the specific contractual rule that triggered the charge. Skipping that step is the single most common reason renewals stall.
Picture a customer disputing 200 Automated Resolutions at renewal. An invoice that reads "200 Automated Resolutions at $1.50" settles nothing. An invoice that ties each one back to a ticket ID, a resolution timestamp, and confirmation that the contractual inactivity window actually elapsed resolves the dispute in one meeting instead of three. Structurally, when the customer holds the measurement data, there's a standing incentive to contest anything near the margin, and the vendor's only real defense is the audit trail sitting behind the invoice.
This is what makes a rule like Zendesk's 72-hour window valuable beyond its role in the outcome definition. It's machine-verifiable, which means it produces a record that a dispute-resolution process can actually check. A vague outcome definition produces a vague record, and a vague record can't be audited consistently no matter how good the underlying product is.
Idempotency deserves specific mention here, because it's the kind of engineering detail that turns into a business problem the moment it fails. Production billing systems have to guarantee each outcome event gets counted exactly once. A duplicate charge on an outcome invoice calls for a real apology, not a shrug: it's a contract violation, and idempotency keys plus deduplication logic are the infrastructure that prevents it, not an optional hardening step for later. Recovery targets matter for the same reason: billing pipelines need something close to a one-hour recovery time (the maximum stretch of downtime before events stop getting metered) and zero event loss on recovery. A gap in the event record is a gap in the invoice, and there's no clean way to reconstruct it after the fact.
Finance teams should watch one signal above all others. If the first week of every month goes to manually reconciling outcome counts, that's evidence the billing architecture producing it is flawed, and the fix is engineering.
What a successful pilot looks like at the 90-day mark, and how to decide whether to scale it
Three signals suggest the model has actually validated. Outcome events get classified consistently, and customers aren't pushing back on the classification. The billing system produces invoices people accept without a call to customer success to walk through the math. And outcome volume is stable enough that revenue can be forecast within a range someone's actually comfortable defending to a board.
Three other signals point the opposite direction, and any one of them should stop a scale-up decision cold. Attribution disputes keep recurring on the same event types, which means the definition itself needs revision, not just better customer communication. Invoice disputes require someone manually pulling logs to resolve, which means the billing architecture needs rebuilding, not patching. Or outcome volume swings too wildly to price with any confidence, which means the model needs a floor, in practice a hybrid structure with a base fee underneath the variable.
Retention is the metric that confirms all this, but it shows up late. Companies running outcome-based components report notably higher retention and satisfaction than peers on flat subscription pricing, retention gains cited at roughly 31 percent. Those numbers won't materialize inside a 90-day window, so treating day 90 as a retention checkpoint misreads what a pilot can actually tell you this early. What the pilot needs to establish by day 90 is the measurement baseline itself, so retention and satisfaction can actually get tracked against it at the first renewal.
Scaling past pilot means freezing the outcome definition so it can't drift mid-contract without a formal amendment process, fully operationalizing the billing pipeline so no step still depends on someone manually checking a spreadsheet, and publishing the Outcome Measurement Agreement as a standard exhibit rather than a bespoke negotiation every time. None of this is static once it's live, either. The PricingSaaS 500 Index tracked more than 1,800 pricing and packaging changes across the top 500 B2B and AI companies in 2025, an average of 3.6 changes per company across the year, which shows outcome-based pricing demands continuous adjustment after launch, not a set-and-forget contract term. It gets recalibrated continuously as outcome volumes shift, AI costs move, and competitors reprice around it.
The last trap catches teams who did everything else right. A pilot that ran cleanly on manual tooling, someone checking classifications by hand, someone reconciling invoices in a spreadsheet before send, hits a wall the moment it scales past a few dozen accounts. The billing system that handled fifty pilot customers without automated event ingestion, real-time classification, and audit-ready invoice generation built into the architecture from the start will not handle five hundred, full stop. That's a reason to trust the pricing model. It's a sign the infrastructure underneath it needed to be built before the pilot proved the model worked, not after.
Sources
- What’s the Endgame for SaaS Pricing Models After the AI panic?
- The Rise of Outcome-Based Pricing in SaaS: Aligning Value With Cost
- Outcomes-based Pricing in B2B Situations
- How to Build an Outcome-Based Pricing Plan - The SaaS CFO
- Understanding Outcome-Based Pricing | Pragmatic Institute
- What is outcome-based pricing? How SaaS companies use it
- Outcome-Based Pricing For B2B SaaS: The Model Where Your Revenue Depends On Whether You Actually Deliver - Ciente
- Outcome-based pricing: A guide for businesses | Stripe


