Overage Invoice Design Experiments and Voluntary Upgrade Rates
Cursor's billing surprise reveals how invoice design drives upgrades or churn.

Most SaaS teams file the overage invoice under accounting rather than product strategy, and that is the first mistake to correct. An overage invoice is a product surface with measurable conversion behavior attached to it, no different in kind from an onboarding flow or a paywall, because the customer reads it and makes a decision. That decision has accelerated into routine territory for AI products specifically: roughly half of formally monetized AI offerings now run on usage-based, consumption, or outcome-based pricing. The overage invoice has become a standard customer touchpoint across AI-native and AI-infused products alike, not an edge case reserved for infrastructure vendors.
The clearest illustration of what happens when this surface is left undesigned comes from Cursor's Pro plan change in June 2025. The plan moved from a fixed number of fast requests to usage billed at API rates. Customers burned through their allowance faster than expected, particularly on complex prompts, and the bills that followed came as a surprise to people who had no clear signal they were approaching a limit. Cursor's CEO issued a public apology and promised refunds within weeks, though many users reported delays in actually receiving them. The plan change itself was defensible on pricing-mechanics grounds. What failed was the invoice: it delivered a number without delivering the context a customer needed to understand why that number existed.
That failure makes visible the two outcomes that branch from every overage moment. A customer who understands what drove the charge and sees a logical next step will upgrade. A customer who feels ambushed and penalized will churn, complain publicly, or both, as happened at scale with Cursor.
The two pricing postures that produce structurally different overage invoices
Before any invoice design experiment can produce a meaningful result, a team has to settle a prior question: what is the overage charge supposed to accomplish? Two strategic postures answer that question differently, and each produces an invoice built to do a different job.
Under what can be called the Nudge strategy, the overage charge exists partly to create a moment of friction that makes upgrading to a higher tier the obvious next step. The invoice under this posture should foreground the gap between the plan a customer is on and the plan that would have avoided the charge. Under what can be called the Meter strategy, the charge is a direct price-for-consumption signal. AWS and Twilio operate this way: customers pay for what they use, and the invoice is built to reflect that usage with precision. There is no higher plan to sell, because the pricing structure is already granular.
Few SaaS or AI products sit purely in either camp. Most operate in a hybrid: a subscription floor that covers a baseline of usage, with variable overages layered on top once that baseline is exceeded. This hybrid territory is where invoice design experiments matter most, precisely because the invoice has to serve two logics at once. It has to communicate consumption with the precision a Meter-style customer expects, while also surfacing the upgrade path a Nudge-style strategy depends on for growth. Deciding which posture governs a given product, or which posture governs a given overage event within a hybrid product, has to happen before any experiment begins, because it changes what counts as success. Under a Nudge posture, success is upgrade rate. Under a Meter posture, success is measured in fewer billing disputes and clearer usage transparency, because there may be no higher tier to sell.
Overage invoice design in current enterprise practice
The typical overage invoice fails for reasons that have nothing to do with arithmetic. The numbers are usually correct. What's missing is the context that would let a customer act on those numbers: why the charge occurred, what specifically consumed the allowance, and what the customer should do differently going forward. An invoice that states a dollar figure without that context forces the customer to reconstruct the story themselves, and most customers will not bother. They will simply conclude the vendor overcharged them.
Credit-based billing systems introduce a sharper version of this problem. When a vendor reserves the right to change credit multipliers unilaterally, a feature that once cost one credit can become substantially more expensive overnight, and the customer has no way to see that shift coming from the invoice alone. The charge becomes structurally uninterpretable: the customer cannot reconstruct it from the information the vendor provides. This is the underlying mechanism behind the Cursor incident, and it reflects a broader pattern documented across enterprise SaaS billing in 2026.
A well-designed overage invoice would show which feature, workflow, or agent actually consumed the overage, broken out as a line-item detail. It would project what the following month's bill looks like if the same usage pattern continues. It would set the cost of the current plan plus overage directly against the cost of the next tier, so the comparison is immediate rather than something the customer has to calculate on their own. And it would surface one specific action to take, rather than a generic link to a pricing page that leaves the customer to figure out the rest.
The absence of that information has consequences that extend past the individual invoice. Zylo's 2026 SaaS Management Index found that 78% of IT leaders experienced unexpected charges tied to consumption or AI features in the past year, in many cases costs that appeared after contracts were already signed and then scaled faster than anyone had planned for. Those charges land differently depending on who is driving them. When a marketing team runs an AI campaign that consumes a budget an IT or finance team is accountable for, or when individual employees adopt tools on their own initiative, an opaque overage invoice turns into an internal accountability problem, with one team unable to explain to another why a number came in the way it did. A customer who cannot explain a charge to their own finance team is not going to quietly absorb it and upgrade. They are going to dispute it, escalate it internally, or assume the vendor made an error, and any of those outcomes is worse for the vendor than the surprise charge itself.
The three invoice design variables that directly influence upgrade decisions
Once the strategic posture is settled and the informational gaps are named, the practical work is identifying which specific design choices actually move the upgrade decision. Three variables account for most of the variance, and each deserves to be treated as its own design problem, not a checkbox on a feature list.
The first is consumption framing: how the invoice represents what the customer actually used. A vendor can show raw units over the plan limit, a percentage of the plan consumed, the dollar cost of the overage in isolation, or the dollar cost of the overage set directly next to the cost of the next plan tier. Percentage framing and side-by-side plan comparison present the charge as a gap between where the customer is and where they could be, and that framing makes the upgrade decision easier for the customer to see than raw units or an isolated dollar figure do. That framing shift, from penalty to gap, is largely what separates an invoice that converts from one that provokes a complaint.
The second is alert timing and sequencing. An invoice by definition arrives after the fact, but the moment when a customer is most open to influence is before or at the point where they actually breach a plan limit. A sequence that sends automated usage alerts as a customer approaches their limit, then triggers an upgrade suggestion at the first significant overage, turns the overage moment into an active sales conversation. When the invoice itself arrives after that sequence has already run, it confirms a conversation the customer has already had, rather than ambushing them with one they were never part of.
The third is upgrade surface placement: whether the upgrade action lives on the invoice itself, inside the alert email, on the product dashboard, or only on a separate pricing page the customer has to go find. The farther the upgrade action sits from the charge that prompted it, the lower the conversion rate tends to be, because every additional click is an opportunity for the customer's attention to move elsewhere. The invoice, or the alert that precedes it, is the highest-intent surface a vendor controls at the overage moment, and placing the upgrade action anywhere else wastes that intent.
Atlassian's Rovo structure shows why these variables have to be designed together. Standard plans include 25 Rovo AI credits per user per month, and the Virtual Service Agent, available on JSM Premium and Enterprise plans, generates additional per-conversation charges once that monthly allowance is exceeded. An invoice that shows the per-conversation overage cost without also showing how many of the included per-user credits were consumed leaves the customer unable to judge whether the overage was reasonable or avoidable. The two numbers only make sense next to each other.
Structuring an overage invoice design experiment
The single most common error in testing overage invoice designs is optimizing for the wrong time window. A design change that lifts upgrade rate in the first thirty days but quietly damages retention over the following months is not a win, even though most testing setups will report it as one, because most testing setups stop measuring before the damage becomes visible.
A properly built experiment measures three things, at three different time horizons. The primary metric is voluntary upgrade rate within the first month following a customer's first overage invoice under the new design. The secondary metric is 90-day retention, measured for customers who received the new design against a control group who received the old one, regardless of whether either group actually upgraded. This metric exists specifically to catch the failure mode where a design converts upgrades by making the overage feel punitive: customers might upgrade in the short term simply to make an uncomfortable experience stop, and then leave a few months later once the discomfort is no longer fresh. The tertiary metric is the rate of billing disputes and support tickets tied to the overage invoice. A well-designed invoice should reduce that rate even among customers who never upgrade at all, because the charge is comprehensible on its own terms.
Because the 90-day retention window only begins once the last customer has entered the experiment, a properly measured overage invoice experiment runs three to five months from start to finish, not the thirty days many teams budget for. A team that reads results at the 30-day mark and calls the experiment complete is working from an incomplete result, no matter how clean that early number looks.
Experiment construction matters as much as the measurement window. Only one variable, consumption framing, alert sequencing, or upgrade surface placement, should change at a time. Changing more than one at once makes it impossible to say which change produced which result. Customer segmentation matters just as much: a customer who overages once behaves very differently from one who overages every billing cycle, and the design choice that moves a habitual overager may do nothing for an occasional one. Running a single experiment across a mixed population risks washing out real effects that would show up clearly in either segment alone.
Before any of this design work begins, a gross margin test comes first. If the overage is tied to an AI feature, the team has to know whether additional usage is actually net-positive for margin before it designs an upgrade path encouraging more of it. For AI-infused SaaS products, a revenue-to-inference cost ratio of 8:1 is the target floor, with 10:1 or higher considered the healthy benchmark. For AI-native products, the target floor is 4:1, with 5:1 or higher considered healthy. If that ratio cannot be calculated, pricing an upgrade path around increased usage is premature, because the vendor may be encouraging the customer toward behavior that loses money.
Invoice design across four AI and SaaS pricing architectures
The design requirements above do not translate into a single universal template, because the pricing architecture behind the invoice determines what information actually needs to appear. Four architectures currently in production among major vendors show the range.
Zendesk bills per automated resolution, defined as a request an AI agent closes without human escalation. Only Verified Resolutions, the highest-value tier, draw against a dollar-denominated allowance pool, while Assisted Escalation and Contained Resolution tiers are free and do not draw against that allowance. An invoice built for this architecture needs to show resolutions consumed against the allowance, distinguish clearly between automated and assisted outcomes, and surface the cost per resolution unit, so a customer can judge whether their current allowance tier is sized correctly for their actual usage pattern.
HubSpot's Breeze Customer Agent and Prospecting Agent, as of April 14, 2026, moved to a model charging $0.50 per resolved conversation, with a separate per-lead rate applying to Prospecting Agent. An invoice under this model should pair the resolution rate with the cost, so a customer who sees their agent resolving the large majority of conversations at a known per-conversation rate is looking at a straightforward expand-or-hold decision.
Credit-wallet architectures, where usage draws down a pool of fungible credits, present a different design problem, and Atlassian's Rovo structure is the clearest illustration available here. Standard plans include a monthly per-user credit allowance, and the Virtual Service Agent generates per-conversation charges once a customer exceeds that included allowance. The invoice has to show included credits consumed, overage conversations incurred, and the per-conversation rate as a single connected picture, because showing any one of those numbers in isolation, as a line item disconnected from the others, reproduces the exact confusion that undesigned overage invoices create everywhere else. The architecture changes. The underlying design obligation, that a customer must be able to reconstruct the charge from what the invoice shows, does not.


