Mato just raised pre-seed Read more

The $230K AI Inference Leak: A Hosts-Only Deep Dive

  • Aug 26, 2026
  • 19 min

Show notes

What the episode covers

Intercom Fin’s AI pricing shift shows why a hidden 23% inference leak can quietly erase B2B SaaS ARR gains, and how tightening cost-per-action before scale protects growth and acquisition economics. This Wednesday episode for Week 35 of 2026 breaks down why inference spend behaves differently from normal COGS, how it pressures gross margin even as revenue climbs, and what that means for founders sitting between $200K and $5M ARR.

In this episode, Derek and Elena unpack how inference costs stay stubborn at scale and why two-thirds of AI-heavy SaaS companies are already feeling margin compression. They walk through buyer pushback and vendor repricing pressure, how outcome-based models like Fin’s change the relationship between unit cost and unit price, and how to work from true cost-per-action up instead of guessing at AI feature pricing. They close with a concrete repricing checklist so you can adjust before enterprise buyers force caps that stall your growth and acquisition roadmap.

Host: Derek and Elena, ARR Autopsy.

Timeline

In this episode

7 moments worth skipping to. The timecodes match the player above.

  1. 0:11Introduction
  2. 1:29The Leak: Why Inference Costs Don't Shrink With Scale
  3. 5:18The Buyer Backlash: Uber's Cap and the Quiet Vendor Repricing
  4. 8:36The Case Study: How Fin Rebuilt Its Pricing (No Guest This Week)
  5. 12:13The Playbook: Pricing From Cost-Per-Action Up
  6. 15:24The Checklist: Repricing Before Buyers Do It For You
  7. 18:23Outro

Quick answers

Straight from the episode

The questions this one settles, without the listen.

Why aren’t AI inference costs shrinking as companies scale?
The episode explains that inference typically grows from about one-fifth of AI costs pre-launch to the low-20% range at scale and then plateaus instead of shrinking. As products scale, companies get better at routing and caching, but they mostly manage around inference rather than eliminating it, so it remains a stubborn, ongoing cost next to talent and infrastructure rather than fading with volume.
How are AI buyers pushing back on unpredictable or usage-based pricing?
Buyers are increasingly wary of usage-only pricing that can explode with successful adoption. Some are imposing hard budget caps on AI tools, while others are re-evaluating vendors whose flat pricing no longer covers real usage. Across the board, there’s growing demand for more predictable, transparent models that don’t turn product success into a budget crisis.
What’s the core pricing lesson from the Fin AI case study discussed in the episode?
Fin’s shift toward outcome-based pricing showed that aligning price with a clear unit of customer value and tracking that unit’s economics can drive strong expansion as usage grows. The hosts contrast this with reactive caps from buyers, arguing that proactive, value-tied pricing gives vendors more control over both margins and customer trust.
How does the episode suggest smaller SaaS founders should start fixing AI pricing and margins?
The hosts recommend starting from the real cost per action—what it costs to deliver a single AI outcome—then building pricing up from there. Rather than treating margin issues as an intractable ‘AI tax,’ founders are urged to pick a first unit of value, quantify its cost, and use that as the basis for experimenting with more resilient pricing models.
What common mistakes are causing AI features to destroy SaaS margins?
Drawing on Conception Labs’ work, the episode highlights patterns like underestimating true per-action costs, bolting AI onto old seat-based pricing without adjustment, and delaying pricing changes until buyers push back. These missteps lead to invisible margin leaks that only become obvious once AI usage scales across the customer base.
What is the main takeaway for SaaS teams adding or scaling AI features?
Treat AI as a unit-economics problem, not just a product feature. The hosts argue you need to understand and regularly revisit cost per action, align pricing with the value of those actions, and proactively reprice before customers impose their own caps or demand concessions that lock in weak margins.

Transcript

The full conversation

Every word of the episode, 3,179 of them, in the order they were said.

Read the transcriptHide the transcript

Derek SimmonsTwo hundred and thirty thousand dollars.

Elena ReyesOkay. No hello, no nothing. Just out of every million bucks of AI revenue, two hundred and thirty grand walks straight out the door to inference.

Derek SimmonsThat's almost a quarter.

Elena ReyesTwenty-three percent, yeah. Welcome to ARR Autopsy. I'm Elena.

Derek SimmonsI'm Derek. And quick heads up before we go further, no founder in the chair with us this week.

Elena ReyesJust us, the data, and a whole lot of margin pain.

Derek SimmonsSaaStr's breakdown of ICONIQ's numbers is where that twenty-three percent comes from, and it's not some one-off outlier.

Elena ReyesAnd wait until you hear what Uber did about their bill.

Derek SimmonsThere's a cap involved, a very specific one.

Elena ReyesPlus, a company that actually fixed this rebuilt their whole pricing around it.

Derek SimmonsRight. And that's the thing. This number doesn't act like a normal cost line at all.

Elena ReyesWhat do you mean?

Derek SimmonsI mean, most costs shrink as a percentage when you scale. This one doesn't. It gets worse.

Elena ReyesOkay, so this is the part that actually keeps founders up at night.

Derek SimmonsYeah. Let's get into exactly why. So if it doesn't shrink with scale, walk me through what SaaStr actually found when they cracked open the ICONIQ chart.

Speaker 3Okay, so get this. Before a product even ships, inference is already eating about a fifth of AI product costs. Twenty percent pre-launch.

Elena ReyesAnd after it scales?

Speaker 3It climbs. Talent sits around twenty-six percent of the cost stack at scale, infrastructure is about seventeen, and inference holds stubborn around twenty-three percent instead of Shrink with Scale.

Elena ReyesWait, so the thing that's supposed to get cheaper as you grow customers is the one line that doesn't cooperate?

Speaker 3That's the mechanism. Every other line you get leverage: more customers, same support team, same office. Inference is metered per call, so more usage is just more usage.

Elena ReyesOkay, but then explain this to me because ICONIQ's own numbers seem to contradict that. Their twenty twenty-six report has AI product gross margins going up forty-five percent in twenty twenty-five, projected fifty-three percent this year.

Speaker 3Right. That's the part people misread.

Elena ReyesBecause if the leak's structural, how are margins improving industry-wide at the same time?

Speaker 3Two-thirds of the companies in that report say they've improved their per query economics through routing and inference management, not by making inference itself cheaper.

Elena ReyesSo they're not draining the leak.

Speaker 3They're building around it. Better routing, smarter model selection, caching repeat calls. The inference tax doesn't go away, they just get more output per dollar of it.

Elena ReyesThat's a different game than the one most founders think they're playing.

Speaker 3Completely different, and a lot of them haven't caught up.

Elena ReyesGive me a concrete example of what that routing actually looks like day to day.

Speaker 3Say a support query comes in. A simple one gets routed to a cheaper, smaller model, and only the genuinely hard cases get sent to the expensive one. You're not cutting inference. You're not wasting your expensive inference budget on questions that didn't need it.

Elena ReyesSo it's not one big switch, it's a hundred small routing decisions stacked up.

Speaker 3Exactly. And caching's the same idea. If five customers ask basically the same question, you don't wanna pay the full inference cost five separate times.

Elena ReyesHow many are we talking?

Derek SimmonsUh, Conception Labs pricing guide put out in June says eighty-four percent of SaaS companies have already lost margin to AI costs.

Elena ReyesEighty-four percent? That's basically everybody.

Derek SimmonsAnd some of them didn't lose it gradually. The guide says some watched gross margins go negative overnight.

Elena ReyesNegative, like the AI feature is actively costing them money on every single customer who uses it.

Derek SimmonsEvery call. Flat-seat pricing was built for a world where marginal cost was basically zero. Add a heavy inference feature under that model, and you're subsidizing usage you can't see coming.

Elena ReyesSo here's what I keep landing on. The winners in that ICONIQ data aren't the ones who found some trick that makes inference cheap. They're the ones managing everything around it.

Derek SimmonsYeah. Nobody's beating the twenty-three percent. They're beating the other seventy-seven.

Elena ReyesWhich is a totally different problem to solve than what most founders think they're solving.

Derek SimmonsMost founders think they have a model cost problem.

Elena ReyesAnd they actually have a pricing architecture problem.

Derek SimmonsExactly the distinction.

Elena ReyesOkay, so if that's true on the seller side, founders quietly eating this, what happens on the buyer side? Because buyers aren't stupid. They see their own bills going up too.

Derek SimmonsOh, they've noticed. They've very much noticed.

Elena ReyesSo what are they doing about it?

Derek SimmonsUber's the one that made it public.

Elena ReyesSo the leak's structural, fine. But buyers aren't just sitting there eating the bill. Uber didn't.

Speaker 3What did Uber do?

Elena ReyesTechCrunch confirmed it citing Bloomberg. Uber blew through its entire annual AI budget in four months. Four months, Derek.

Speaker 3On what? Coding tools?

Elena ReyesCursor, Claude Code, agentic coding software. So now every engineer gets capped at fifteen hundred dollars a month per tool, and it's tracked on an internal dashboard.

Speaker 3A dashboard. So somewhere there's a spreadsheet watching every engineer's token habit?

Elena ReyesBasically, yeah. And it's not a suggestion, it's a hard ceiling.

Speaker 3Okay, fifteen hundred a month sounds like nothing until you stack two tools.

Elena ReyesThat's exactly what Simon Willison did the math on. He calculated an engineer maxing two tools at the cap lands around thirty-six thousand dollars a year.

Speaker 3Thirty-six grand on top of salary.

Elena ReyesRoughly eleven percent of a typical engineer's total pay by his estimate.

Speaker 3So the AI tooling is basically a second junior hire's worth of comp.

Elena ReyesAnd Willison's take wasn't outrage. It was, "This is a rational policy response to overspending." Uber isn't being cheap, they're stopping the bleeding.

Speaker 3Right. Which means every enterprise buyer watching that headline is now asking their own vendors the same question.

Elena ReyesAnd this is where it stops being an internal Uber story and starts being everyone's pricing problem.

Speaker 3Because if the buyer caps their own usage, they're going to cap what they'll pay you too.

Elena ReyesThere's a post going around. Someone made the point that usage-based pricing sounds fair on paper. Use more, pay more.

Speaker 4Sounds reasonable to me.

Elena ReyesExcept flip it around. It means you get charged more precisely at the moment your product becomes indispensable to them. That's not a bug from the buyer's chair. That's the scary part.

Speaker 4So the better it works, the more it costs them.

Elena ReyesRight. And for some buyers, predictable beats fair. They'd rather overpay a flat number than get a surprise invoice when their usage spikes.

Speaker 4That tracks with something I saw too, a post pointing out that as of July sixth, flat ChatGPT Enterprise seat pricing quietly stopped covering everything.

Elena ReyesWait, what changed? Workspace Agent and ChatGPT for Excel and Sheets got pulled out into token-based credits, so the seat price you thought you locked in, it doesn't cover the new stuff. So even the biggest player in the category is admitting flat pricing and AI features don't survive contact with each other.

Speaker 4If OpenAI can't hold a flat seat price on its own product, nobody's holding one on a startup's roadmap.

Elena ReyesIt's a good gut check for anyone still pricing like it's twenty nineteen.

Speaker 4Right. Flat forever isn't a strategy anymore. It's a liability waiting for its invoice.

Elena ReyesOkay. So now we've got two things sitting on the table. One, Inference eats your margin no matter how big you get. Two, your buyers are capping spend before you even get the chance to reprice.

Speaker 4So the real question isn't whether to reprice.

Elena ReyesIt's what a model looks like that survives both pressures at the same time, the cost side and the buyer revolt side simultaneously.

Speaker 4Which conveniently somebody's already tried.

Elena ReyesThere's an actual company that ran this exact experiment, restructured the whole thing around real Cost-Per-Action instead of guessing.

Speaker 4Let's open that file.

Elena ReyesOkay. So get this. There's an actual company that ran this exact experiment, and it's not some hypothetical. Intercom, their AI product, Fin.

Speaker 4Right. And we should say up front, this isn't a founder sitting across from us walking us through it. No guest this week. We're doing this off a published case study.

Elena ReyesA writeup on Mostly Metrics lays out the whole thing. Fin used to be priced like every other seat based add on, flat fee bolted onto your plan. Done.

Speaker 4Classic problem. You're selling a thing whose cost moves with usage and pricing it like it doesn't.

Elena ReyesSo they ripped that out. Fin moved to charging ninety-nine cents per resolution, not per seat, not per message, per actual outcome.

Speaker 4Wait. Per resolution? Meaning they only get paid when the AI actually solves the customer's ticket?

Elena ReyesThat's the structure the case study describes. Yeah. That's the whole thing right there. If your cost is inference per answer and your price is dollars per resolution, those two numbers are chained together. They move as one line. Compare that to Uber capping spend after the fact. Uber saying stop before you hit the ceiling. Intercom saying every unit of cost already has its matching unit of revenue baked in.

Speaker 4One's a brake pedal, the other's a transmission.

Elena ReyesSure. I'll take it. And it apparently worked. The case study says net revenue retention hit one hundred and forty six percent as Fin scaled from a million to a hundred million in ARR.

Speaker 4Hold on. A hundred x the revenue on that one product?

Elena ReyesThat's what's reported. And one hundred and forty six NRR means existing customers were spending more over time, not just staying flat.

Speaker 4Because the more resolutions they ran, the more they paid, and presumably the more value they were getting since it's outcomes, not seats sitting idle.

Elena ReyesRight. Which is the trap Flat-Seat pricing never solves. You can't expand revenue on a static seat count once the product actually starts working harder.

Speaker 4Now I wanna flag something because I went looking for the other number I actually wanted.

Elena ReyesThe margin swap.

Speaker 4Yeah. Before and after gross margin company wide. The piece doesn't have it. We get the price point. We get the NRR. We don't get a clean margin percentage stapled to it.

Elena ReyesSo we shouldn't treat this like a locked one for one margin trade.

Speaker 4No. It's directional. It tells you the pricing model survived and grew. It doesn't hand you the exact Leak that got closed in dollars.

Elena ReyesStill, Ninety-nine cents a resolution tied straight to the cost of generating that resolution. That's not a coincidence. That's design.

Speaker 4Design is the word. Somebody sat down and matched a unit of price to a unit of cost instead of matching price to a seat that has nothing to do with what the model's actually doing under the hood.

Elena ReyesAnd it's worth saying this only works because they know that ninety-nine cent number cold. They didn't guess at it.

Speaker 4Right. You can't price a unit of value if you don't actually know what the unit costs you first.

Elena ReyesSo if you're a founder listening and you don't have Intercom's scale or Intercom's data team, what do you actually do with this on a Tuesday?

Speaker 4That's the real question because most people can't just flip a switch to outcome pricing overnight.

Elena ReyesSo walk me through it. If I'm running a fifteen person SaaS company and I want Fin's math without Fin's headcount, where do I even start?

Speaker 4That's the real question because most people can't just flip a switch to outcome pricing overnight. So walk me through it. If I'm running a fifteen person SaaS company and I want Fin's math without Fin's headcount, where do I even start?

Elena ReyesSo if outcome pricing is the destination, the on-ramp is a guide from Conception Labs that just came out. The whole pitch is build up from your real cost per action, then price on top of that, not down from a seat price you made up in a spreadsheet. Exactly backwards from how most of us learned pricing. They lay out five different models: usage tiers, credit packs, hybrid seat plus metered, outcome-based like Intercom, the whole menu with actual companies doing each one. Which one's the training wheels version for somebody who's never touched their COGS line? Hybrid. Seat gets you predictability for finance. A metered tier on top catches the power users before they eat your margin. That's basically Uber's cap just baked into the product instead of bolted on after a budget blew up. Same dashboard logic earlier in the story. Okay, walk me through the actual conversation. A customer's paying the flat seat, they're about to hit the metered tier. What do you say to them? You don't frame it as a penalty. You say, "Your usage tells us you're getting more value out of this than the average account. So here's a tier that scales with that value instead of capping it." And if they push back? Then you show them the per action cost next to what they're actually generating from it. Nobody argues with their own ROI math. Right, because the alternative is Uber's version, cap first, explain never. That same Conception Labs piece also runs through the mistakes that quietly wreck margins. Companies skip the cost per action step entirely, or they price the feature once at launch and never touch it again while usage patterns shift underneath them. So the guide isn't just, "Here's the model," it's, "Here's how you'll screw it up." Pretty much. It's basically a highlight reel of every mistake we've already made on this show, just labeled clearly this time. Cheaper to read about it here than live it on your own P&L. And this isn't some isolated fire drill either. SaaS Mag's been tracking this as a full on twenty twenty-six trend, gross margins getting squeezed across public SaaS companies, not just the AI native ones. Yeah. This is the whole category adjusting at once. Which brings us right back to where we opened the show.

Speaker 5The two-thirds of companies actually managing their per query economics instead of just eating the hit.

Elena ReyesThey're the ones who did some version of this work, found the real cost per action, built a structure around it before a customer or a Slack thread forced their hand.

Speaker 5The founders who wait get Ubered. The ones who move first get to write the terms.

Elena ReyesSo if you're listening and you haven't run this yet, pull your actual inference cost per customer this week, not the blended average, the real number. Compare it against what that customer pays you. If the gap's closing, you've got the same leak we opened with.

Speaker 5And if it's already gone negative on a segment, that's not a future problem.

Elena ReyesThat's a Tuesday morning problem. Go build the tier before someone else builds the cap for you. So we've got the leak, we've got the buyer backlash, we've got a working proof point. What do you actually do with that Monday morning? Okay, three things. Write them down. Bossy. First, go find your real cost per action, not your blended margin, not your average deal size. What does one query, one resolution, one generation actually cost you fully loaded? And most founders don't have that number. Most founders have never pulled it. That's the whole problem in one sentence. Second? Second, once you know the real cost, set a floor and a ceiling. Floor protects your margin on the light users. Ceiling protects you from the one customer who runs your model into the ground. A cap with a name on it, not a surprise on an invoice. Right. And third, script the renewal conversation before a buyer hands you their own version of that cap. Because that conversation is coming either way. It's coming either way. You just get to pick whether you're the one talking or the one reacting. Okay, here's a stat I actually like for this. ICONIQ's own twenty twenty-six report says companies are now blending an average of one point seven pricing models. One point seven. So basically everybody's running two structures stitched together. Which tells you hybrid isn't the weird experimental thing anymore. It's just what the companies ahead of this are already doing. Right. It's not the exception, it's the baseline now. And ICONIQ's same report is where that two-thirds figure comes from. Two out of three companies actually moved their unit economics through routing and Inference management. Which gives you a benchmark. If you reprice and your unit economics don't budge at all, you didn't do the thing, you just relabeled the invoice. You wanna land somewhere in that two-thirds. Not everybody gets there, but that's the group you're trying to join. So audit the real cost, cap it top and bottom, script the talk track. That's the whole checklist. And it all traces back to that number we opened the show with, Two hundred and thirty thousand dollars walking out the door on every million in AI revenue. Yeah, that number doesn't move because we said it twice. It doesn't, but the rule underneath it does apply every time. Reprice before your buyers do it for you. Uber already showed everybody what that looks like from the buyer's side.

Speaker 5Exactly. You'd rather be the one setting the cap than the one reading about it in a cap someone else imposed on you.

Elena ReyesSo don't wait for your Uber moment. And to be clear, that's not a five-minute exercise, but it's not a quarter-long project either. A few hours with your billing data and your Inference logs, honestly. No, no, no. Go find your Cost-Per-Action this week. Actually pull the number. That's the assignment. All right, I think that's everything we've got for this one. Elena's gonna take us out with where all this actually lands. So if you're grabbing one number out of this whole hour, no guest, no origin story, just us and the math, mine's simple. Inference doesn't shrink with scale, so your pricing has to do the work your usage curve won't.

Speaker 5Mine's shorter than that.

Elena ReyesGo.

Speaker 5Reprice before your buyers do it for you.

Elena ReyesThere it is.

Speaker 5That's the whole thesis of this show, honestly. We just spent this whole episode proving it with somebody else's balance sheet.

Elena ReyesNo guest in the room today, and it still held up.

Speaker 5Better, maybe. Nobody was defending their own numbers.

Elena ReyesFair.

Speaker 5All right, here's the ask. If you know a founder staring at their AI margin line right now, send them this episode. Not the highlights, the whole thing.

Elena ReyesThey need the math, not the vibes.

Speaker 5Exactly. Subscribe on YouTube or wherever you get your podcasts, and if you've got thirty seconds, leave us a review. It's how other operators find this stuff.

Elena ReyesWe read every one.

Speaker 5We do, even the mean ones.

Elena ReyesEspecially the mean ones.

Speaker 5Go check your Cost-Per-Action before your next renewal. That's the whole ask.

Elena ReyesSeriously, that's the one homework assignment we're both giving out this week.

Speaker 5Do it before the renewal, not after.

Elena ReyesSee you next week.

More episodes

Keep listening

Other episodes of ARR Autopsy, newest first.

All episodes of ARR Autopsy

Sources

Where this came from

9 reports behind the episode. Every one of them opens where it was published.