AI SDR vs. Human SDR: One Founder's 90-Day Controlled Test
Show notes
What the episode covers
This week's guest reached $1.4M ARR with two human SDRs before running a controlled 90-day AI versus human outbound test on identical lists — and the results reframed the entire cost argument once show rates entered the equation. This is the Wednesday episode for Week 23 of 2026.
- Why AI-booked meetings show at 40-60% versus 70-85% for human-booked meetings, and how that single gap inverts the cost-per-meeting math in B2B SaaS outbound
- The SQL conversion rate spread — 15-20% for AI versus 25% for humans — that compounds the show rate problem and inflates true cost per qualified opportunity
- How the hybrid fix created three operational failures: CRM data decay, a burned warm domain, and no handoff rule between AI and human reps
- The three conditions required before building hybrid infrastructure across the $200K to $5M ARR growth window, and the one weekly metric that signals system degradation
Derek Simmons and Elena Reyes, ARR Autopsy — covering acquisition, pipeline math, and B2B SaaS revenue decisions that hold up when someone finally checks the numbers.
Timeline
In this episode
7 moments worth skipping to. The timecodes match the player above.
- 0:11Introduction
- 2:06The Number That Made Us Book This Episode
- 4:30$1.4M ARR and a Broken Outbound Motion
- 7:41The Scoreboard: 90 Days of Data, Head to Head
- 11:22The Hybrid Fix: One Workflow Change That Changed the Math
- 14:50What the Dashboard Looks Like Now and What They'd Do Differently
- 18:14Outro
Quick answers
Straight from the episode
The questions this one settles, without the listen.
- What show rates did AI-booked meetings achieve compared to human-booked meetings?
- In the 90-day test, AI-booked meetings had show rates of 40-60%, while human-booked meetings ran 70-85%. That gap inverted the cost-per-meeting math once someone divided total cost by actual shows at the end of the quarter.
- How does the SQL conversion rate differ between AI and human SDRs?
- AI-booked meetings converted to qualified opportunities at 15-20%, compared to 25% for human-booked meetings. Combined with the lower show rate, this compounded the cost disadvantage of the AI outbound stack.
- What is the 30-minute handoff rule for a hybrid AI and human SDR model?
- The fix was a strict SLA: the AI owns the first five touches, and any reply containing a question mark or a pricing signal gets routed to a human within 30 minutes. This recovered show rates to the mid-seventies and cut overall SDR spend significantly.
- What are the three conditions a company should meet before deploying a hybrid AI SDR setup?
- ACV must be under $20K, the company needs a specific ICP with distinct outbound campaigns, and there must be a dedicated manager committing at least 10 hours per week to oversee the system.
- What single metric should founders track weekly to know if their AI SDR system is healthy?
- Revenue per held meeting is the weekly dashboard number that signals whether the hybrid system is working or quietly degrading.
- Should founders under $500K ARR use an autonomous AI outbound stack?
- No. The episode advises founders below $500K ARR to start with a human SDR augmented by AI tools rather than building a fully autonomous outbound stack, since the infrastructure and management overhead of a hybrid model requires more revenue base to justify.
Transcript
The full conversation
Every word of the episode, 2,846 of them, in the order they were said.
Read the transcriptHide the transcript
Derek SimmonsForty to Sixty percent. That's what your AI SDRs show rate is like. Your human reps, Seventy to Eighty-five. Nobody ran the math
Elena ReyesMm
Derek Simmonsuntil
Elena ReyesMm.
Derek Simmonsend of quarter.
Elena ReyesWait, wait, wait. The whole cost argument flips once you divide by show rate.
Derek SimmonsCompletely flips. And that is exactly where today's episode starts.
Elena ReyesWelcome to ARR Autopsy. I'm Elena Reyes.
Derek SimmonsDerek Simmons, we dissect the revenue moves that worked. The ones that blew up and the ones that looked great until someone finally checked the actual math.
Elena ReyesToday's guest hit 1.4M ARR with two human SDRs, pipeline that looked fine on paper, and a nagging feeling something was wrong. So they ran a real 90-day AI versus human test on identical lists.
Derek SimmonsControlled experiment. Same lists. That detail matters.
Elena ReyesRight. And the scoreboard is a split verdict. AI won on volume and follow- And follow-up consistency lost on reply quality, deliverability, and cost per qualified opportunity once SQL conversion rates hit the picture.
Speaker 3Mm-hmm.
Derek SimmonsOkay, so get this: the hybrid fix that solved the quality problem created three new fires. Resentful human SDRs letting CRM data rot, a warm domain burned by volume, no handoff route.
Elena ReyesWow. Love when the solution has its own bugs. Classic. And we close out with the decision framework. 200K ARR to 5 million. What conditions have to be true before you build hybrid infrastructure and the one weekly metric that tells you if the system is quietly degrading?
Derek SimmonsSo what does this mean for the person listening right now? Probably a lot. Let's get into it. The cost per meeting number looked incredible on paper, and then someone divided by show rate.
Elena ReyesOh no, what happened when they divided?
Derek SimmonsThe math inverted. Here's the setup. Founder at 1.4 million ARR, two human SDRs, one AI SDR agent running in parallel for 90 days. Identical ICP lists, identical copy briefs. Clean test.
Elena ReyesOkay, so actually controlled. That matters.
Derek SimmonsIt really does because the AI looked great right up until Q2. two. Cost-per-meeting, way down. Volume, way up. Salesmotion published benchmarks showing AI SDRs deliver outreach of 20 to 60 percent of the cost of a human. The headline is real.
Elena ReyesYeah, the vendor deck math always checks out, until it doesn't.
Derek SimmonsUntil it doesn't. So here's where it gets good. The show rate on AI-booked meetings were running 40 to 60 percent. Human-booked meetings, 70 to 85 percent.
Elena ReyesWait, so you're telling me nearly half the AI meetings were no-shows?
Derek SimmonsAt the bottom of that range, yes.
Elena ReyesWow.
Derek SimmonsYour AEs calendar fills up, looks totally healthy, and roughly half of the slots are ghosts.
Elena ReyesThat is a silent budget leak, the kind that doesn't show up until someone pulls the actual held meeting number.
Derek SimmonsWhich nobody pulled until end of quarter.
Elena ReyesRight, right, right. And that's the whole thing. Cost-per-meeting is a vanity stat if you don't divide by show rate. Great. You're really buying cost-per-held-meeting. Exactly. And once you do that math, the gap between AI and human narrows a lot faster than the pitch deck suggested. So help me stress test this for a second. Was the show rate problem the AI or was it the ICP, the copy, the sequence? Because those feel like separable problems. Great question. And honestly, that's what the next 90 days became. But to even get there, first you need to understand where this founder was. was before the test started. Right. What was the actual state of the business, the ARR, the team? What had already broken?
Derek SimmonsBecause 1.4 million ARR with two human SDRs and a pipeline problem is a very specific situation. And the reason they ran a controlled test instead of just switching everything over, That's worth sitting with.
Elena ReyesYeah, why not just flip the switch?
Derek SimmonsThat's the question. So here's the actual situation before the test started. At 1.4 million ARR, the founder had two human SDRs. Fully loaded, each rep was running about 130K a year. That's salary, benefits, tools, the whole stack.
Elena ReyesAnd what was the meeting to opportunity rate? Like, what were those reps actually producing?
Derek SimmonsThat's the thing. On paper, solid. About 15 booked meetings per rep. rep per month. But the underlying pipeline felt thin. SQLs weren't closing at the rate the ARR growth required.
Elena ReyesOkay, but let me stress test that a little bit. Was the pipeline actually thin or were they looking at the wrong number? Because 15 meetings sounds fine on the surface.
Derek SimmonsRight. And that's exactly what they couldn't answer cleanly. Nobody had broken it down to cost per qualified opportunity yet. They were watching cost per meeting. Which,
Elena ReyesSo
Derek Simmonsas we covered, is where the trouble starts.
Elena Reyesthe reps were hitting activity targets. The math just hadn't been done on what those meetings were actually worth downstream.
Derek SimmonsExactly. And then you layer on the tenure problem. According to Bridge Group data, your average SDR stays about 14 to 16 months. Three of those months are ramp. So you're getting roughly a year of peak output, then you're restarting. A year.
Elena ReyesYou pay to recruit, pay to ramp for a quarter. Get twelve months of real production then do it all over again.
Derek SimmonsI've seen this movie
Elena ReyesYeah.
Derek Simmonsbefore, and it's brutal on pipeline predictability. Every time someone walks, you're three or four months from getting back to steady state.
Elena ReyesSo what did they already tried before reaching for AI? Because I doubt the first instinct was "let's automate this.
Derek SimmonsOh no, they'd tried the standard playbook: third party leads lists, a fractional SDR, even ran one rep on a pure cold call motion for a quota. For a quarter, nothing got the meeting to SQL rate they needed.
Elena ReyesAnd the fractional SDR--specifically what happened there?
Derek SimmonsThree months, eight meetings booked, two showed. That's your whole problem in miniature.
Elena ReyesOh, man, two showed. So, by the time they're looking at AI SDRs, they're not doing it because a vendor had a good deck, they're doing it because the human motion keeps resetting, and they can't afford another ramp cycle eating into runway. Which is why the controlled test matters so much they didn't just flip the switch.
Derek SimmonsNo!
Elena ReyesAnd this is where it gets good. Same lists, same ICP, same sequences handed to both the AI and the two humans.
Derek Simmons90 days, clean split. The whole point was to make sure the data would actually mean something at the end.
Elena ReyesThat's the setup that makes this worth talking about, because without the controlled structure, you're just getting a vendor case study.
Derek SimmonsExactly. And 90 days later, they had a spreadsheet. That's where we're going next. The actual numbers, side by side, what the AI booked, what the humans booked, and what happened when someone finally checked who actually showed up. Thanks for watching. So the scoreboard is open; ninety days, identical lists. Walk me through the raw numbers first.
Elena ReyesOkay, so the AI booked more meetings, full stop. Volume was up roughly three to four times what the human reps produced over the same window. And the follow up cadence? The AI hit every single touchpoint. Humans skipped follow ups constantly.
Derek SimmonsShocker.
Elena ReyesRight? But here's where it gets interesting. Auto Interview AI's benchmarks have human SDRs responding to inbound leads in 42 to 47 hours on average. The AI sub 60 seconds every single time.
Derek SimmonsThe founder tracked that Explicitly?
Elena ReyesExplicitly. And on inbound, that speed advantage translated directly to booked meetings. No debate on that one.
Derek SimmonsOkay, so volume up, speed up. Now give me the number that actually matters.
Elena ReyesCost per qualified opportunity-that's the verdict number, not cost per meeting-and when you divide all the way down to SQL, the gap narrows dramatically.
Derek SimmonsBecause the show rate hurt them earlier but the SQL conversion rate hurt them again at the next stage.
Elena ReyesExactly. According to data from Apollo, AI-booked meetings convert to fifteen to twenty percent qualified opportunities. Human-booked meetings run around twenty five percent. So you book more, fewer show, and fewer of those convert. That math compounds fast.
Derek SimmonsThat's the volume up, quality down pattern showing up at every funnel stage, but let me ask the uncomfortable one—what did the AI actually say when a prospect pushed back on pricing?
Elena ReyesOoh, that's where it got ugly. The auto reply handling was... not good.
Derek SimmonsDefine not good.
Elena ReyesThe AI responded to a pricing objection... objection with a generic feature dump. No context, no acknowledgement of the objection, just a wall of capability bullets. The prospect replied back asking who they were actually talking to.
Derek SimmonsAnd that's not an edge case. According to Autopsy AI's research, AI handles the top 10 common objections acceptably, but humans handle the other 90% that need real thinking. Pricing questions live in that 90%.
Elena ReyesAnd the deliverability problem compounded everything. AI emails got spam flagged at 8% versus 3% for human-written outreach on a 100,000 email analysis.
Speaker 3Wow.
Elena ReyesThat's nearly three times the spam rate on the same domain.
Derek SimmonsSo the reply gap was narrowing, which is the story the vendors lead with, but the domain was quietly burning underneath it.
Elena ReyesThat's the part that decided the verdict. Not cost per meeting, not even show rate. It was the cost per qualified opportunity once you factored in conversion and domain damage.
Derek SimmonsBridge Group's data puts hybrid pods at 54% lower cost per qualified opportunity versus human only, but pure AI, worse.
Elena ReyesWorse, the founder's own spreadsheet matched that directionally. Pure AI motion failed at quality at every downstream stage.
Derek SimmonsAnd that is exactly the question the next 90 days had to answer. Answer, if the pure AI motion broke at quality, what's the specific fix because they didn't shut it down?
Elena ReyesNo, they didn't, and the hybrid structure they built is the piece worth unpacking.
Derek SimmonsSo the scoreboards said hybrid, but building the hybrid, That's where it got messy.
Elena ReyesFast, because the first problem wasn't the AI, it was the humans.
Derek SimmonsOf course it was.
Elena ReyesThe two SDRs on staff saw the AI tool as a threat, and when people feel threatened, they're not exactly rushing to clean up the CRM data that feeds the system they resent.
Derek SimmonsThis is the part that never shows up in vendor case studies. Dirty CRM means the AI targets worse over time. Over time, Apollo actually flags this explicitly: duplicate records, missing job titles, stale company data, all degrade AI output quality. The garbage in problem compounds at volume. And volume was exactly what made it worse. The AI was running cold outbound at scale on their main sending domain before anyone had a real warm-up protocol. DevCommX has written about this. Sender reputation, once damaged, can take months to recover.
Elena ReyesCover.--You don't notice it happening until open rates are already in the floor.
Derek SimmonsSo the CRM's dirty, the domain's getting cooked and the humans are quietly routing around the system, three fires at once.
Elena ReyesClassic scaling debt: pay now or pay later. And they paid later.
Derek SimmonsOkay, so walk me through the actual fix, because I don't want vague process talk here. What specifically changed? The rule they wrote was simple enough to put on a sticky note.
Elena ReyesAI owns touches one through five and handles any inbound response within 60 seconds. The moment a reply comes in with a question mark or the word pricing or anything that signals real intent, it routes to a human within 30 minutes.
Derek SimmonsHold on, thirty minutes not sixty seconds?
Elena ReyesFor the handoff, Yeah. The AI fires an instant acknowledgement, buys the human time, but the human has to be in the thread within thirty minutes with actual context, not a template.
Derek SimmonsOkay, and did that move the needle on show rates?
Elena ReyesThat's fair to push on. The show rate improvement on hybrid booked meetings came about six weeks of running that rule consistently. They got back to the mid seventies range, which is closer to what the human human only meetings had been doing.
Derek SimmonsSo that one hand off rule basically closed the show rate gap
Speaker 4Yes.
Derek Simmons-the thing that was bleeding them out at the start of the episode.
Elena ReyesMost of it. The other piece was restructuring the pod-one human SDR per roughly two to three AI sequences running simultaneously. The humans' job title basically became reply triage and relationship escalation. No cold list work at all.
Derek SimmonsWhat did that cost versus the old model?
Elena ReyesOne human SDR running oversight on AI sequences costs somewhere in the $60,000 to $80,000 fully loaded range.
Derek SimmonsCompare that to the $260,000 they were burning on two fully independent human SDRs. The AI stack itself, a mid-range platform, was running around $900 to $1,500 a month on top.
Elena ReyesSo, real talk. The math works if the handoff rule holds. Break the rule, the show rates slip, and the whole cost argument falls apart.
Derek SimmonsExactly. The SLA isn't a nice to have, it's the load-bearing wall.
Elena ReyesThe system is only as good as the worst handoff week. Which sets up the real question, what does a founder do if they're watching this from $500,000 ARR, not $1.4M? Is this even worth building yet? So, what does all this mean for the person sitting at $500K ARR right now, genuinely wondering if they should build this whole hybrid pod?
Derek SimmonsHonestly, the honest answer is probably not yet. And here's why: at 500K ARR, you likely haven't locked in your ICP tightly enough to clone the playbook. The AI clones what works. If you're still figuring out what works, you're automating noise.
Elena ReyesThat's the thing nobody says in the vendor demo. They show you volume, they don't ask whether your messaging is actually proven.
Derek SimmonsOkay, but let me stress test that assumption with some specifics. The ACV question is real. The AI-SDR playbook piece puts it pretty bluntly: $1 to $20,000 ACV is the sweet spot. AI qualifies and books, humans close. Above $20,000, you want a human doing most of the work, with AI on first touch and research only.
Elena ReyesAnd the math backs that up: the product growth.blog piece also flags that if revenue per meeting from your AI pipeline is below 50% of your human pipeline, you need to fix segmentation before you add any volume. Which is exactly the trap our founder walked into in segment three: high volume, low SQL conversion. Those two numbers compounded badly. Turns out multiplying bad math by 10 just gives you more. More Bad Math Faster, at Scale So the decision framework I'd hand someone right now? Three conditions have to be true before you spend a dollar on AI SDR infrastructure: (one) your ACV is under twenty thousand; (two) your ICP is specific enough that you can write five distinct campaigns for distinct segments, not one campaign blasted at everyone; (three) you have someone who can actually manage the system.
Derek SimmonsMinimum ten hours a week.
Elena ReyesThat last one is the silent killer. The product growth to unplug playbooks cites SaaStr spending 15 to 20 hours weekly just managing their AI SDR deployment,
Derek SimmonsWow.
Elena Reyesand performance dipped when that person got busy with other work.
Derek SimmonsRight. This isn't a set it and forget it channel.
Elena ReyesReal talk for a second. If you're below 500K ARR and you're thinking about spinning up an AI SDR tool, The better move is probably a strong human SDR with AI augmentation on the research and sequencing side. You get the signal without burning your domain. And Landbase had data on this. Sales tech companies running AI-augmented human reps are hitting pipeline velocity gains without the deliverability risk of a full autonomous outbound stack.
Derek SimmonsShort pause.
Elena ReyesShort pause. One number to watch on your dashboard weekly, whatever stage you're at. you're at. Revenue per held meeting, not meetings booked. If that number's dropping while volume is climbing, the system is quietly degrading.
Derek SimmonsThat's the one metric that doesn't lie: cost-per-booked-meeting is vanity stat without it.
Elena ReyesI've seen this movie before: the dashboard looks great, pipeline feels thin, and Nobody divided by show rate until end of quarter.
Derek SimmonsWhich is exactly how this whole episode started, full circle.
Elena ReyesAnd that's your Autopsy right there. That's a wrap on this one. And honestly, Elena, if there's one thing from today that I keep coming back to, the calendar that looks full but half the slots are ghosted.
Derek SimmonsMm-hmm.
Speaker 3Right. Cost per meeting as a vanity stat. The moment you divide by show rate, the whole math flips.
Elena ReyesThat reframe alone is worth a listen: cost per held meeting. Write it on your whiteboard. And the fix wasn't ditch the AI; it was a 30-minute SLA and a clean handoff rule. Boring answer: works.
Speaker 3Boring answers usually do.
Elena ReyesIf this one saved you from a bad bet, do us a favor: share it with one founder who needs it.
Derek SimmonsSubscribe on YouTube or wherever you're listening and drop a review. That's what keeps real operators talking to us.
Elena ReyesWe'll see you next time on ARR Autopsy.
Speaker 3Take care, everyone.
More episodes
Keep listening
Other episodes of ARR Autopsy, newest first.
- AI Overview Tax: 82% Coverage, Half the ClicksSep 16, 2026 · 21 min
- AWS Marketplace: How 40% of SaaS Revenue Skips CFOsSep 9, 2026 · 21 min
- Procurement: The 8 Weeks Before 2027 Budgets LockSep 2, 2026 · 14 min
- The $230K AI Inference Leak: A Hosts-Only Deep DiveAug 26, 2026 · 19 min
Sources
Where this came from
23 reports behind the episode. Every one of them opens where it was published.
- AI SDR vs Human SDR: Real Cost, Performance, and When to Use Each (2026 Guide) | Auto Interview AIautointerviewai.com
- The AI SDR Playbook: What Actually Works in 2026productgrowth.blog
- AI SDRs vs Human SDRs: The Real ROI Comparison for 2026 | Salesmotionsalesmotion.io
- AI SDR Pricing 2026: Real Costs Explaineddevcommx.com
- How Do I Calculate SDR vs. AI Cost Savings? | Apolloapollo.io
- Outbound SDR Statistics 2025 | AI, Metrics & Performance Data - Sales Sosalesso.com
- 50 Key AI SDR Statistics You Should Know in 2026devcommx.com
- AI SDR Statistics 2026: 100+ Outbound Sales Data Pointsdigitalapplied.com
- 47 AI Calling Statistics Every Sales Leader Needs to Know in 2026 | Auto Interview AIautointerviewai.com
- AI SDR Real Performance: 100K Email Analysis 2026digitalapplied.com
- AI SDR vs Agency: Which Actually Books Meetings in 2026?prospeo.io
- Artisan AI - Crunchbase Company Profile & Fundingcrunchbase.com
- Artisan AI - Wikipediaen.wikipedia.org
- Artisan AI ARR Hits $5M | Artisan AI ARR Milestone | ARR Clubarr.club
- Artisan AI: $5M/yr, 120 customers, $46M raised, replaces ...linkedin.com
- Artisan raises $11.5M seed roundartisan.co
- Artisan raises $11.5M to deploy AI 'employees' for sales teams | VentureBeatventurebeat.com
- Artisan raises $25M series Aartisan.co
- Artisan Raises $25M Series A to Scale AI Sales Agents and Expand Product Linetheaiinsider.tech
- Best AI SDR Software 2026 (and Why You Might Not Need One)unifygtm.com
- How AI SDR Agents Boost Conversions by 70% (2026) | Landbaselandbase.com
- How Many Meetings Should an SDR Book? (2026 Data Inside)tamtotarget.com
- Your AI SDR that books meetingsaisdr.com
