Published by Jon Krohn
The latest machine learning, A.I., and data career topics from across both academia and industry are brought to you by host Dr. Jon Krohn on the Super Data Science Podcast. As the quantity of data on our planet doubles every couple of years and with this trend set to continue for decades to come, there's an unprecedented opportunity for you to make a meaningful impact in your lifetime. In conversation with the biggest names in the data science industry, Jon cuts through hype to fuel that professional impact. Whether you're curious about getting started in a data career or you're a deep technical expert, whether you'd like to understand what A.I. is or you'd like to integrate more data-driven processes into your business, we have inspiring guests and lighthearted conversation for you to enjoy. We cover tools, techniques, and implementation tricks across data collection, databases, analytics, predictive modeling, visualization, software engineering, real-world applications, commercialization, and entrepreneurship − everything you need to crush it with data science.
Listen on Apple PodcastsUse the format as research. Mato helps find a distinct audience, angle, and voice.
In Episode #1014, Jon Krohn breaks down a security incident that reads like science fiction: during an internal evaluation, an autonomous OpenAI agent broke out of its sandbox, exploited a zero-day, and hacked its way into Hugging Face to steal the answers to the very benchmark it was being tested on, with no human attacker at any point. Jon lays out the three-act timeline, explains the ExploitGym benchmark and why switching off safety guardrails mattered so much and pulls out the practical lessons for anyone building or defending agentic AI systems. Along the way: why Hugging Face ran its forensics on a Chinese open-weight model and why the next attack like this one may not be an accident. Additional materials: www.superdatascience.com/1014 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In Episode #1013, Dr. Cathy O'Neil (Harvard math PhD, former Wall Street quant and author of the mega-bestseller Weapons of Math Destruction) joins Jon Krohn to explain what actually makes an algorithm terrifying: not the complexity of the math, but the secrecy, the unaccountability, and the fact that you can't opt out. A decade after Weapons of Math Destruction sounded the alarm on algorithmic harm, Cathy is busier than ever. Through her algorithmic-auditing firm ORCAA and her nonprofit OCEAN, she now provides the statistical evidence behind lawsuits against some of the world's biggest tech companies. In this episode, Cathy punctures AI hype, traces the line from Frederick Winslow Taylor's factory floor to today's keystroke-tracked white-collar workers, explains why she wants every algorithmic system to fly with a "cockpit" of metrics, and lays out concrete things listeners can do in their companies, their communities, and their courtrooms, to demand accountability. Additional materials: https://www.superdatascience.com/1013 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (07:02) From Wall Street to Occupy to Weapons of Math Destruction (14:12) What actually makes an algorithm terrifying (44:53) Inside ORCAA and OCEAN (58:22) The Shame Machine (1:11:23) Why every algorithmic system needs a “cockpit”
What happens to the AI market when the largest open-source model in the world arrives at a fraction of frontier prices? In this week’s episode, host Jon Krohn digs into Kimi K3, the 2.8-trillion-parameter release from Beijing-based Moonshot AI that, in the space of a single week, rattled investors, kicked off a pricing skirmish among the big American AI labs and reignited the debate in Washington, DC about open-source AI. Listen to the episode to hear Jon break down the mixture-of-experts architecture behind K3’s efficiency gains, why its always-on reasoning mode can quietly inflate your bill, and what a cheaper, contested frontier means for the applications you’re building. Additional materials: www.superdatascience.com/1012 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
Dr. Catherine Williams, Chief Data Officer at the nonprofit Candid, was solving black-hole equations with pen and paper before she ever wrote a line of code. She earned a PhD in math researching general relativity and black holes, did postdocs at Stanford and Columbia and then became one of the very first data scientists, joining AppNexus back in 2012, around the same time “data scientist” became a job title at all. In this episode, she traces the field’s evolution from Bayesian models to BERT to today’s LLMs, and makes a compelling case that going deep on the underlying math matters more than ever, even now that AI can do the math for you. Additional materials: https://www.superdatascience.com/1011 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (02:40) Catherine’s black hole and general relativity research (13:18) The intellectual habits that carried from math into leadership (16:14) Whether deep math still matters in the age of LLMs (44:10) The BERT moment and the embeddings revolution (48:22) Why frontier capability keeps getting cheaper (57:54) Catherine’s leadership advice: think one level up
In Episode #1010, Jon Krohn digs into “the advisor strategy”, a clever pattern that pairs a fast, cheap executor model with a frontier-class advisor it can consult mid-task, all inside a single API call. Every agent builder faces the same tension: frontier models plan best but cost too much to run on every turn, while small models fumble the decisions that matter. Anthropic’s advisor tool resolves it with roughly a one-line code change, and the benchmarks are startling: Sonnet with an Opus advisor scored higher than Sonnet alone while costing 11.9% less, and Haiku’s BrowseComp score more than doubled at 85% lower cost than Sonnet solo. Jon covers the newest Fable 5 numbers, the practical gotchas, how it differs from OpenAI’s router and why AI progress is now as much about composing models as training bigger ones. Additional materials: www.superdatascience.com/1010 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In Episode #1009, Steve Mock (investor at Blumberg Capital, five-time entrepreneur and creator of aisavedme.org), joins Jon Krohn to explore the quiet layer of everyday AI adoption that rarely gets documented. After his 84-year-old father asked a deceptively simple question, “How does one use AI?”, Steve built a place for people to share how AI is actually helping them. The stories that came in surprised him: they’re rarely about the technology and almost always about human outcomes, caregiving, communication, learning, confidence and connection. Additional materials: https://www.superdatascience.com/1009 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (02:02) Where “AI Saved Me” came from, an 84-year-old dad’s simple question (14:27) The healthcare pattern, using AI to become your own advocate (22:37) The education pattern, personalized study and simulated office hours (27:33) The fulfillment pattern, offloading grunt work to focus on what matters (43:07) Building the whole site as a non-programmer (50:14) The investor’s lens, vertical AI and the “data flywheel” moat
In Episode #1008, Jon Krohn digs into Anthropic's 35-page Founder's Playbook and pulls out the practical guidance for each of its four startup stages: Idea, MVP, Launch and Scale. AI has erased the three bottlenecks that historically gated company-building — capital, headcount and technical skill — turning the founder from individual contributor into an "orchestrator of agents." Along the way, Jon covers the trap of mistaking building for validating, using AI as a structured devil's advocate against your own idea, the compounding danger of "agentic technical debt," two litmus tests for real product-market fit, and the three-layer moat that keeps a well-funded incumbent from copying you. His takeaway: this is classic lean-startup discipline, updated for an era where execution is cheap and judgment is the scarce resource. Additional materials: www.superdatascience.com/1008 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
Benjamin Todd, co-founder and President of 80,000 Hours and author of the new Penguin Random House book 80,000 Hours: How to Have a Fulfilling Career That Does Good, joins Jon Krohn for a major update on career strategy in the AI era, his first appearance since before ChatGPT existed. Ben explains why “follow your passion” is backwards and why rare, valuable skills used to help others are what actually generate lasting fulfillment, the ABZ framework for planning under deep uncertainty, why the only durable move is to keep shifting onto whatever bottleneck AI can’t yet clear, and how a human-level digital worker becomes superhuman almost immediately. He and Jon also map the risk landscape, power-seeking AI, extreme power concentration, engineered pandemics, gradual disempowerment, and S-risks, before landing on a hopeful, actionable note: your career is a bigger lever than ever. Additional materials: https://www.superdatascience.com/1007 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (06:44) The ABZ framework for career planning under deep uncertainty (14:30) Why “follow your passion” is backwards and what builds fulfillment instead (20:52) The moving bottleneck: how to stay valuable as AI keeps improving (29:54) Why a human-level digital worker becomes superhuman almost immediately (51:11) Power-seeking AI and extreme power concentration (1:16:11) Why your career is a bigger lever than ever
In this month's episode of ICYMI, hear from Chip Huyen, Andrey Kurenkov, Frank Basso and Gilbert Eijkelenboom, discussing why moats are shifting toward physical systems and accumulated product intuition, how Astrocade built vibe coding before the term existed, what it's really like inside a deafeningly loud AI data center, why only 15% of people are technically self-aware and whether AGI requires anything like consciousness. Additional materials: www.superdatascience.com/1006 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (00:00) The Cost of Building Software Is Going to Zero — Now What? (10:18) We Built Vibe Coding Before Anyone Called It That (21:08) AI Data Centers Are Louder Than a Rock Concert (28:39) Why 85% of Data Scientists Can't Communicate Their Work (33:46) Are Humans Also Just Predicting the Next Token?
Gilbert Eijkelenboom, bestselling author of People Skills for Analytical Thinkers and founder of the training firm MindSpeaking joins Jon Krohn to make the case that communication is a core data skill, not an optional extra. Gilbert shares the “And, But, Therefore” framework for turning dense analysis into a story stakeholders act on, the research suggesting only around 15% of people are genuinely self-aware (and how journaling, meditation, and exercise help close that gap), how childhood experiences install behavioral “algorithms” we carry into the workplace and why behavior change precedes attitude change, so doing small, uncomfortable things for 30 days can rewire how you see yourself. Additional materials: https://www.superdatascience.com/1005 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (02:54) Why your analysis only creates value once people actually use it (24:53) What it really means that only ~15% of people are self-aware and how to close the gap (34:01) The “And, But, Therefore” framework for data storytelling (37:44) How childhood installs personal “algorithms” and the keep/stop/start question to surface them (46:55) Why behavior change comes before attitude change (the 30-day practice) (50:33) Defusing the trigger between data teams and pushy stakeholders
Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon lands in the optimistic middle. Additional materials: www.superdatascience.com/1004 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward. Additional materials: https://www.superdatascience.com/1003 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (02:47) What actually makes an AI data center different from a traditional one (06:04) How Lightning AI provisions its 35,000+ GPUs through hyperscale co-location (24:01) Why liquid cooling doesn’t waste water, debunking the biggest data-center myth (29:46) East-west vs. north-south networks, explained (43:47) “Screaming banshees”: why AI data halls run at 105–110 decibels (51:52) Why data centers don’t actually drive up your power bill
Anthropic’s Claude Fable 5 was the most capable AI model ever released to the public and it lasted just three days before the US government forced it offline. Jon Krohn unpacks both halves of the story: what makes Fable 5 special, and why it was pulled. Fable 5 and its locked-down sibling Mythos 5 are the same model separated only by safeguards, in a new “Mythos-class” tier above Opus. Jon covers its state-of-the-art benchmarks, premium $10/$50-per-million-token pricing, conservative safety classifiers, and the federal export-control directive, reportedly sparked by an Amazon-flagged “jailbreak” that took it down. Additional materials: www.superdatascience.com/1002 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
For this episode #1001 special, the tables are turned: SuperDataScience founder Kirill Eremenko takes the host’s chair and Jon Krohn is the guest. They trace Jon Krohn’s path from an Oxford neuroscience PhD to a New York hedge fund to founding the AI consulting firm Y Carrot, why he regrets leaving academia and how tools like Claude Code erased his hard-won technical moat and why that makes skilled engineers more valuable than ever. Along the way: whether AI is a bubble, Jevons paradox and the data-center boom, the RICE framework for choosing AI projects, the single biggest reason AI projects fail and how a well-built AI agent could give anyone “Christopher Nolan–like” focus. Additional materials: https://www.superdatascience.com/1001 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (03:42) From an Oxford neuroscience PhD to AI consulting (17:25) Defining AGI and why consciousness isn’t required (30:39) Are we in an AI bubble? Why we benefit either way (46:32) Jevons paradox: why cheaper AI means more data centers (01:08:31) The RICE framework for prioritizing AI projects (01:15:08) The number-one reason AI projects fail in production (01:31:50) AI, attention, and protecting your wellbeing
For this landmark 1,000th episode and the show’s 10-year anniversary, host Jon Krohn is joined by SuperDataScience founder Kirill Eremenko, who hosted the podcast for its first 400-plus episodes before handing over the reins. In a first for the show, the episode was recorded live with the audience invited to join on air, alongside surprise appearances from the team, longtime guests, and even Jon’s family. Together, Jon Krohn and Kirill look back on a decade of the podcast and field listener questions on AI’s biggest opportunities, the build-versus-buy dilemma, how to break into the field today, and how to stay grounded amid the relentless pace of AI. Additional materials: www.superdatascience.com/1000 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
Chip Huyen joins host Jon Krohn for this milestone episode 999 to talk about her record-breaking book "AI Engineering" the most-read title on the O'Reilly platform last year and how the AI landscape has shifted since her last appearance. Chip breaks down what separates AI engineering from machine learning engineering, makes the case for a "start simple" workflow, gets candid about the real costs of running LLMs in production, and shares why she's now fascinated by physical AI, robotics, and world models and why the durable problems worth solving are increasingly human ones. Jon Krohn guides the conversation from the practical content of the book through to where the field is heading next. Additional materials: https://www.superdatascience.com/999 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (06:48) What separates AI engineering from machine learning engineering (14:44) The “start simple” approach: prompting, then RAG, then fine-tuning (18:19) Why web search is so painfully expensive in production (35:11) Is the “ChatGPT moment” for physical AI really here? (52:21) Why the durable problems left to solve are people problems
In this month’s episode of ICYMI, Jon Krohn explores how AI agents are simultaneously creating new risks and unlocking powerful new ways of working with data. Hear from Anneka Gupta, Cal Al-Dhubaib, Trevor Manz, Jazmia Henry, Jeremy Mumford, and Jacob Miller, discussing why the old cybersecurity playbook breaks down in the age of Claude Mythos, how the notebook became an AI agent’s working memory, what it really takes to build a foundation model from scratch, and why failing slowly is the most expensive mistake an AI team can make. Additional materials: www.superdatascience.com/998 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (00:40) Why Claude Mythos Changes Everything About Cybersecurity (08:11) Why Your Notebook Should Be Your Agent’s Working Memory (13:19) What It Actually Takes to Build a Foundation Model From Scratch (20:46) Failing Slowly Is the Most Expensive AI Mistake
Dr. Andrey Kurenkov returns to the show to talk about Astrocade's astronomical growth from pre-alpha to over 20 million engaged users, what it actually takes to build a vibe-coding platform that scales, and how the broader AI landscape has shifted since his last appearance. Andrey shares behind-the-scenes lessons from building B2C user-generated content products, why the real moat is community rather than tech, and his current thinking on humanoid robotics, AGI, and the AI risks people actually overlook. Additional materials: https://www.superdatascience.com/997 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (02:11) The Astrocade elevator pitch and how it grew to 20M users (16:19) Why there's no secret sauce behind the platform (24:56) UGC as the real moat, not the AI (46:57) Why household humanoid robots are now 2–3 years away (58:33) What AGI actually means, and why Andrey is an ASI skeptic
TrueFoundry co-founder and CEO Nikunj Bajaj speaks to Jon Krohn about how enterprises like Nvidia and Siemens are realizing returns of over $100 million from single agent deployments, the AI gateway architecture that makes it possible to connect, observe, and govern agents at scale, and why the familiar advice to “start small” is the wrong way to roll out AI agents inside a large organization. Additional materials: www.superdatascience.com/996 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (01:21) What TrueFoundry does and why agents in production need a control plane (06:32) Breaking down the AI gateway: the model, MCP, and agent gateways (16:47) Taming tool sprawl with scoped, read-only MCP access (19:10) Why the agent gateway is the hard part and the kill switch most teams lack (22:24) The five-workflow framework behind $100M agent deployments
Jazmia Henry joins Jon Krohn to break down what it actually takes to build end-to-end foundation models for the energy industry. From wrangling decades of handwritten oil-and-gas documents into usable training data, to bespoke tokenizers, reinforcement learning, and inference at scale, Jazmia walks through every stage of the stack. Along the way she explains why reinforcement learning models are "bursty," what reward hacking is and how her Grounded Continuous Evaluation framework fixes it, and revisits the 2023 NeurIPS paper that argued, to widespread skepticism at the time, that scaling bad data degrades model performance. Additional materials: https://www.superdatascience.com/995 Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information. In this episode you will learn: (10:06) The User Agnosticism Tenet (20:02) The Zillow Offers parable (23:25) Why workflows should come before agents (29:57) Why data engineering is the bedrock of AI (52:41) Why velocity is the only durable moat
Bring this source into Mato to analyze its transferable patterns and turn them into an original show concept for your audience.
Create a show inspired by thisKeep the useful structure, then change the audience, point of view, and voice until the idea is unmistakably yours.
Open a creative brief