Podcast charts
Published by Conor Bronsdon
AI is reshaping infrastructure, strategy, and entire industries. Chain of Thought is the podcast where builders reason through what's changing. Host Conor Bronsdon sits down with the engineers and founders shipping AI in production to get past the hype into what's working and what isn't. Episodes cover model infrastructure, inference, agent frameworks, evaluation, and developer tools. Guests have come from NVIDIA, Google DeepMind, AMD, Databricks, Vercel, and more. Every episode carries a full transcript and show notes at chainofthought.show. New episodes weekly. Conor Bronsdon is an independent consultant and angel investor in AI infrastructure and developer tools. He led technical ecosystem at Modular, acquired by Qualcomm in 2026; led developer awareness at Galileo, acquired by Cisco; and ran developer marketing at LinearB, where he was GM of the Dev Interrupted podcast and community. Views expressed by the host and guests are their own.
On the charts
Every published chart this podcast appears in, in the snapshot behind this page. Each one links to the chart it came off.
From the feed
The latest episodes published to this podcast’s own RSS feed. Titles and descriptions are the publisher’s.
Genspark went from launch to $250 million in ARR in about a year. Along the way it shipped a card-thin meeting recorder, open sourced an office suite that Wen Sang says one engineer prototyped in a week, and started running product triage with agents instead of product managers. Wen's bet is that agents, not people, become the next users of software. Wen Sang is co-founder and COO of Genspark. In this episode he walks through the company's three-layer architecture (models, tools and premium data as the execution layer, a memory layer he calls the second brain, and a collaboration layer called Gen Team), why a meeting note should be the start of work rather than the end of it, the engineering behind the SecondBrain Note, and where he thinks knowledge work goes once agents absorb the busy work. Disclosure: Genspark provided the SecondBrain Note recorder discussed in this episode at no cost. Genspark is not a sponsor of this episode. We cover: Why Genspark builds the self-driving car around the frontier labs' engines, and what that means for people who cannot code How Genspark's mixture-of-agents architecture routes work across 70+ models, 150+ in-house tools and paid data sets Evals that grade whether the sales proposal answered the RFP, not whether the model can solve a differential equation What a meeting turns into a week later when an agent needs it: proposals, pricing models, research, follow-ups The SecondBrain Note's microphone array, battery decisions, and consent in a two-party state Why GenOffice went open source, and the one-week prototype story behind it How Genspark runs product feedback triage with agents and no dedicated PMs Chapters: (0:00) Cold open (0:32) Geniuses with goldfish memories (3:47) Engines and vehicles: Genspark builds the self-driving car (6:25) Mixture of agents: models, in-house tools and premium data (10:03) Grade the work output, not the intelligence (11:22) A meeting note is where the work starts (13:25) The second brain: Genspark's memory layer (14:45) A thousand recorders, one question for the revenue team (16:33) Execution, memory and collaboration layers (19:19) Ten days in Bora Bora without a laptop (20:07) Gen Team, Slack, and meeting customers where they are (21:57) Agents become the users of software (24:17) Keeping memories current when the deal changes (26:41) Engineering the SecondBrain Note (30:18) The note as an API for the room (31:19) What deserves hardware and what stays software (33:33) Learning hardware supply chains at a two-year-old company (35:05) Why GenOffice went open source (38:01) What knowledge workers do once the busy work is gone (40:07) Building on Genspark with the CLI (42:01) Consent, two-party states and the surveillance line (43:45) Genspark Claw (47:09) Cheaper hardware, deeper integration, and model welfare (51:13) Eighty people and a lot of agents (53:03) What 2027 looks like (59:26) Where to find Wen, and product triage without PMs Connect with Wen Sang: LinkedIn: https://www.linkedin.com/in/wen-sang/ Twitter/X: https://x.com/sang_wen Genspark: https://www.genspark.ai GenOffice on GitHub: https://github.com/genspark-ai/genoffice Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Inngest, presenting sponsor of season four of Chain of Thought. Agents in production run long - they call models and wait on APIs and people. But the longer agents run, the more they break. Inngest handles that with durable execution. You build your agent as steps in TypeScript, Python, or Go. When a step fails, Inngest retries it with exponential backoff, and completed steps are saved and skipped. Try it out: https://inngest.link/cot-pod Thanks to G2i for sponsoring this episode - for over a decade, they vetted and placed engineers at other companies, from startups to FAANG. Two years ago, they turned that same judgment inward, building their own bench to review RL environments, evals, and training data,that models are trained on. Get access: https://fandf.co/3SFxVm6
Watch the full conversation on YouTube AI agents can keep tuning GPU workloads after you step away from the keyboard. Anush Elangovan, Corporate VP of AI Software at AMD, returns to Chain of Thought to explain how that works with Hyperloom and ROCm 10. Anush and Conor Bronsdon trace the process from installing ROCm through Claude Code or Codex to profiling workloads, finding slow kernels, and testing optimizations while preserving numerical accuracy. Anush shares a Hyperloom run spanning 14,000 models and explains why clear goals and feedback matter when agents are doing the tuning. They also explore what comes next for engineers: keeping skills and frameworks reliable, managing the security and accountability of autonomous agents, and applying AI to the last mile of useful software. We cover: How agents help install ROCm and serve models through natural language How Hyperloom profiles workloads and uses LLMs to explore optimizations The role of GEAK in tuning kernels while preserving numerical accuracy Anush’s account of optimizing 14,000 models in one pass How software improvements get more performance from existing GPUs Keeping agent skills current and testing across AI frameworks Security, accountability, and the next bottlenecks in agent-driven development Chapters: (0:25) A decade of ROCm, now agent native (3:03) What agentic ROCm looks like in practice (6:17) Installing ROCm then versus now (9:14) An order of magnitude more CI across every framework (10:44) Anush’s workflow: agents and deployment (12:21) Speed is the moat (15:00) Success is a stranger who cannot spell ROCm serving an LLM (17:09) Keeping agent skills from going stale (21:38) Co-designing kernels with the frontier labs (24:06) Hyperloom, GEAK, and 14,000 models in one pass (26:45) Managing autonomous agents: control and liability (32:21) Security at the speed of agent swarms (36:03) ROCm performance gains on the same hardware (37:36) Where enterprises hit walls in production (40:24) Why coding was the right reward function for AI (44:42) Which industries get the next software scale unlock (47:02) The last mile of AI (50:31) Closing thoughts Connect with Anush Elangovan: LinkedIn: https://www.linkedin.com/in/anushelangovan/ Twitter/X: https://x.com/AnushElangovan ROCm.AI: https://rocm.ai AMD AI blog: https://www.amd.com/en/blogs/by-author/anush-elangovan.html AMD AI Developer Program: https://www.amd.com/en/developer/ai-dev-program.html Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot
Jaime DeLanghe has spent nine years at Slack turning search, machine learning, and now agents into product. Her team just shipped Slack Code: tag a coding agent like Claude Code, Devin, Codex, or the GitHub agent in a conversation, and it spins up a code channel where everyone in that conversation gets a live development environment, diffs post as artifacts, and the channel winds down when the task is done. Slack's bet is that AI at work is multiplayer. Agents belong in the channels where teams already work, not in a private chat with one person. Jaime explains why Anthropic pushes so much of its code through Slack, how the channel permission model became the agent context model, and what has to change in engineering culture when the branch is public and the whole team is steering the same agent. The bigger question is whether Slack becomes the context harness where enterprise agents actually run. In this conversation: What happens mechanically when an agent creates a code channel, from authentication to diffs as artifacts Why engineering at Slack now looks like delegating discrete tasks to agents instead of copy-pasting from a chat How Slack's channel permission model doubles as the context and access model for agents Why Anthropic ships code through Slack: the conversation is where the issue emerges How culture decides whether a multiplayer coding session converges or splits Why solo-terminal coding with an "army of Claudes" reinforces bias, and what social spaces fix Slack as an accidental knowledge management system that ranks recency and engagement over correctness (0:00) Slack as an IDE and a GitHub for your team (0:29) Who is Jaime DeLanghe (1:21) The reaction to the Slack Code launch (5:30) Why coding agents belong in a context-rich environment (6:08) Engineers now manage agents, not copy-paste code (7:24) The permission model: agents get the channel's context (11:44) What happens when a code channel is created (15:00) Why Anthropic pushes so much code through Slack (19:14) Steering one agent with many people: culture decides (24:54) Slackbot, skills, and MCPs: agents go where the work is (30:53) The solo terminal vs. agents in social spaces (33:53) Org charts and ownership when agents join the team (39:33) Learning loops and shared agent memory (42:39) Citations, recency, and accidental knowledge management (46:50) Context bloat and multi-pass search for agents (50:01) How Jaime uses Slackbot as CPO (52:38) Slack Code is V1 of multiplayer AI Connect with Jaime DeLanghe: LinkedIn: https://www.linkedin.com/in/jaime-delanghe-aba59b1a/ Slack Code announcement: https://slack.com/blog/news/slack-code-channels-for-agents Introducing Slack Code (Salesforce): https://www.salesforce.com/introducing-slack-code/ Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Walrus Memory, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot . Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000.
Tormod Ree puts 11 or more cameras and dozens of microphones into a single meeting room, then runs computer vision on all of it to figure out who is present, who is talking, and who is looking at whom. He is the chief product and engineering officer at Neat, the Zoom-backed video hardware company. Before Neat, Tormod co-founded AVA, a computer vision security company Motorola acquired, and spent close to eight years at Cisco running the Spark Board. He explains how Neat turns a room into a system that directs the meeting instead of just framing whoever talks, why almost all of the AI has to run at the edge, and how the company builds computer vision models without ever collecting a customer's audio or video. In this conversation: Why the real advantage is the harness that combines signals from different detectors, not the models themselves How Neat draws a hard line: no customer audio or video ever trains its models The "captain" device that orchestrates a room full of cameras and mics, directing from metadata rather than re-running inference on every stream Why the meeting "director" is still deterministic today, and what changes when it becomes a trained model Building AI under a five-year device lifetime and phone-class compute Agentic fleet management, an MCP server, and what Tormod calls "agentic healing" How Neat gets its own non-technical teams building with AI (0:00) Reading the room: 11 cameras, dozens of mics (0:27) Who is Tormod Ree (2:07) Turning a meeting room into a system that directs itself (3:39) What it takes to actually read a room (5:13) Why the media path has to run at the edge (6:27) Open models, in-house models, and the harness that matters (7:51) The data problem when you can't touch customer meetings (11:09) The captain device: distributing compute across the room (15:38) Two users: the people in the room and the IT admin (17:00) Agentic management and "agentic healing" for device fleets (20:12) Why a meeting device has to stay useful for five years (23:44) How agentic workflows evolve on an open platform (25:53) When the meeting director becomes a trained model (28:54) Building AI under five-year, phone-class hardware limits (37:26) Where silicon caps what you can run locally (39:18) AI pendants and other form factors (41:48) How Neat adopts AI across its own teams (49:20) Non-technical teams building their own MCPs (50:27) Where Neat is headed Connect with Tormod Ree: LinkedIn: https://www.linkedin.com/in/toree/ Neat: https://neat.no Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Walrus Memory, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot .
Behind Thomson, the new legal AI model from Thomson Reuters, is a $40 million investment in people, compute, and evaluation methods. The final training run cost just $450,000. CTO Joel Hron, whose teams build Westlaw, Practical Law, and CoCounsel for millions of professionals in more than 100 countries, joined us for the launch to break down why the 175-year-old company chose to own its model layer instead of solely renting frontier intelligence. We cover: Why Thomson Reuters trained its own model instead of relying only on Claude, GPT, or Gemini The compute, data, and expertise flywheel behind the Thomson model How rebuilding CoCounsel around agent-native tools took one-shot accuracy from 25% to over 70% The rent-versus-buy case for owning model weights and compounding expert feedback over time The dangers of AI inaccuracies in legal work Citation ledgers, deep research, and verifying legal work with no ground-truth oracle Joel's advice to CTOs weighing open models and training on their own data Chapters: (0:00) Cold open: a $40M model and 25% to 70% (0:27) Why Thomson Reuters built the Thomson model (2:46) From information services to an AI company (5:14) The flywheel: compute, data, and expertise (8:21) The oldest company to ship a model? (9:47) Training for users without catastrophic forgetting (14:37) Continuous pre-training on Westlaw and Checkpoint (15:29) Fine-tuning, DPO, and agentic reinforcement learning (17:37) Rebuilding CoCounsel: 25% to 70% overnight (21:48) Capturing expert judgment: own versus rent the model (28:59) Managing lawyer time and protecting customer IP (31:45) Eval results and avoiding catastrophic forgetting (34:34) Tabular analysis and legal deep research (36:26) Benchmarks, Harvey, and frontier comparisons (39:26) Verifying legal work with no ground-truth oracle (41:54) Citation ledgers and the hallucinations that matter (45:41) Rebuilding the platform and the Trust in AI Alliance (48:59) Advice to CTOs on open models and owning intelligence (51:31) The compounding flywheel and what comes next Connect with Joel Hron: LinkedIn: https://www.linkedin.com/in/joel-hron-90a3421a/ Thomson Reuters AI: https://www.thomsonreuters.com/en/artificial-intelligence The Thomson model and next-gen CoCounsel Legal: https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-launches-next-generation-of-cocounsel-legal-the-ai-ecosystem-built-for-legal-professionals Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Our sponsors: Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Walrus Memory, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot .
Attackers used to take months, sometimes 270 days, to weaponize a disclosed vulnerability. Now it happens in weeks, minutes if the incentive is there, and independent reports from Mandiant and CrowdStrike show the average time to exploit has gone negative. Dan Lorenc's conclusion: finding flaws is no longer the hard part. Fixing them first is. Dan is the co-founder and CEO of Chainguard. Before that he spent years at Google building the backbone of software supply chain security and created Sigstore. In June his team launched Athena, a coalition of more than two dozen companies including JPMorgan, Cloudflare, Cisco, and Kyndryl, built for the era where AI finds vulnerabilities faster than maintainers can patch them. Last month alone it processed more than 40,000 AI-discovered findings. In this conversation: Why the average time to exploit went negative, and what collapsed the fat tail of never-exploited bugs How models chain "low severity" flaws into working exploits, like Project Zero's zero-click iPhone takeover Inside Athena: what happens between a member submitting a finding and a fix landing upstream Why 40,000 findings is not 40,000 CVEs: validation, dedup, and vulnerability archaeology Fuzzing outpaced patching for a decade, and AI is the first tool that speeds up the fixing side The Log4j thought exercise for solo maintainers and enterprise CISOs alike Fork economics, "state your intentions," and defense in depth for agent infrastructure Chapters: (0:00) Cold open: how time to exploit goes negative (0:31) The 20-year assumption that just died (3:21) What a negative time to exploit actually means (6:04) Two new realities: attacks democratized, more bugs than anyone knew (8:55) Chaining tiny flaws: the Project Zero iPhone story (11:26) Why fixing AI-found vulnerabilities takes a coalition (13:27) 40,000 findings in one month: submission to upstream fix (18:06) The agentic pipeline: as few human eyes as possible (19:03) Fuzzing outpaced patching for a decade (20:50) The Log4j thought exercise for maintainers and CISOs (23:49) When no maintainer answers: the new economics of forking (27:40) Deleting dangerous code to slow the treadmill (29:35) How security kills entire vulnerability classes (31:08) Agent infrastructure: defense in depth or nothing (34:35) Regulation: maintainer liability, frontier labs, DC's busy year (37:01) Open model economics (39:55) Gas Town, multiclaude, and going back to normie (42:47) Why Dan turns the AI memory system off (45:53) Closing: still the most fun time to build software Connect with Dan Lorenc: LinkedIn: https://www.linkedin.com/in/danlorenc/ X: https://x.com/lorenc_dan Chainguard: https://www.chainguard.dev Athena: https://www.chainguard.dev/athena Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000. Thanks to Walrus, presenting sponsor of season four of Chain of Thought. Walrus Memory gives AI agents portable, verifiable memory that carries context across apps, sessions, and other agents. Get started at https://walrus.xyz/cot .
Agents are like teenagers: profoundly intelligent, extremely resourceful, no fear of consequence, and missing the judgment to know right from wrong at all times. That's how Cisco President and Chief Product Officer Jeetu Patel thinks about securing AI agents, and it's why he says static allow/block rules are already obsolete. Agents are smart enough to route around them. Jeetu returns to Chain of Thought to map cyber's third phase: an agent trust platform where security and observability fuse into one discipline. He explains how Cisco Cloud Control spins up a digital twin to test every agent-recommended fix before it touches production, why Cisco moved from unlimited tokens to rationing them like headcount, and why the gap between people who are fluent with AI and people who aren't is now a 10x differential, not 10%. He also makes the contrarian case that AI will create more jobs than it destroys - and of course, we talk infrastructure for this new era. We cover: Why agents need dynamic runtime guardrails instead of static allow/block rules The agent trust platform: how security and observability are fusing into one discipline Agentic ops in Cisco Cloud Control: ambient agents, digital twins, and human-in-the-loop How Cisco rations AI tokens the way it rations headcount The intelligence, cost, and control trade-offs behind open vs closed models Action control vs access control: what rights you actually grant an agent Why every automation step creates a new human bottleneck, and more jobs Chapters: (0:00) Agents are like teenagers: cold open (0:25) Welcome back Jeetu Patel (1:18) Open weight vs closed models (2:26) Intelligence, cost, and control: the model trade-off triangle (10:13) Shrinking model half-life and the economics of frontier training (11:38) Why token costs still outrun token value (14:15) Build your own evals and route intelligently (17:11) Rationing tokens like headcount at Cisco (19:54) Agentic ops: ambient agents and digital twins in Cisco Cloud Control (23:01) When to take the human out of the loop (25:13) Cyber's third phase: the agent trust platform (28:17) Parenting AI agents: dynamic boundary conditions, not static rules (31:10) LiveProtect and baking security into the network fabric (34:40) Action control vs access control for agents (36:42) What the industry is getting wrong (37:48) The case for AI creating more jobs and the 10x fluency gap (42:36) Career paths, entry-level hiring, and upskilling at scale (45:40) Closing thoughts Connect with Jeetu Patel: LinkedIn: https://www.linkedin.com/in/jeetupatel/ X: https://x.com/jpatel41 Cisco: https://www.cisco.com Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Svix, presenting sponsor of season four of Chain of Thought. Svix delivers billions of reliable webhooks for startups and the Fortune 500. Get started at https://link.svix.com/cot . Qualified startups get $12,000 in credits, and YC companies get $50,000.
Jitender Aswani was customer zero for Presto at Meta, where a billion daily active users generated queries that took hours to return. He watched that drop to minutes, scaled the same technology at Netflix across 300 million subscribers, and now runs engineering and security at Starburst, the $3.35 billion platform built on Trino. His argument: every enterprise AI project that stalls is fighting the same hidden battle. The agents can query the model fine. They just can't reach the data. The average enterprise runs 52 to 200 data sources, and a decade of moving all of it into one lake produced ETL debt, governance problems, and pipelines that break whenever a SaaS vendor adds a column. Federation is the only model that scales with entropy. We cover: Why Presto changed what Meta could experiment on, and how that compounded product velocity What broke when Jitender took the same technology to enterprises running 52 to 200 data sources Why centralization stopped working once data grew faster than the ability to move it What happened to Starburst's query volume the day they shipped an MCP server The FinOps agent that fired queries for 30 minutes against data it never had How AIDA turns ad hoc analysis into workflows using skills and MCP servers Why a context graph is different from a knowledge graph, and why ontology decides agent accuracy (0:00) Enterprises run on 52 to 200 data sources (0:25) Intro (2:18) Customer zero for Presto at Meta (9:50) Scaling to trillions of events at Netflix (15:11) Taking Trino from Silicon Valley to 10,000 enterprises (20:24) The 2011 research that predicted conversational analytics (28:54) Why centralization can't scale with entropy (32:26) The agent query explosion and what MCP did to volume (41:44) Inside AIDA, Starburst's conversational analytics product (46:56) Context graphs versus knowledge graphs (51:17) Where to follow Jitender's work Connect with Jitender Aswani: LinkedIn: https://www.linkedin.com/in/jitenderaswani/ Starburst: https://www.starburst.io/ Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Charles Guillemet is CTO of Ledger and the founder of the Donjon, Ledger's internal offensive security lab whose job is to break the company's own products before attackers do. He spent a decade in cryptography and hardware security before Ledger, including designing secure integrated circuits. His argument is blunt: you cannot secure an AI agent with software alone. As agents start moving real money, API keys and trust scopes leave no physical verification layer, and Charles makes the case that hardware has to sit in the loop. This one turned into a wide-ranging thought piece (and some debate) on what the agentic economy actually looks like, and how to stay safe inside it. We cover: Why Charles thinks "securing an AI agent" with software permissions and API keys is a false promise The economic asymmetry between attackers and defenders, and how AI is collapsing it How a policy engine plus a hardware-enforced signature can delegate rights to an agent safely Why Charles thinks the agentic economy settles on blockchain rails over Visa and Mastercard Secure elements, HSMs, and zero-knowledge proofs as execution-integrity guarantees How Ledger uses hardware authorization internally for passkeys, signed releases, and multisig A practical way to classify assets by threat model and match security to value (0:00) Why securing an AI agent in software alone is impossible (0:30) Delegating execution power inside your security perimeter (2:28) The attack-defense asymmetry AI is erasing (6:00) The alignment problem and delegating rights to agents (9:24) Policy engines, intents, and hardware-enforced signatures (13:19) From developer experience to agent experience (15:12) Secure elements, HSMs, and execution integrity (20:00) Zero-knowledge proofs, proving without revealing (27:24) Convincing the skeptics on agent-driven payments (34:49) Why Ledger bet on dedicated hardware (36:15) Hardware as a determinism layer for agents (38:52) How Ledger uses hardware authorization internally (43:42) Classifying assets by threat model (46:55) When attack and defense become symmetric (48:44) Deepfakes, voice cloning, and the scam wave (50:04) Closing thoughts on staying safe in the agentic economy Connect with Charles Guillemet: LinkedIn: https://www.linkedin.com/in/charles-guillemet/ Twitter/X: https://x.com/P3b7_ Ledger: https://www.ledger.com Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Jiaona Zhang(JZ) is the Chief Product Officer at Laurel, where the team runs its own product on itself to see exactly where AI helps and where it doesn't. Before Laurel, JZ built products at Airbnb, Dropbox, Webflow, and Linktree, and she has taught product management at Stanford for nearly a decade. Companies are spending billions on AI tooling, but most still can't say where it returns time or revenue. Jiaona breaks down how to get that visibility, why blanket AI mandates backfire, and what it takes to re-architect a team so anyone can ship. Her argument is simple: stop token maxing and start measuring time back. We cover: Why most organizations can't see where AI is actually working, and how Laurel uses time data to fix it The token max trap that "use AI everywhere" mandates create, and how to drive efficient use instead Why former managers make the best operators of agent fleets How Laurel lets PMs, designers, and customer success ship features end to end The bottom-up plus top-down playbook for re-architecting a team around AI Why technology moats are falling away while brand and data moats endure Laurel's bet on returning time to people instead of replacing them (0:00) The token max trap (1:47) Why companies can't see where AI is working (5:03) What Laurel does: turning time into data (8:53) Agents as an extension of the workforce (13:43) Why former managers make the best AI users (18:23) Lean teams and shipping end to end (22:29) Enabling non-engineers to ship features (28:30) Re-architecting teams: bottom-up and top-down (32:09) Keeping your professional identity as AI shifts work (38:53) The context layer is the new race (42:06) Fundamentals plus tinkering: how to learn (48:45) Brand and data moats when tech moats fall away (54:31) Laurel's movement: returning time to people Connect with Jiaona Zhang(JZ): LinkedIn: https://www.linkedin.com/in/jiaona/ Laurel: https://www.laurel.ai/ JZ's Linktree: https://linktr.ee/jz Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Most of the web will never get APIs for AI agents. School district sites, small business pages, government offices, and the long tail of e-commerce were built for humans, and they will keep working that way for years. So how do agents actually get things done across the web? Dhruv Batra is co-founder and chief scientist of Yutori, the company building specialized browser and computer-use agents. He previously led embodied AI at Meta's FAIR lab, training robots in simulation and shipping the image question-answering model on Ray-Ban Meta glasses. His bet: the web is a shared roadway, much like roads split between human drivers and self-driving cars, and agents will be built to use it the way people already do. Pixels in, clicks out. That is the API. In this conversation: Why the long tail of the web won't re-architect itself for agents How Yutori's Navigator perceives pixels and writes JavaScript on the fly to shorten task trajectories Why Navigator runs 2-3x faster and 4-5x cheaper than Opus 4.7 and GPT-5.5 on browser tasks Learning from live websites, and using URL query parameters as privileged verifiers instead of cloning sites What the shift from American to Chinese open-weight models means for startups How smart glasses and robots share the same perception-action loop Why demand for inference compute is pushing models smaller and onto devices Chapters: (00:00) Pixels in, clicks out (01:37) Why most of the web will never get APIs (08:47) Aggregation, specialization, and human friction (11:39) Digital niches and specialized models (16:41) The web's heavy tail and where browser agents win (20:40) Inside Yutori's Navigator and Scouts (24:08) N1.5: writing JavaScript to cut trajectory length (27:45) Training on live websites (33:29) Open source: FAIR's legacy and the Chinese frontier (37:22) Agent frameworks: OpenClaw, Hermes, heartbeats (40:57) How non-technical users adopt agents (44:25) Smart glasses, robotics, and embodied AI (50:57) Compute demand and smaller on-device models (53:12) Why the company is called Yutori Connect with Dhruv Batra: LinkedIn: https://www.linkedin.com/in/dhruv-batra-dbatra/ X/Twitter: https://x.com/DhruvBatra_ Yutori: https://yutori.com Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Kristin "Kris" Lovejoy has spent her career inside the systems the global economy runs on: banks, hospitals, energy grids, governments. Today she is Global Head of Strategy at Kyndryl, the world's largest IT infrastructure services provider, working with mission-critical enterprises across more than 60 countries. Before that she ran security businesses at EY and IBM, founded the AI security company BluVector (acquired by Comcast), and now sits on the board of Dominion Energy. Her prediction: the first fully autonomous AI attack, where an AI takes down an enterprise network with no human driving it, lands within 18 months. Conor and Kris dig into why 62% of enterprise AI initiatives are still stuck in pilots even as spend climbs 33% year over year, why attackers chaining low-risk vulnerabilities changes the patching math, and why she has a fraught relationship with policy as code. We cover: The electricity analogy: we can build the models, but the transmission lines for industrial AI don't exist yet Productivity AI vs mission-critical AI, and why banks and healthcare systems aren't running agentic AI at production scale Why deterministic policy as code clashes with autonomous systems, and "human on top" vs human in the loop The 18-month prediction: chaining low-risk vulnerabilities, outcome-oriented agents that take systems down by accident, and insiders armed with AI attack tools The data center build-out from a Dominion Energy board member: PJM load forecasts that miss by double digits every year, water use, density, and rack optimization Privacy as a double-edged sword: data combinations that suddenly become PII and the shift to continuous compliance What's next: open source everywhere, sovereignty as control, autonomous robotics, and quantum Chapters: (00:00) Meet Kris Lovejoy: Kyndryl, EY, IBM, and Dominion Energy (02:09) Why 62% of AI initiatives are stuck in pilots (03:07) The electricity analogy: models without transmission lines (04:23) Productivity AI vs mission-critical AI (06:53) Vintage systems, hybrid data, and the risk gap (11:03) Policy as code and "human on top" (16:25) Data centers, energy, and the grid build-out (24:44) Data center design: density, cooling, rack optimization (26:54) Privacy, continuous compliance, and sovereignty as control (32:06) The first fully autonomous AI attack: 18 months away (38:06) Predictions: open source, robotics, and quantum (42:32) Control planes for agentic AI: closing thoughts Connect with Kris Lovejoy: LinkedIn: https://www.linkedin.com/in/klovejoy/ Kyndryl: https://www.kyndryl.com Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Jerry Liu built one of the most installed pieces of AI plumbing of the last three years. LlamaIndex became the indexing and retrieval layer a whole generation of RAG apps were stitched together with. Then he started arguing that the framework era he helped create is over. Jerry is co-founder and CEO of LlamaIndex. In this conversation he walks through the company's pivot from open-source framework to managed document infrastructure with LlamaCloud and LlamaParse, and why he is betting that context quality is the one moat that compounds as agent loops get good enough to absorb the scaffolding. If you are a founder worried a frontier lab or a coding agent is about to eat your product, this is the playbook for reinventing your ICP without losing the thread. In this conversation: Why Jerry says the AI framework era is over, and what actually survives How agent harnesses like Claude Code collapsed the old framework patterns into the model Why context quality is the durable moat, not the agent loop How LlamaParse beats legacy OCR and frontier models on document accuracy and cost Why 95%+ accuracy is the real bar for legal, insurance, and financial document work How LlamaIndex disrupted its own product and reinvented its ICP to stay alive Jerry's take on agent memory, model personalities, and why LLMs are still bad writers (0:00) Is the AI framework era over? (1:56) What died and what survived (6:31) Why context quality is the moat (8:12) Defining the context layer (13:18) Coding and vision as the abstraction layer (18:13) The bet that context compounds (23:59) Which verticals are adopting (25:14) Why 95%+ accuracy is the real bar (29:49) The file system as an agent primitive (34:33) Surviving your own pivot (37:15) Reinventing strategy and hiring (42:00) Agent memory as persistent context (44:41) Model personalities and cultural memory (47:51) Writing with AI (50:19) Closing thoughts Connect with Jerry Liu: LinkedIn: https://www.linkedin.com/in/jerry-liu-64390071/ Twitter/X: https://x.com/jerryjliu0 LlamaIndex: https://www.llamaindex.ai LlamaIndex careers: https://www.llamaindex.ai/careers Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Tyler Akidau spent 12 years on streaming systems at Google and five years at Snowflake before joining Redpanda as CTO. He wrote the O'Reilly Streaming Systems book most of the field has on its shelf. His new piece on O'Reilly Radar (Post-Human: We All Built Agents, Nobody Built HR) argues that enterprises are stuck in the prototype-to-production gap because they're applying human-era identity, auth, and observability tools to a workforce that's unpredictable in structurally novel ways, runs at machine speed, and follows bad instructions to a fault. Inline guardrails like CLAUDE.md work until they don't. Governance has to be enforced through channels the agent can't see, modify, or override. We cover: Why AI agents are a new kind of co-worker (unpredictable, machine-speed, directable to a fault) and what that means for enterprise infrastructure The four pillars of agent governance: identity, authorization, observability and explainability, accountability and control Why task-scoped, short-lived identity is the foundation everything else builds on Authorization that's deny-capable and intersection-aware (Tyler's "guest badge" model) Why OpenTelemetry is the right starting point for recording every prompt, tool call, and response How Redpanda's Agentic Data Plane combines streaming topics, Oxla SQL, and Postgres under the hood Tyler's academic paper with a psychologist on the neurobiological systems humans have that AI agents are missing Chapters: (00:00) Why nobody built HR for AI agents (02:12) Three ways agents differ from human employees (07:53) The four pillars of out-of-band governance (10:29) Identity: task-scoped, short-lived, chained to humans (14:40) Authorization: deny-capable and intersection-aware (18:57) Observability: record everything via OpenTelemetry (24:24) Redpanda's agents and the $1,000 trade limit example (30:10) Accountability and the kill switch (34:02) The Agentic Data Plane: streaming, Oxla SQL, Postgres (41:20) Should we stop chasing model alignment? (44:04) Building human-like value systems into agents (47:25) Tyler's 12-24 month outlook for agent governance Connect with Tyler: LinkedIn: https://www.linkedin.com/in/takidau/ Redpanda: https://www.redpanda.com/ Post-Human article: https://www.oreilly.com/radar/posthuman-we-all-built-agents-nobody-built-hr/ Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show
Loïc Houssier leads engineering at Superhuman, the email client Grammarly acquired for ~$825 million in July 2025. Before Superhuman he was CTO of OpenTrust (acquired by DocuSign), ran engineering at ProductBoard, and started his career in applied cryptography for France's defense industry, including work on nuclear submarine systems. Loïc joined Superhuman in early 2024 and within 30 days was leading a six-week sprint to ship AI Inbox. Superhuman's brand is built on speed: every interaction under 100 milliseconds. LLMs do not run in 100 milliseconds. So Loïc walks Conor through how his team retrofitted AI into a product that was already winning without it: pre-caching context for the mobile voice feature, starting every feature on the smartest available model and only then fine-tuning down to cheap dedicated infrastructure, treating "look foolish" as a P0 bug class, and refusing to auto-send any email even when their agents could. This is a practitioner's tour of what it actually takes to put AI on top of a product that has to stay fast, stay quiet, and never embarrass the user. We cover: The model-routing strategy: Opus and frontier models to prove a feature, then fine-tuned BERT classifiers on dedicated inference Pre-caching voice and tone context separately from dictation to keep the mobile voice feature feeling fast Why eval engineering at Superhuman is owned by PMs, and how a single "how much time did I spend in Waymo last month" query exposes the eigenvectors a feature has to cover Why "look foolish" is a P0 bug class, and where the boundary between agent agency and agent laziness actually sits How Superhuman's pod structure (PM, tech lead, designer) and a central AI platform team support aligned autonomy Hiring for AI fluency: how interview questions are changing and what self-augmenting engineers look like Pattern detection as the leadership skill that transfers from nuclear submarines to AI email Chapters: (00:00) Cold open: pattern detection beats new tools (00:18) Loïc's path: cryptography, OpenTrust, ProductBoard, Superhuman (02:13) Retrofitting AI into a 100ms product (04:08) Voice on mobile: pre-caching LLM context to keep the feel fast (07:46) Frontier first, then fine-tune: model strategy across features (11:04) The "double-dipping" trick that worked on GPT-4 and stopped working (12:25) Cognitive load and staying current as a leader (16:59) Balancing YC founder urgency with peer CTO grounding (19:28) Pods, AI Guild, and aligned autonomy (23:15) Managing models vs. managing people: delegation in reverse (28:27) The Waymo example: eigenvectors of evaluation (32:15) Day 30 onboarding: leading the AI Inbox sprint (35:04) Why email is the killer agent use case (38:51) Auto-draft, never auto-send (39:57) Agent agency vs. agent laziness (43:07) Hiring for AI fluency (45:55) Pattern detection is the leadership skill (47:21) Nuclear submarines as engineering reference points (48:37) Closing thoughts (49:38) Superhuman is hiring Connect with Loïc: LinkedIn: https://www.linkedin.com/in/houssier/ Superhuman careers: https://superhuman.com/careers Superhuman: https://superhuman.com Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Galileo — download their free 165-page guide to mastering multi-agent systems at galileo.ai/mastering-multi-agent-systems
Job applications are up 239% since ChatGPT launched, tech layoffs show no signs of slowing down, and the market for technical talent is a topsy turvy mess. Greenhouse has a unique vantage point to understand all of this: they process 22 million job applications a month across 7,500+ companies including HubSpot, Anthropic, Coinbase, and the NFL. CEO Daniel Chait has had a front-row seat to the strangest hiring market in decades, and he's here to advise us all on how to navigate it. Daniel coined the term "AI doom loop" for what's happening: applications up 239% since ChatGPT launched, resume hacks like white-fonting and prompt injection up 500%, and 75% fewer applications reaching the hire stage. 91% of recruiters have spotted candidate deception. 38% of job seekers walk away from processes that include an AI interview. It's the worst job market for candidates and the hardest hiring market for recruiters. Daniel explains how technical talent can break the loop. We cover: Why software engineers, according to Greenhouse data, are the worst auto-appliers and what to do instead The North Korean infiltration problem: deepfakes, laptop farms, and why companies are flying candidates in for in-person interviews again How AI screener interviews open up the funnel when companies are transparent about using them, and break it when they aren't Greenhouse Dream Jobs: how a single high-signal application a month converts at 5x the rate Why take-home assignments don't survive contact with AI and what Greenhouse uses instead What a coding interview looks like when leetcode is dead and engineers run 10+ Claude Code sessions in parallel The case for killing the resume entirely and rebuilding hiring around AI conversations Chapters: (00:00) Cold open: 239% more applications, 75% fewer hires (02:14) Galileo (03:05) The AI doom loop, defined (04:01) How we got here: remote work, ZIRP, and ChatGPT (07:51) Are software engineering jobs really in trouble? (12:46) The trust crisis: 91% of recruiters spot deception (15:52) North Korean spies, deepfakes, and laptop farms (19:34) Can AI fix the problem it created? (20:52) AI screener interviews and the uncanny valley (26:33) Greenhouse Dream Jobs: one signal, 5x conversion (28:31) Why auto-apply doesn't work (and what does) (30:18) Communities, building in public, and the early-mover advantage (37:08) Gen Z lost trust, and the bias problem (39:04) Kill the resume: rethinking hiring from scratch (43:34) How Greenhouse changed its own interview process (48:47) Coding interviews in the agent era: leetcode is dead (51:33) Predictions: more proof, more conversations, less noise (54:34) Where job seekers and hiring teams should start Connect with Daniel: Greenhouse: https://www.greenhouse.com My Greenhouse (for job seekers): https://www.mygreenhouse.com LinkedIn: https://www.linkedin.com/in/dhchait/ Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Galileo — download their free 165-page guide to mastering multi-agent systems at galileo.ai/mastering-multi-agent-systems
Alex Ratner co-founded Snorkel AI out of Chris Ré's Stanford lab and helped establish data-centric AI as a field. Today, Snorkel is a $1.3B company shipping thousands of data sets and environments a week to frontier labs and vertical AI teams like Harvey. In this conversation, he argues our ability to build AI agents has outpaced our ability to measure them. That gap is what's keeping most enterprise agents stuck in demo purgatory. If you can't measure it, you can't improve it. And you can't deploy it. In this conversation: The three axes of the evaluation gap: input complexity, autonomy horizon, and output complexity Big Law Bench: how Snorkel and Harvey benchmarked legal agents on deep-research tasks that take lawyers 10-15 hours What Snorkel's $3M Open Benchmarks Grant is funding, and why "benchmaxxing" critiques don't kill the case for public benchmarks Why 40-50% of Snorkel's data work is still review and labeling, even with the best models in the loop The "expert-agentic" era, where domain expertise (law, finance, coding, even woodworking) is the new bottleneck Why self-supervision is a dead end outside narrow cases like distillation The false dichotomy between data and environments, and why pure-environment vendors miss how AI actually works Chapters (00:00) Intro: Alex Ratner and Snorkel AI (02:50) What the evaluation gap actually is (06:05) Moravec's paradox and the jagged frontier (08:46) Where AI agents fall down in enterprise work (10:40) Big Law Bench: benchmarking Harvey's legal agents (12:00) The three axes: input, autonomy horizon, output (18:31) Snorkel's $3M Open Benchmarks Grant (22:33) From "janitorial" to epicenter: 15 years of data-centric AI (29:26) The expert-agentic data era (34:54) The false dichotomy between data and environments (40:05) DoorDash Tasks and expert data at scale Connect with Alex Ratner: X/Twitter: https://x.com/ajratner Snorkel AI: https://snorkel.ai Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Galileo — download their free 165-page guide to mastering multi-agent systems at galileo.ai/mastering-multi-agent-systems
What happens when a VP of AI Software at a major chip company goes all-in on AI coding agents for his own team's work? Anush Elangovan runs 10–12 Claude Code agents across three machines, burns 6.5 billion tokens a week, and rewrote a 25-year-old project (Slurm → Spur in Rust) in a single night. He does it all on dangerously-skip-permissions. About Anush Anush Elangovan is Corporate VP of AI Software at AMD. He founded Nod.ai, where his team built SHARK and was a primary contributor to Torch-MLIR and IREE. AMD acquired Nod.ai in 2023, and Anush now leads AI software strategy across AMD's full silicon portfolio. Before Nod.ai, he shipped the graphics stack on the first ARM Chromebook and led Chrome OS's migration to Gentoo. We cover: How Anush runs 10–12 parallel agents with a geo-distributed AMD hardware rig Why the test harness is the new code review (and why agents are "sneaky and dumb") Rewriting a 25-year-old project in Rust overnight, without opening the editor Why every new project is in Rust specifically because he refuses to learn it The "HR partner fixing engineering bugs" moment and what it says about upskilling Why normal SDLC is dead and speed is the only durable moat AMD's fully open-source software stack and how community contributions are accelerating ROCm "Software is just tokens" and what that means for AMD's bet against CUDA lock-in Connect with Anush LinkedIn: linkedin.com/in/anushelangovan Twitter/X: @AnushElangovan AMD AI blog: amd.com AMD AI Developer Program: amd.com/developer Connect with Chain of Thought host Conor Bronsdon: Newsletter: newsletter.chainofthought.show Twitter/X: @ConorBronsdon LinkedIn: linkedin.com/in/conorbronsdon YouTube: @ConorBronsdon More episodes: chainofthought.show Chapters 0:00 Cold open 0:21 Welcome + guest intro 3:43 250K lines a week, 10–12 parallel agents 7:34 Agent architecture + geo-distributed test rig 9:57 When does AI-generated code become a liability? 14:12 80% tests first: the test harness philosophy 18:24 Dangerously-skip-permissions + testing as code review 19:52 "Normal SDLC is dead in the agentic world" 20:44 Advice for engineers and leaders who feel behind 24:51 Tokens, throughput, and what happens next 26:29 Block layoffs, uneven AI gains, the 25-year Slurm rewrite 32:55 Galileo sponsor break 34:24 When agents go off the rails: sneaky and dumb 37:52 Orchestrator agents vs. focused multi-threading 40:45 Open source, ROCm, AMD's software bet 44:19 "Software is just tokens" 45:24 AMD Developer Program + community contributions 47:09 Where to start with AMD 48:39 Heterogeneous compute 50:13 Outro Thanks to Galileo. Download their free 165-page guide to mastering multi-agent systems at galileo.ai/mastering-multi-agent-systems Full show notes: newsletter.chainofthought.show Disclaimer from our host: All views, opinions and statements expressed on this account are solely my own and are made in my personal capacity. They do not reflect, and should not be construed as reflecting, the views, positions, or policies of my employer. This account is not affiliated with, authorized by, or endorsed by my employer in any way.
Sudhir Hasbe is President and Chief Product Officer at Neo4j, the graph database company powering 84 of the Fortune 100 (Walmart, Uber, Airbus) at $200M+ ARR and a $2B+ valuation. Before Neo4j, he ran product for all of Google Cloud's data analytics services: BigQuery, Looker, Dataflow, and led the Looker acquisition. His thesis: the hallucinations we blame on AI models are really a data architecture problem. LLMs weren't trained on your enterprise knowledge, so handing them a data lake with 10,000 disconnected tables and asking them to reason is the wrong design. The fix is knowledge graphs: feeding the model a structured map of relationships, entities, and context so it can reason over meaning, not just vector similarity. Sudhir breaks down the five capabilities knowledge graphs unlock for enterprise AI: GraphRAG (moving accuracy from 60% to 97%), semantic mapping across siloed systems, context graphs, agent memory, and multi-hop reasoning. He explains three architecture patterns customers are actually shipping, why giving an LLM hundreds of tools makes it worse, and what Uber, EA Sports, Klarna, and Novo Nordisk are doing differently. This is the case for treating knowledge as infrastructure. We cover: Why enterprise AI needs a different playbook than consumer AI The five data asset types every agentic system needs: system of record, historical, memory, context, and reference How GraphRAG combines vector search and graph traversal to move from 60% accuracy to 95%+ Three architecture patterns: semantic layer only, semantic map plus domain data, full consolidation (the Klarna/Kiki model) What context graphs capture that Salesforce doesn't: the Slack and email negotiation behind every deal Why giving an LLM hundreds of tools drops accuracy, and how Uber uses knowledge graphs as a business validation layer What Neo4j's Aura Agent, MCP server, and A2A support mean for developers starting today Chapters: (0:00) Why building a self-driving car is hard (0:22) Intro (2:03) Hallucinations as a data architecture problem (4:31) From models-as-core to systems-of-knowledge (6:13) Why data lakes fail AI agents (9:15) The five data asset types enterprise agents need (11:46) Where basic RAG breaks down: the Spotify metadata lesson (16:00) GraphRAG: 3x accuracy, easier development, explainability (18:47) Semantic mapping across the enterprise estate (19:23) Three knowledge-graph architecture patterns (22:42) Context graphs: capturing the "why" behind decisions (25:33) Individual vs. organizational agent memory (28:40) Multi-hop reasoning for fraud rings and AML (31:52) Why there are no shortcuts in enterprise AI (36:38) What happens when you give an LLM 100 tools (39:19) The Uber example: knowledge graph as business validation (44:42) First mile of a 26-mile marathon (48:32) Aura Agent, MCP server, and the A2A protocol (50:43) Where developers should start Connect with Sudhir Hasbe: LinkedIn: https://www.linkedin.com/in/shasbe/ Neo4j: https://neo4j.com/ Neo4j Aura: https://neo4j.com/product/auradb/ Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Galileo — download their free 165-page guide to mastering multi-agent systems at: galileo.ai/mastering-multi-agent-systems
Every few weeks at Microsoft, someone would build an AI prototype that blew everyone's minds. Three months later? Dead. "We can never ship that." Dan Klein watched this happen for five years before he decided to do something about it. Dan is co-founder and CTO of Scaled Cognition, a professor of computer science at UC Berkeley, and winner of the ACM Grace Murray Hopper Award. His previous startups include adap.tv (acquired by AOL for $405M) and Semantic Machines (acquired by Microsoft in 2018), where he spent five years integrating conversational AI. His PhD students now run AI teams at Google, Stanford, MIT, and OpenAI. At Scaled Cognition, Dan's team built APT1 (the Agentic Pre-trained Transformer) for under $11 million. It's a model designed for actions, not tokens, with structural guarantees that go beyond prompt-and-pray. Dan makes the case that current LLMs are plausibility engines, not truth engines, and that the gap between demo and production is where most AI projects die. Why prompting is a fundamentally unreliable control surface for production AI How APT1's architecture gives actions and information first-class status instead of treating everything as tokens The specific failure modes that kill enterprise AI prototypes within three months Why stacking multiple models to check each other produces correlated errors, not reliability How Scaled Cognition applied RL to conversational AI when there's no zero-sum winner Why every S-curve in AI gets mistaken for an exponential — and what comes after the current plateau The societal risk of systems that produce output indistinguishable from truth Chapters (0:00) Cold open: RL is about doubling down on what works (0:28) Introducing Dan Klein and Scaled Cognition (2:53) The demo-to-production gap: why AI prototypes die (5:40) Why prompting is not a real control surface (8:06) Modular decomposition vs. end-to-end optimization (10:55) Are LLMs fundamentally mismatched with how we use them? (14:26) What's wrong with benchmarks today (20:27) APT1: building a model for actions, not tokens (24:14) What makes data truly agentic (28:02) Hallucinations as an iceberg — visible vs. undetectable (34:16) Building a prototype model for under $11 million (39:57) Applying RL to conversations without a zero-sum winner (43:31) LLMs as a condensation of the web — and what happens when it runs out (50:07) Reasoning models: where they work and where they don't (53:04) Early deployments in regulated industries (57:14) Why multi-model checking fails (1:00:34) The minimum bar for trustworthy agentic systems (1:04:07) Societal risk: when AI output is indistinguishable from truth (1:13:33) Where Dan is inspired in AI research today Connect with Dan Klein: Scaled Cognition: https://scaledcognition.com LinkedIn: https://www.linkedin.com/in/dan-klein/ UC Berkeley NLP Group: https://nlp.cs.berkeley.edu Connect with Chain of Thought host Conor Bronsdon: Newsletter: https://newsletter.chainofthought.show/ Twitter/X: https://x.com/ConorBronsdon LinkedIn: https://www.linkedin.com/in/conorbronsdon/ YouTube: https://www.youtube.com/@ConorBronsdon More episodes: https://chainofthought.show Thanks to Galileo — download their free 165-page guide to mastering multi-agent systems at galileo.ai/mastering-multi-agent-systems
Ranking source
Apple Podcasts rankings via the Mato Topic Intelligence Platform.
Observed September 19, 2026.
Apple and Apple Podcasts are trademarks of Apple Inc., registered in the U.S. and other countries.
Pairs with
Bring this source into Mato to read its transferable patterns, then turn them into an original show for your own audience.