Episode · Super Data Science: ML & AI Podcast with Jon Krohn
1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak
25 Sep 2026 · 23 min
Episode · Super Data Science: ML & AI Podcast with Jon Krohn
25 Sep 2026 · 23 min
During their #sponsored discussion, Senior Vice President Product Management AI and Metadata at Salesforce, Gaurav Pathak talks to Jon Krohn about why AI agents need well-labeled, high-quality data to deliver reliable answers in the enterprise. Listen to the episode to hear Gaurav Pathak talk about the difference between a “data brawl” and “garbage in, gospel out”, who the “sin eaters” of enterprise AI are and the three skills that matter most for AI engineers today! Additional materials: www.superdatascience.com/1030 Interested in sponsoring a SuperDataScience…
by Jon Krohn · English · Tech & Science
The latest machine learning, A.I., and data career topics from across both academia and industry are brought to you by host Dr. Jon Krohn on the Super Data Science Podcast. As the quantity of data on our planet doubles every couple of years and with this trend set to continue for decades to…
6 Oct 2026 · 1 hr 15 min
In Episode #1033, Prof. Jeff Hancock (Professor of Communication at Stanford) and Dr. Kate Niederhoffer (Chief Scientist at BetterUp) join Jon Krohn to explain the hidden cost of AI-generated work. A year ago they coined "workslop" in a Harvard Business Review article that went viral and landed the term among Merriam-Webster’s words of the year: content that masquerades as real work but quietly shifts the burden onto whoever receives it. Their research finds that 40% of workers have been sent workslop and 53% admit to producing it, at a cost running to millions of dollars a year for a large…
2 Oct 2026 · 28 min
During their #sponsored discussion, Chief Data Officer at Salesforce Michael Andrew talks to Jon Krohn about what changes for a data team when its customers are AI agents as well as people. Listen to the episode to hear Michael Andrew talk about why agents need ten times more trusted data than humans, how Salesforce untrapped its own customer data with Data 360 and what practitioners should be learning to stay effective in the agentic era! Additional materials: www.superdatascience.com/1032 Interested in sponsoring a SuperDataScience Podcast episode? Email…
29 Sep 2026 · 1 hr 15 min
In Episode #1031, Ish Shah and Tyler Cox (Distinguished Engineers in the Office of the CTO for Dell Technologies' client group) join Jon Krohn to work out why agentic AI bills are exploding and what can be done about it. Over one weekend Ish burned roughly two billion tokens on a side project, and that is the ordinary shape of agentic work now: agents spawn sub-agents, the pie of work grows, and cheaper tokens only invite more ambitious projects. Tyler runs a small Dell lab that pushes hundreds of millions of tokens a day through local hardware instead. In this episode, they define what…
22 Sep 2026 · 1 hr 10 min
In Episode #1029, Dr. Katie Malone (Host of Linear Digressions) joins Jon Krohn to explain how AI brought her podcast back from the dead. After nearly 300 episodes, Katie shut down Linear Digressions due to burnout, but better tools helped her relaunch it six years later. Along the way she has taught machine learning at Udacity and the University of Chicago and led the development of agentic AI platforms inside a company of tens of thousands of people. In this episode, she argues that people management and agent management are the same skill in different clothing, works through what AI slop…
18 Sep 2026 · 30 min
In Episode #1028, Anton McGonnell (VP of Product at SambaNova) joins Jon Krohn to explain why the chips running most AI inference today were never designed for the job. Agentic AI has changed the computational profile of inference, with much larger inputs and far heavier caches feeding the token generation that follows, and that shift has exposed where GPU architecture struggles. SambaNova has raised over $2 billion to build an alternative, the reconfigurable dataflow unit, which lays a whole model out spatially across the chip rather than executing it kernel by kernel. In this episode,…
15 Sep 2026 · 1 hr 3 min
In Episode #1027, Dr. Dilani Kahawala (Co-Founder and CEO of Anna) joins Jon Krohn to explain what it takes to build an always-on AI assistant that busy parents will trust with their inboxes. Anna watches the email, school apps, WhatsApp messages and calendars flowing into a family's life and surfaces what matters, over text and voice, with barely any app to speak of. Dilani came to it by way of a Harvard physics PhD, McKinsey, and a decade of product leadership at Etsy, Meta and Atlassian, and says she has had to throw away most of what that decade taught her about how products get built.…
9 Oct 2026 · 44 minNew
In ICYMI Episode #1034, Jon Krohn moves from what sits underneath AI systems to what it takes to live with them every day. Hear from Aishwarya Srinivasan, Luis Serrano, Katie Malone, Ish Shah and Dilani Kahawala, discussing why the bottleneck in shipping software has moved from execution to judgment, how a physicist recast attention in transformers as words bending space, why years of managing people may be the best preparation for managing AI agents, how one weekend side project burned through billions of tokens and which three problems are hardest to solve when a product has no user…
11 Sep 2026 · 20 min
In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the safety story, which for this release is unusually intertwined with capability. He weighs the AGI claim against Anthropic’s Fable 5.1 and lands, as ever, in a measured middle. Additional materials: www.superdatascience.com/1026 Interested in sponsoring a…
8 Sep 2026 · 1 hr 11 min
In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than…
4 Sep 2026 · 34 min
In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the…
Every episode of Super Data Science: ML & AI Podcast with Jon Krohn →