Episode 511 · Data Engineering Podcast
Text to Data Products: Kaarvi’s End-to-End AI for Ingestion, Quality, and Dashboards
8 Jun 2026 · 53 min
Episode 511 · Data Engineering Podcast
8 Jun 2026 · 53 min
Summary In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams. He explores Kaarvi’s multi-agent architecture that runs queries across seven LLMs in parallel for reliability, its synthetic data generator that mirrors source schemas for quick testing, and “Hey Kaarvi” chat for text-to-SQL, text-to-transformations, and text-to-dashboard workflows. He also digs into on-prem versus SaaS deployments, domain-specialized agents for privacy and accuracy, code…
Tap a chapter to play from there.
by Tobias Macey · English · Tech & Science
This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.
E514 · 2 Aug 2026 · 1 hr 2 min
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for traversal workloads, and how that perspective shaped OmniGraph’s design on top of object storage, Lance, Arrow, and DataFusion. Ragnor explained the motivation for combining graph semantics with Git-style branching and merging so that teams can manage probabilistic writers such as AI agents with stronger governance, shared…
E513 · 6 Jul 2026 · 1 hr 1 min
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance depends on contextual intelligence: institutional knowledge, semantic meaning, procedural know-how, and access to the right tools. She also dug into how metadata catalogs are evolving into broader context layers that serve both humans and agents, and why agentic systems are changing the economics of metadata and…
E512 · 18 Jun 2026 · 50 min
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync at scale highlighted both the strengths and common misuses of Kafka. He digs into using events as the source of truth, materialized views with KTables, and how schema registries and type safety prevent downstream breakage. Jevin explains why teams often reach for heavyweight Kafka clusters without leveraging Streams,…
E510 · 1 Jun 2026 · 54 min
Summary In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources. He explores how PuppyGraph lets you run Cypher and Gremlin traversals and graph algorithms directly on data in Iceberg, Delta, Hudi, Hive, and even MongoDB—without loading into a separate graph store. Weimo explains their edge-sharded, vectorized, MPP architecture that tackles hub nodes, multi-hop traversals, and shuffle at scale, targeting sub-second to single-digit-second workloads. He digs into practical graph data…
E509 · 6 May 2026 · 59 min
Summary In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads. He explores Ray’s evolution alongside Kubernetes and PyTorch, and why consolidation at these layers has enabled a new generation of complex, heterogeneous workloads. Robert explains how data preparation has shifted to GPU- and inference-heavy, multimodal pipelines; where Ray fits compared to Spark and workflow orchestrators; and why Ray excels at composing heterogeneous pools of compute, handling failures, and scaling complex…
E508 · 7 Apr 2026 · 59 min
Summary In this episode, I sit down with Gleb Mezhanskiy, CEO and co-founder of Datafold, to explore how agentic AI is reshaping data engineering. We unpack the leap from chat-assisted coding to truly agentic workflows where AI not only writes SQL and dbt models but also executes queries, debugs, runs tests, and ships production-ready outcomes. Gleb explains why teams that master this AI-first loop can see 10–50x gains, how security/compliance concerns can be addressed with platform-native LLM endpoints, and why the role of data engineers is shifting from code authors to operators of…
E517 · 24 Sep 2026 · 50 min
Summary In this episode Christopher Doidge talks about his Agile Ledger Architecture (ALA) approach to data warehousing and how it aims to reduce data debt while shortening the path from raw data to trustworthy business insight. Christopher explained that ALA is not a replacement for existing warehouse patterns like medallion architecture, star schemas, or other modeling approaches, but a complementary discipline focused on pushing business definitions upstream, enforcing cleaner ledger-style transformations, and producing gold-layer tables that stakeholders can actually use without relying…
E516 · 15 Sep 2026 · 53 min
Summary In this episode Soham Mazumdar, co-founder and CEO of Wisdom.ai, talks about what “context” really means in data engineering and AI systems. He explores why context has become such an overloaded term, spanning everything from semantic layers and data catalogs to tribal knowledge, query logs, dashboards, and even agent memory. Soham explained that the big shift is that context is no longer being prepared primarily for human analysts, but for LLMs and agents that can’t reliably fill in missing gaps on their own. That change raises the bar for how context is represented, validated,…
E515 · 27 Aug 2026 · 46 min
Summary In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers. She explored why generic coding assistants often fall short in data workflows, how Otto adds the missing context around Airflow, Astro, upgrades, and troubleshooting, and why Astronomer focused first on high-leverage use cases such as DAG authoring, investigation of pipeline failures, version migrations, and legacy scheduler modernization. She also discussed the practical realities of introducing agents into…
E507 · 29 Mar 2026 · 50 min
Summary In this episode Himant Goyal, Senior Product Manager at Salesforce, talks about how data platform investments enable reliable, accurate metering for consumption-based business models. Himant explains why consumption turns operations into a real-time optimization problem spanning metering, cost attribution, billing, governance, and cross-functional ownership. He explores the richness required in usage data to support sophisticated pricing, the importance of treating metering like a financial system, and the architectural foundations - event schemas, durable ingestion,…