Episode notes
In this episode of Data Skeptic, we dive deep into the technical foundations of building modern recommender systems. Unlike traditional machine learning classification problems where you can simply apply XGBoost to tabular data, recommender systems require sophisticated hybrid approaches that combine multiple techniques. Our guest, Boya Xu, an assistant professor of marketing at Virginia Tech, walks us through a cutting-edge method that integrates three key components: collaborative filtering for dimensionality reduction, embeddings to represent users and items in latent space, and bandit…
Data Skeptic
by Kyle Polich · English · Tech & Science
The Data Skeptic Podcast features interviews and discussion of topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of…
More from Data Skeptic
-
26 Jan 2026 · 50 min
Fairness in PCA-Based Recommenders
In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users…
-
26 Dec 2025 · 38 min
Video Recommendations in Industry
In this episode, Kyle Polich sits down with Cory Zechmann , a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares…
-
18 Dec 2025 · 52 min
Eye Tracking in Recommender Systems
In this episode, Santiago de Leon takes us deep into the world of eye tracking and its revolutionary applications in recommender systems. As a researcher at the Kempelin Institute and Brno University, Santiago explains the mechanics of eye tracking technology—how it captures gaze data and processes it into fixations and saccades to reveal user browsing patterns. He introduces the groundbreaking RecGaze dataset, the first eye tracking dataset specifically designed for recommender systems research, which opens new possibilities for understanding how users interact with carousel interfaces like…
-
23 Nov 2025 · 37 min
Designing Recommender Systems for Digital Humanities
In this episode of Data Skeptic, we explore the fascinating intersection of recommender systems and digital humanities with guest Florian Atzenhofer-Baumgartner, a PhD student at Graz University of Technology. Florian is working on Monasterium.net , Europe's largest online collection of historical charters, containing millions of medieval and early modern documents from across the continent. The conversation delves into why traditional recommender systems fall short in the digital humanities space, where users range from expert historians and genealogists to art historians and linguists,…
-
13 Nov 2025 · 33 min
DataRec Library for Reproducible in Recommend Systems
In this episode of Data Skeptic's Recommender Systems series, host Kyle Polich explores DataRec, a new Python library designed to bring reproducibility and standardization to recommender systems research. Guest Alberto Carlo Maria Mancino, a postdoc researcher from Politecnico di Bari, Italy, discusses the challenges of dataset management in recommendation research—from version control issues to preprocessing inconsistencies—and how DataRec provides automated downloads, checksum verification, and standardized filtering strategies for popular datasets like MovieLens, Last.fm, and Amazon…
-
5 Nov 2025 · 35 min
Shilling Attacks on Recommender Systems
In this episode of Data Skeptic's Recommender Systems series, Kyle sits down with Aditya Chichani, a senior machine learning engineer at Walmart, to explore the darker side of recommendation algorithms. The conversation centers on shilling attacks—a form of manipulation where malicious actors create multiple fake profiles to game recommender systems, either to promote specific items or sabotage competitors. Aditya, who researched these attacks during his undergraduate studies at SPIT before completing his master's in computer science with a data science specialization at UC Berkeley,…
-
5 Oct 2026 · 45 minNew
Implicit Interactions
How do we design robots and autonomous vehicles that understand the unwritten rules of human behavior? Kyle speaks with Cornell Tech professor Wendy Ju about implicit interaction, "Wizard of Oz" prototyping, and what studying pedestrians, self-driving cars, and even robotic furniture can teach us about designing technology that behaves the way people expect.
-
25 Sep 2026 · 34 min
The Lived Informatics Model
The data we collect about ourselves can tell us a lot—but only if the technology collecting it actually fits into our lives. Daniel Epstein explores personal informatics, from fitness trackers and food journals to baby tracking and AI, and explains why abandoning a tracking tool doesn't necessarily mean it failed.
-
9 Sep 2026 · 23 min
Recommender Systems Today and Tomorrow
In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control. From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at what happens when recommender systems must answer not only for what they recommend, but for the consequences of those choices.
-
1 Sep 2026 · 31 min
Recommender Systems Optimization Goals
In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for? Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, fairness, human curation, embeddings, and the growing role—and risks—of large language models in shaping what gets recommended to us.
