Episode notes
Anas Buhayh discusses multi-stakeholder fairness in recommender systems and the S'mores framework—a simulation allowing users to choose between mainstream and niche algorithms. His research shows specialized recommenders improve utility for niche users while raising questions about filter bubbles and data privacy.
Data Skeptic
by Kyle Polich · English · Tech & Science
The Data Skeptic Podcast features interviews and discussion of topics related to data science, statistics, machine learning, artificial intelligence and the like, all from the perspective of applying critical thinking and the scientific method to evaluate the veracity of claims and efficacy of…
More from Data Skeptic
-
27 Mar 2026 · 39 min
Book Ratings and Recommendations
Goodreads star ratings can be misleading as measures of "book quality," and research from Hannes Rosenbusch suggests that for many professionally published books, differences between readers often matter more than differences between books. The episode also explores how to model reader preferences, why reviews often reveal more about the reviewer than the text, and how LLMs can aid computational literary research while still falling short of human editors in creative writing.
-
10 Mar 2026 · 31 min
Disentanglement and Interpretability in Recommender Systems
Ervin Dervishaj, a PhD student at the University of Copenhagen, discusses his research on disentangled representation learning in recommender systems, finding that while disentanglement strongly correlates with interpretability, it doesn't consistently improve recommendation performance. The conversation explores how disentanglement acts as a regularizer that can enhance user trust and interpretability at the potential cost of some accuracy, and touches on the future of large language models in denoising user interaction data.
-
27 Feb 2026 · 55 min
Collective Altruism in Recommender Systems
Ekaterina (Kat) Fedorova from MIT EECS joins us to discuss strategic learning in recommender systems—what happens when users collectively coordinate to game recommendation algorithms. Kat's research reveals surprising findings: algorithmic "protest movements" can paradoxically help platforms by providing clearer preference signals, and the challenge of distinguishing coordinated behavior from bot activity is more complex than it appears. This episode explores the intersection of machine learning and game theory, examining what happens when your training data actively responds to your…
-
2 Feb 2026 · 27 min
Healthy Friction in Job Recommender Systems
In this episode, host Kyle Polich speaks with Roan Schellingerhout, a fourth-year PhD student at Maastricht University, about explainable multi-stakeholder recommender systems for job recruitment. Roan discusses his research on creating AI-powered job matching systems that balance the needs of multiple stakeholders—job seekers, recruiters, HR professionals, and companies. The conversation explores different types of explanations for job recommendations, including textual, bar chart, and graph-based formats, with findings showing that lay users strongly prefer simple textual explanations over…
-
26 Jan 2026 · 50 min
Fairness in PCA-Based Recommenders
In this episode, we explore the fascinating world of recommender systems and algorithmic fairness with David Liu, Assistant Research Professor at Cornell University's Center for Data Science for Enterprise and Society. David shares insights from his research on how machine learning models can inadvertently create unfairness, particularly for minority and niche user groups, even without any malicious intent. We dive deep into his groundbreaking work on Principal Component Analysis (PCA) and collaborative filtering, examining why these fundamental techniques sometimes fail to serve all users…
-
26 Dec 2025 · 38 min
Video Recommendations in Industry
In this episode, Kyle Polich sits down with Cory Zechmann , a content curator working in streaming television with 16 years of experience running the music blog "Silence Nogood." They explore the intersection of human curation and machine learning in content discovery, discussing the concept of "algatorial" curation—where algorithms and editorial expertise work together. Key topics include the cold start problem, why every metric is just a "proxy metric" for what users actually want, the challenge of filter bubbles, and the importance of balancing familiarity with discovery. Cory shares…
-
5 Oct 2026 · 45 minNew
Implicit Interactions
How do we design robots and autonomous vehicles that understand the unwritten rules of human behavior? Kyle speaks with Cornell Tech professor Wendy Ju about implicit interaction, "Wizard of Oz" prototyping, and what studying pedestrians, self-driving cars, and even robotic furniture can teach us about designing technology that behaves the way people expect.
-
25 Sep 2026 · 34 min
The Lived Informatics Model
The data we collect about ourselves can tell us a lot—but only if the technology collecting it actually fits into our lives. Daniel Epstein explores personal informatics, from fitness trackers and food journals to baby tracking and AI, and explains why abandoning a tracking tool doesn't necessarily mean it failed.
-
9 Sep 2026 · 23 min
Recommender Systems Today and Tomorrow
In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control. From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at what happens when recommender systems must answer not only for what they recommend, but for the consequences of those choices.
-
1 Sep 2026 · 31 min
Recommender Systems Optimization Goals
In part two of the Data Skeptic Recommender Systems season finale, Kyle asks a deceptively difficult question: what should recommender systems actually optimize for? Drawing on conversations from across the season, the episode explores engagement, filter bubbles, popularity bias, fairness, human curation, embeddings, and the growing role—and risks—of large language models in shaping what gets recommended to us.
