Skip to content
Melo Podcasts Home
CategoriesLanguagesFollowing

Episode notes

From a professional development perspective, you should always be learning: listening to podcasts, reading books, connecting with internal colleagues, following useful people on Medium and LinkedIn, and so on. Did we mention listening to podcasts? Well, THIS episode of THIS podcast is not really about that kind of learning. It's more about the sort of organizational learning that experimentation and analytics is supposed to deliver. How does a brand stay ahead of their competitors? One surefire way is to get smarter about their customers at a faster rate than their competitors do. But what…

Transcript

Read the transcript · about 12,630 words, follows along as you listen

[Announcer]: Welcome to the Analytics Power Hour. [Announcer]: Analytics topics covered conversationally and sometimes with explicit language. [Tim Wilson]: Hi, everyone. [Tim Wilson]: Welcome to the Analytics Power Hour. [Tim Wilson]: This is episode 290. [Tim Wilson]: I'm Tim Wilson, and I'm joined for this episode by Val Kroll. [Tim Wilson]: How's it going, Val? [Val Kroll]: Fantastic. [Val Kroll]: Excited for today. [Tim Wilson]: Outstanding.

[Tim Wilson]: Unfortunately, we were supposed to also be joined by Michael Helbling for this show, but he's gone all on brand for the winner and gotten the flu. [Tim Wilson]: Luckily, as we're into our 11th year of doing this show now, we've learned a thing or two about rolling with the punches. [Tim Wilson]: And as it turns out, learning is the topic for today's show. [Tim Wilson]: I mean, it's implicit in all forms of working with data. [Tim Wilson]: We're looking at analysis or research or experimentation results and hoping, just hoping that we come out of the experience with a deeper knowledge of something.

[Tim Wilson]: I mean, and hopefully it's something useful, more knowledge than we had before. [Tim Wilson]: It's a simple idea. [Tim Wilson]: Sometimes though, it's a little harder to execute in practice. [Tim Wilson]: That's why we perked up when we came across an article from some folks at Spotify called Beyond Winning, Spotify's experiments with learning framework. [Tim Wilson]: We're excited to welcome one of the co-authors of that piece to today's show.

[Tim Wilson]: Mårten Schultzberg is a product manager and staff's data scientist at Spotify. [Tim Wilson]: He has a deep background in experimentation and statistics, including actually teaching advanced statistics in a prior role for a number of years. [Tim Wilson]: So who better to chat with about learning? [Tim Wilson]: Welcome to the show, Mårten. [Tim Wilson]: Thank you so much. [Tim Wilson]: Excited to be here. [Tim Wilson]: Oh, right. [Tim Wilson]: It's a borderline giddy about the topic as we were diving into our excitement before we hit the record button.

[Tim Wilson]: Yeah. [Val Kroll]: We definitely fought over who got to be on this one. Yeah. [Tim Wilson]: Mårten, in the article that I referenced in the opening, which we're definitely going to link to in the show notes, it's a great read, you and your co-authors make the distinction between a win rate and a learning rate for experimentation. [Tim Wilson]: That's the premise of the article is this win rate. [Tim Wilson]: this learning rate as a proposed metrics or a metric that's actually in use.

[Tim Wilson]: That seems like a good place to start. [Tim Wilson]: Maybe can you explain what you were seeing as a drawback to too much focus on win rate as a metric for experimentation programs? [Mårten Schultzberg]: Yes. [Mårten Schultzberg]: I think it needs to take a little step back. [Mårten Schultzberg]: I think it started with [Mårten Schultzberg]: When we rolled experimentation out at Spotify properly, like at scale 2019-2020, we quite quickly realized that one of the biggest wins that we made over and over again was to detect bad things early and being able to avoid them.

[Mårten Schultzberg]: So using it as a sort of dodge-bullets type of mechanism. [Mårten Schultzberg]: And we have used it like that since. [Mårten Schultzberg]: It's one of the biggest reasons why we run so many experiments. [Mårten Schultzberg]: We want to avoid shipping bad things that happens, you know, unintentionally. [Mårten Schultzberg]: Side effects and stuff like that. [Mårten Schultzberg]: And at the same time, I've seen over the years a lot of blog posts and papers published about win rates from other companies.

[Mårten Schultzberg]: Win rates as in the rate of experiments where you find a variant that is better than the previous variant and you ship it. [Mårten Schultzberg]: So a clear winner. [Mårten Schultzberg]: And so I just felt that it was sort of under celebrated all of the other types of wins that you can make besides finding something that was better than the current version. [Mårten Schultzberg]: And I also think that it doesn't really reflect how most companies, at least the companies I'm familiar with, are actually using experimentation.

[Mårten Schultzberg]: They're using experimentation [Mårten Schultzberg]: partly to optimize things. [Mårten Schultzberg]: So to find winners, to continuously improve something and optimize it. [Mårten Schultzberg]: But that's only one part of that puzzle. [Mårten Schultzberg]: The other part of using it as a mechanism for safety and safety net is something that wasn't, I think, talked about enough. [Mårten Schultzberg]: And so that's sort of where this sprung from.

[Val Kroll]: I love that. [Val Kroll]: And the one thing though that I think is, I would love for you to talk a little bit about more, that I think even if an organization was like, yes, like in spirit, I completely agree with that premise, Mårten. [Val Kroll]: It seems like using a metric like learning rate seems squishy. [Val Kroll]: Like win rate is objective. [Val Kroll]: We can tally that in a column and calculate that percentage. [Val Kroll]: But can you talk a little bit about how you thought about the criteria for determining [Val Kroll]: how you say, yes, we learned something from this experiment or how it's defined.

[Mårten Schultzberg]: Yeah, and so yeah, I want to firstly call out that this was a team effort. [Mårten Schultzberg]: It was a lot of people involved. [Mårten Schultzberg]: It was driven by the central experimentation team at Spotify, but there was also a lot of other data scientists that are actually doing product work that was involved in this discussion. [Mårten Schultzberg]: We had a lot of really good discussions actually about what learning means and when you actually get value from an experiment.

[Mårten Schultzberg]: And so I just want to call that out. [Mårten Schultzberg]: And I think [Mårten Schultzberg]: We see it as there are essentially three ways that you can learn from a Navy test. [Mårten Schultzberg]: One is that you find an obvious winner. [Mårten Schultzberg]: So what other people refer to as win rate. [Mårten Schultzberg]: So you find a version that is better than the current version. [Mårten Schultzberg]: The other one is that you find that the current version was worse somehow, that you detect something bad, that you detect the regression or [Mårten Schultzberg]: Often that can be, you know, not only that users didn't like it, but more that maybe something went wrong with some integration somewhere, so you get latencies increasing or crashes increasing.

[Mårten Schultzberg]: And so those are quite obvious wins, so the finding better stuff and avoiding worse stuff. [Mårten Schultzberg]: And then there is the middle one, which is more nuanced, which is when you run a well-planned experiment and you find nothing. [Mårten Schultzberg]: So a neutral experiment, which is, I guess, vague. [Mårten Schultzberg]: But what we count there as a win is an experiment that actually had a sample size calculation before that did a proper power analysis and said, hey, I want to have a certain power of finding an effect if it exists.

[Mårten Schultzberg]: And then they ran that experiment according to that plan, and they found nothing. [Mårten Schultzberg]: We also view that as a learning, because at that point, they can actually, with the certainty that they hoped for, say, no, there was no effect from this change. [Mårten Schultzberg]: So the neutrality in that case is informative, because you can say, hey, maybe this is not worth pursuing, because we actually ran a proper experiment.

[Mårten Schultzberg]: If there was an effect of the size that we were interested in, or that we hypothesized, we would have found it. [Mårten Schultzberg]: So there are those three cases. [Mårten Schultzberg]: And obviously that middle one, the neutral one, is a little bit more complicated. [Mårten Schultzberg]: It's more complicated to implement or to instrument because you need to know what sample size calculations were run and if the experiment actually met the planned sample size and all of those things.

[Mårten Schultzberg]: Fortunately for us, in our tool, it's fairly easy to do. [Mårten Schultzberg]: But yeah, take some thinking to get that right. [Val Kroll]: I'm literally writing those because there's so many things I want to dig into. [Val Kroll]: But before going to the 5,000-foot view, I guess I'm just curious about the culture change internally. [Val Kroll]: with so many people with access to run experiments and this appetite for experiments, what was it like to get them to shift away from the win rate to this other new metric that you rolled out?

[Val Kroll]: I'm just very curious what that experience was like if there was resistance, if there was excitement, or some people were really questioning it. [Mårten Schultzberg]: There's always people questioning everything at Spotify, which is one of the things that I love about Spotify. [Mårten Schultzberg]: So that's a constant. [Mårten Schultzberg]: But yeah, I think because of the fact that we so early realized that experiment was such a powerful tool to avoid mistakes and to detect bad things early, I think that the sort of common definition of learning was already incorporating that aspect of experimentation.

[Mårten Schultzberg]: I think a lot of people has [Mårten Schultzberg]: sort of over the years learned to, I should not use the word learned, come to appreciate that, yeah exactly, come to appreciate that avoiding something bad is a great learning and something that is super valuable for product development. [Mårten Schultzberg]: So I think that part was not so controversial when we developed this metric. [Mårten Schultzberg]: I think the neutral one is trickier and there's also [Mårten Schultzberg]: It's a much more room there for discussions about what should count, should you be super strict about that it should be exactly powered, should you allow some wiggle room, there's a lot of things that you can discuss there.

[Mårten Schultzberg]: We were eager to get a very clear and explicit definition out and we were also eager actually to write about it externally because we were hoping that other companies would, and I guess this podcast is a good example of that too, that we could have this discussion because I think it's been [Mårten Schultzberg]: I'm really curious how other people think about this. [Mårten Schultzberg]: I'm not convinced that our definition of learning is like the ultimate one or the final one or anything, but I think it's a good first step away from the more naive, only wins count definition.

[Tim Wilson]: The raging cynic in me would be, well, gee, if people realized that's a way to game the metric would be to run really inconsequential small [Tim Wilson]: tests, which at the same time, the analyst in me thinks that, yeah, that happens with analytics a lot, that you're kind of digging in and trying to find something. [Tim Wilson]: You're like, well, somebody thought there would be some relationship here and we're just not seeing it. [Tim Wilson]: And that can be equally unsatisfying for the analyst.

[Tim Wilson]: So like, how do you think about [Tim Wilson]: neutral being, we were trying something that did have a legitimate chance of being meaningful. [Tim Wilson]: And maybe this kind of bridges to another article that you wrote, which is, you know, like, how do you say neutral, but not have neutral become a crutch for, yeah, we're essentially doing AA tests and, you know, giving ourselves two thumbs up on the learning rate. [Mårten Schultzberg]: That's a great question.

[Mårten Schultzberg]: I think we've been thinking a lot about what a healthy distribution should look like. [Mårten Schultzberg]: A healthy distribution of different types of winds and also the proportion of neutral experiments. [Mårten Schultzberg]: And I think that's actually a super interesting topic. [Mårten Schultzberg]: I think depending on what kind of strategy you have here from a product side, you can want to have different distributions.

[Mårten Schultzberg]: So for example, if you take the [Mårten Schultzberg]: If we wait with the neutral one, because it's maybe a trickier one, but if we think about how many experiments should you find regressions in that you dodge versus the win rate, how should that distribution look? [Mårten Schultzberg]: Well, that will depend on a lot of things. [Mårten Schultzberg]: But if you're a company that has everything to win and little to lose, then maybe you can afford to have a high rate of just trying stuff.

[Mårten Schultzberg]: Because whenever you find a win, it's going to be quite big because you're still in early stages, whereas if you're [Mårten Schultzberg]: If you're a product that is already very mature, then maybe you have other goals for those things. [Mårten Schultzberg]: It's a super interesting discussion to have, and that's one of the discussions we're having now with teams at Spotify and other people that are using our experimentation tooling.

[Mårten Schultzberg]: What should we do with this information? [Mårten Schultzberg]: And what's good and what's bad? [Mårten Schultzberg]: And I think it's different for different parts, even of organizations within Spotify. [Mårten Schultzberg]: What's good, depending on how they're looking at it. [Mårten Schultzberg]: But for sure, we wouldn't look at the learning rate only. [Mårten Schultzberg]: So we would say we want the learning rate to be reasonable.

[Mårten Schultzberg]: But then we, of course, should probably aspire for having a high win rate. [Mårten Schultzberg]: That's nothing bad in itself. [Mårten Schultzberg]: But at least if we have a high learning rate, we know that we're not wasting our experimentation efforts. [Mårten Schultzberg]: We know that experiments we're running, we're actually learning from. [Mårten Schultzberg]: If we're running a ton of experiments that are not powered and neutral, then we will never be able to say these things didn't have an effect.

[Mårten Schultzberg]: We can't separate between if these things didn't have an effect or if we just didn't run a good enough experiment to detect it. [Mårten Schultzberg]: So on the one hand, you look at the learning rate and say like, hey, we want to utilize our experimentation bandwidth really well. [Mårten Schultzberg]: So we want to have a high learning rate at all times. [Mårten Schultzberg]: But then at each quarter, you can look at this metric and the distribution of these outcomes and say like, hey, you know what, we're dodging a lot of bullets, but we're almost never finding something good.

[Mårten Schultzberg]: Should we rethink our strategy or even more, if we're finding a ton of neutral results and we see more and more neutral results in some part of the organization, maybe we're hitting diminishing returns and we should [Mårten Schultzberg]: try something different. [Mårten Schultzberg]: Maybe we found some kind of local optima, maybe, or something like that. [Mårten Schultzberg]: So I think it can be a quite strategic instrument if you have all of these, the distribution of all of these outcomes as part of the learning metric.

[Michael Helbling]: You know what's worse than writing SQL? [Michael Helbling]: Probably writing that same SQL for the third time because you forgot where you saved it. [Tim Wilson]: or explaining to an LLM for the 10th time that your GA4 medium field is a mess because three different interns had three different naming conventions. [Michael Helbling]: Yeah, like organic, organic underscore social or, I mean, it's like a crime scene of good intentions.

[Tim Wilson]: Which is why Askwise Skills feature really helps. [Michael Helbling]: Record that data cleaning nightmare once as a skill, reuse it across different datasets, portable expertise, and their jam memory system remembers context, like the July data is doubled or use the product table, not staging. [Michael Helbling]: Exactly. [Tim Wilson]: It's context focused, not just code focused. [Tim Wilson]: Plus your data never touches the LLM. [Tim Wilson]: Semantic layer generates code that runs locally.

[Michael Helbling]: where your data presumably won't judge you for that medium field situation. [Michael Helbling]: We can hope. [Michael Helbling]: We're going to ask-y.ai. [Michael Helbling]: That's ask-y.ai. [Michael Helbling]: Use code APH to jump the wait list and stop paying the context switching tax. [Val Kroll]: That's making me think as you were talking about that, that like even within an organization, like you were saying like companies who have everything to gain or you know, I think everything to [Val Kroll]: Nothing to lose.

[Val Kroll]: I forget exactly. [Val Kroll]: I never get that right. [Val Kroll]: Well, apparently I can't either. [Val Kroll]: But even within Spotify, thinking about the different product teams that if it's a group that's working on the cancellation flow and thinking about retention, they're probably having very different distribution of those outcomes as their goals or targets versus [Val Kroll]: playlist creation which is like such an established user pattern is that like how you customize some of those conversations from like the center of excellence experience like perspective to kind of consult with those teams.

[Mårten Schultzberg]: Yeah, let's say so, but I also add that there's a lot of centers of excellence when it comes to experimentation at Spotify. [Mårten Schultzberg]: Fortunately, we have many parts of the organization that have super strong experimentation organizations or champion groups or [Mårten Schultzberg]: nerds. [Mårten Schultzberg]: I like to think about it. [Mårten Schultzberg]: I mean, look who's talking. [Mårten Schultzberg]: But anyway, no, but so I think, so that discussion happens locally in a lot of places and a lot of people are having those discussions.

[Mårten Schultzberg]: So it's not like sometimes we get, you know, questions about how to think about things. [Mårten Schultzberg]: And also, one interesting aspect of this metric is that sometimes you might find that [Mårten Schultzberg]: You know, if you're actually, we didn't talk about, there's one outcome here that we didn't talk about, which was the, when you get an invalid experiment where something is wrong with the setup of the experiment.

[Mårten Schultzberg]: That's the final sort of outcome in this learning framework. [Mårten Schultzberg]: So you didn't learn because something went wrong. [Mårten Schultzberg]: For example, something went wrong with integration. [Mårten Schultzberg]: Maybe you got imbalanced treatment and control group assignment for some reason, or you don't get all of the data that you should get or something like that. [Mårten Schultzberg]: And that's of course an outcome that [Mårten Schultzberg]: is the least fun one, so to speak.

[Mårten Schultzberg]: It's just like, yeah, we couldn't get this integration to work well enough. [Mårten Schultzberg]: So we have used that one and worked really hard on getting that to as close to zero as possible. [Mårten Schultzberg]: We want it to be possible for anyone to run a really high quality experiment. [Mårten Schultzberg]: With Spotify running experiments on so many different devices and apps and combinations of those, it's really tricky to always nail those things, but it's obviously an important signal.

[Mårten Schultzberg]: So whenever that one is high, that's something that teams come to us with and say like, hey, we don't get our integration to work as well as we want to, how can we improve these things? [Mårten Schultzberg]: And also when it comes to the neutral aspect, [Mårten Schultzberg]: the quality of the sample size calculator starts mattering a lot. [Mårten Schultzberg]: So whenever someone sets up an experiment and we try to predict what sample size they need, it's a prediction, right?

[Mårten Schultzberg]: We're looking at historical data saying like, yeah, well, given how historical data has moved, the variation in that data and the means and the treatment effects that you say you're interested in finding, we think that you need to run your experiment for this long to reach this many users. [Mårten Schultzberg]: And that's a prediction that takes a lot of things into account, but it can always be improved probably. [Mårten Schultzberg]: So that's also a conversation that we sometimes have when people are like, in our use case, the sample size calculator is not good enough.

[Tim Wilson]: But that is a case where you, that's one where you would come back. [Tim Wilson]: Like what is the scenario where you run it, they've got a MDE, they've got the estimates, you've got the sample size calculator, it says run this. [Tim Wilson]: If it comes back, I'm trying to understand the distinction between, actually, we probably just didn't run this long enough versus, well, for what we ran and what parameters we put in, it's a neutral result.

[Tim Wilson]: Is there a distinction there? [Mårten Schultzberg]: I can speak a little bit to it. [Mårten Schultzberg]: In practice, when we do the sample size calculation, I don't know how technical and nerdy I'm allowed to get here, but given the name of this podcast, I'm going to go deep. [Tim Wilson]: We don't want to hit if Matt Gershauf would [Tim Wilson]: have to think about it for a minute, that's a little bit too technical. [Mårten Schultzberg]: No chance, no chance.

[Mårten Schultzberg]: This is bread and butter for him, promise. [Mårten Schultzberg]: No, so we never know the variance of the treatment group, right, before we run the experiment. [Mårten Schultzberg]: We can always just think like maybe it will be a homogeneous treatment effect, or we could, I suppose, speculate about how the treatment will affect the variance, but it's always [Mårten Schultzberg]: gnarly, it's difficult to do. [Mårten Schultzberg]: So what we do always in practice, I think everyone essentially is saying like, let's presume that the treatment effect is homogeneous.

[Mårten Schultzberg]: In practice, of course, when we start running the experiment, maybe the treatment effects only part of the treatment group, which will then disperse the distribution. [Mårten Schultzberg]: If we have a beautiful distribution to start with, but some people get the large treatment effect, you will make the variation of that distribution larger. [Mårten Schultzberg]: So the variance in that group will be larger. [Mårten Schultzberg]: So the required sample size will go up.

[Mårten Schultzberg]: We do, in confidence in our experimentation tool, we do both. [Mårten Schultzberg]: So we have the pre-experimentation sample size calculator, which uses historical data to make this prediction. [Mårten Schultzberg]: And then during the experiment, we're also collecting the data from the experiment and running the subsize calculation continuously. [Mårten Schultzberg]: I actually wrote a paper about that. [Mårten Schultzberg]: I think there is a blog post about that too.

[Mårten Schultzberg]: If someone wants to nerd in on that, that it's actually valid to look at the [Mårten Schultzberg]: power during the experiment. [Mårten Schultzberg]: It's a peaking that is non-problematic. [Mårten Schultzberg]: You can look at that. [Mårten Schultzberg]: Anyway, so you have those and you might have a big discrepancy then. [Mårten Schultzberg]: So when you start the experiment, you might think that, hey, I can run this for two weeks.

[Mårten Schultzberg]: I will reach my whatever 10,000 users that I need. [Mårten Schultzberg]: But then when you run it for a week, you realize that like, no chance. [Mårten Schultzberg]: I will reach much less or I will need much more, maybe more likely. [Mårten Schultzberg]: I thought I needed 10,000, maybe I need 40,000. [Mårten Schultzberg]: And that's just not possible given the traffic that I have on this page. [Mårten Schultzberg]: And in that case, it might be done a conversation about like, hey, how can we make this better?

[Mårten Schultzberg]: And so one way that we do it in practice is that we say like, okay, maybe instead of us trying to predict it, you can point to a similar experiment. [Mårten Schultzberg]: If you know you have a similar experiment, we're changing the same kind of thing. [Mårten Schultzberg]: But yeah, it's a tricky thing. [Mårten Schultzberg]: It's a truly difficult problem to make good sample size estimations. [Val Kroll]: And one thing that I found interesting, because there's definitely like two different camps here, is that I hopefully I'm not putting, correct me if I've interpreted this incorrectly, that you do allow for multiple success metrics in this, which I know makes that a little bit more complicated.

[Val Kroll]: And I think it also talked about adequately powered guardrail metrics, deterioration metrics, quality metrics, which [Val Kroll]: not a lot of organizations do or have the capability to do, but that was like, oh, well, definitely enough to talk a little bit about that. [Val Kroll]: But how do you handle the multiple success metrics, especially if you're looking at things further into the funnel that have a lower incidence? [Val Kroll]: How do you think about that layer?

[Mårten Schultzberg]: Yeah. [Mårten Schultzberg]: This is a rich topic. [Mårten Schultzberg]: We have a framework for this that we have developed over the years. [Mårten Schultzberg]: And it's also a paper that is, I think, about to be published. [Mårten Schultzberg]: It's an archive, at least, where we go through exactly all of the details of how we're handling the multi-metric that we call decision framework, statistically. [Mårten Schultzberg]: But I can give the short version of it.

[Mårten Schultzberg]: So essentially, what we're saying is that we have an explicit decision rule for the multi-metric setup. [Mårten Schultzberg]: So we have success metrics and guardrail metrics. [Mårten Schultzberg]: So success metrics are metrics that you want to improve, and guardrail metrics are metrics that you don't want to harm. [Mårten Schultzberg]: And so, for example, at Spotify, maybe we want to improve the music consumption, but we don't want to harm the podcast consumption.

[Mårten Schultzberg]: We don't want to do it at the expense of podcast, for example. [Mårten Schultzberg]: So if you're making a new music recommendation algorithm, you don't want to harm any other consumption. [Mårten Schultzberg]: And so the decision rule is essentially that at least one of the success metrics should have improved and none of the guardrail metrics should have been harmed. [Mårten Schultzberg]: There are a lot of nuance here, because for the garter and metrics we're using so-called non-inferiority tests, which makes everything much more complicated to talk about, but leaving that aside, it means that when we're talking about power and false positive rate, we're talking about the false positive rate and the power for that decision rule.

[Mårten Schultzberg]: So we're saying we want that decision that we would make based on this rule. [Mårten Schultzberg]: So at least one of the success metrics are significantly better, and none of the garter and metrics are worse. [Mårten Schultzberg]: We want that to be the false positive rate of intent, and we want to have the power to detect given the sample size. [Mårten Schultzberg]: So we have to make the adjustments for multiple testing corrections accordingly, and then we have to make the power and sample size calculations accordingly.

[Mårten Schultzberg]: things to fiddle with there. [Mårten Schultzberg]: But in principle, since the guardrail metrics all have to be not harmed, they are not giving you additional chances of succeeding, so you don't have to correct for them in the same sense. [Mårten Schultzberg]: But at the same time, you have to power them simultaneously, because all of them has to show simultaneously that they weren't harmed if you're using non-inferiority tests.

[Mårten Schultzberg]: I'm deliberately avoiding going in too much to know if you were to test because it's like such a tongue twister to talk about. [Mårten Schultzberg]: But if you're interested in... You still said it eight times. [Tim Wilson]: Good. [Mårten Schultzberg]: Yeah, no, but that was... No, but it's tricky. [Mårten Schultzberg]: So, yeah, so that's how we do those things. [Mårten Schultzberg]: So it's a bit messy, but... [Val Kroll]: So back to the culture side of this, how do you coach product teams to not just pick 50 success metrics?

[Val Kroll]: Because they are so excited about this new feature. [Val Kroll]: It came from up high, and we really want this to, we want to find some success. [Val Kroll]: And obviously, there's a statistical part of it, like the correction, but culturally, how do you guide that conversation away from? [Val Kroll]: No, it shouldn't be like a pick list of up to 75 metrics to find something that went quote unquote up. [Mårten Schultzberg]: Yeah, yeah, yeah.

[Mårten Schultzberg]: No, I mean, this is a conversation that we have. [Mårten Schultzberg]: I think it's Spotify. [Mårten Schultzberg]: It has settled, but like this is a conversation that we have from time to time. [Mårten Schultzberg]: And I think it's a [Mårten Schultzberg]: It's a sort of healthy discussion to have because it's not... I think this is more tricky than it might seem. [Mårten Schultzberg]: I want to give the answer that, no, but of course you should just have a discussion and decide on the metrics.

[Mårten Schultzberg]: I'll come back to that because that's ultimately what we do a lot at Spotify, but there is more to it. [Mårten Schultzberg]: There is also the fact that we're making a lot of changes and we are truly interested in any kind of effect that it has. [Mårten Schultzberg]: It's a true statement that if actually this change that I made affected a metric that I didn't think about, like some weird metric, weird from my perspective metric, if that was truly the case, I would want to know.

[Mårten Schultzberg]: So from one perspective, I can really understand this. [Mårten Schultzberg]: I want to look at all of the metrics and just see which one that I affected. [Mårten Schultzberg]: But then on the other hand, you get this obviously super hard problem of like, [Mårten Schultzberg]: cursor dimensionality type issues here where you're looking from too much, so you're either just going to find noise, or you're going to find noise, and then you have to control that, and then you're going to have, instead, very low power to find things.

[Mårten Schultzberg]: But I think there is merit to [Mårten Schultzberg]: the type of experiments where you're just like, I just want to see what happens when I do this. [Mårten Schultzberg]: And I don't really care. [Mårten Schultzberg]: Of course, I care what it is that happens, but I am ultimately interested in all things. [Mårten Schultzberg]: But in practice, of course, this is hard. [Mårten Schultzberg]: So again, at Spotify, it's not like the central experimentation team, which I'm part of.

[Mårten Schultzberg]: building the tooling, we are not dictating these things. [Mårten Schultzberg]: It's rather the other way around that we are, I like to think about it as that we are sort of cultivating what the teams that are doing experimentation are thinking about this. [Mårten Schultzberg]: So we have a lot of discussions with them. [Mårten Schultzberg]: So the way it works as Spotify is that we don't decide the defaults and how things should work in the platform.

[Mårten Schultzberg]: It's rather that we talk to all of the product teams that are experimenting the 300 teams in various forms and then [Mårten Schultzberg]: We collect what they're saying, and we're refining it, and then we're putting that into the tool. [Mårten Schultzberg]: So when it comes to this, how many metrics you should have, there's not one answer at Spotify. [Mårten Schultzberg]: It's different in different parts of the organization.

[Mårten Schultzberg]: But in most of these parts, there have been very explicit conversations where people have talked about, like, hey, how should we trade off here? [Mårten Schultzberg]: actually getting super high precision in the things that we know we're interested in versus getting interesting insights and stuff that we could be interested in. [Mårten Schultzberg]: And this is sort of traded off in various parts of the organization and in various projects, depending on [Mårten Schultzberg]: how and what stage those projects are.

[Mårten Schultzberg]: If it's like a very new product, then you probably see, or you often see experiments with much more metrics because you're just interested in understanding what happens when we ship something like this, what kind of behavioral changes does this cause? [Mårten Schultzberg]: Whereas when we're optimizing something, then we're like, okay, we know pretty well what we need to measure here to do this and to optimize this in a healthy way.

[Tim Wilson]: to Spotify, massive user base, a lot of the ability to design, to try to cover and still be sufficiently powered seems doable. [Tim Wilson]: I'm thinking of a client we had that was in that same boat. [Tim Wilson]: It still feels like the risk, the slippery slope, [Tim Wilson]: fishing expedition of let me tell myself a story that I just want to see if it impacted anything. [Tim Wilson]: And the understanding required that if you go on a fishing expedition, you are, I think, if I understand correctly, your false positive rate can go way up because you detect noise as a signal, which then when you detect it, you get really excited.

[Tim Wilson]: Nobody can rationalize why this metric changed. [Tim Wilson]: It turns out it was noise. [Tim Wilson]: Now we've wound up doing negative. [Tim Wilson]: We've learned something incorrect potentially, unless you have the discipline to say, if we're going to chase that, we need to come up with a theory, and we need to have the rigor to validate that theory before we accept it as fact. [Tim Wilson]: That just feels [Tim Wilson]: coming from an analytic side, similar sort of thing.

[Tim Wilson]: If I just point the machine at all the data and it finds anomalies or finds patterns, there's a very good chance that it's detecting noise that just happened to hit at a point where it can show some statistical merit. [Tim Wilson]: Somehow, some part of me is just terrified. [Tim Wilson]: While I love getting comfortable with [Tim Wilson]: We looked for X, we did not find X. That is still a learning and let's work with our business partners to acknowledge that's a learning and not have them just chasing for everything.

[Tim Wilson]: That also feels like a challenge, you know? [Mårten Schultzberg]: Yeah, no, no, I mean, I agree with everything you say, but I also feel like, I mean, I have the same uncomfortable feeling in my body when I think about this, like, let's look at all of the metrics from a statistics perspective. [Mårten Schultzberg]: But I also just like, I really want to, I also think it's a cop out, not projecting on you now, Tim, but for myself to say like, you know, to say like, you know, [Mårten Schultzberg]: We can only look at the metrics that we decided before because we decided that we found nothing.

[Mårten Schultzberg]: Let's move on because it's also obviously true to me somehow, even though I can't come up with this is how you should do it and this is how it won't lead to these incorrect learnings that you mentioned. [Mårten Schultzberg]: But it feels like it's a hard argument to make when someone says, yeah, but I looked at some other metrics and I learned something. [Mårten Schultzberg]: And then you're like, maybe you did, maybe you didn't.

[Mårten Schultzberg]: And I can think about ways that you could do this. [Mårten Schultzberg]: You could do sample splitting and stuff. [Mårten Schultzberg]: You could take one part of the sample and look for groups. [Mårten Schultzberg]: And then you could validate those findings in another part of the sample and stuff like that to make it much more plausible. [Mårten Schultzberg]: Again, you would have the issue then of having lower powers actually find things, or lower precision at least.

[Mårten Schultzberg]: I just don't want to be too much of a... [Mårten Schultzberg]: Curious? [Mårten Schultzberg]: Yeah, or like a grumpy statistician kind of person. [Mårten Schultzberg]: But I do, I mean, I agree. [Mårten Schultzberg]: I have the same feeling and I haven't seen anyone do it well. [Mårten Schultzberg]: So what I've seen is that people have used the argument of saying like, yeah, we must be able to be able, you know, it must be possible to learn more and then just throw all of the metrics at it.

[Mårten Schultzberg]: And then I think they're just as well. [Mårten Schultzberg]: Like that's just as bad as not doing it, I think. [Mårten Schultzberg]: So I don't have an answer to it, but [Mårten Schultzberg]: Maybe someone smart listens and then they can call me. [Val Kroll]: Yeah. [Val Kroll]: Let us know in the comments. [Tim Wilson]: I mean, I, I mean, I, and I don't know that this is the answer, but I'm, it does feel like, well, if you, if you throw that at it and you find something figuring out how to have the, the step, which is probably a combination of a data scientist or a statistician with the product manager to say, we need to come up with a, [Tim Wilson]: plausible theory as to what's causing that surprising thing.

[Tim Wilson]: And we need to have somebody with their bullshit meter turned on. [Tim Wilson]: Cause I mean, I've certainly watched people find things and they come up with a bullshit theory. [Tim Wilson]: They're like, well, this is clearly happening. [Tim Wilson]: Cause obviously like left-handed people, when they're in the Southern hemisphere, it makes sense that they would prefer the color blue, you know, and something that's, [Tim Wilson]: It's a theory that fits the data, but it's not a theory that holds up to human scrutiny.

[Mårten Schultzberg]: I think one thing that I'm excited about is replication. [Mårten Schultzberg]: I think if you have a streamlined enough way to run experiments and you have your velocity throughput for experimentation high, then one true possibility here is to replicate, to just say like, okay, [Mårten Schultzberg]: I looked broad and deep here and I found something. [Mårten Schultzberg]: I believe in it. [Mårten Schultzberg]: I think I've made my people in the sudden hemisphere argument, but I believe it.

[Mårten Schultzberg]: And then for anyone who would say, I believe in it to the extent that I will now launch a new experiment, take 10% other people or a new random sample and run it again with only that metric or only the new metrics that I care about. [Mårten Schultzberg]: And if I can repeat it, then I will ship it. [Mårten Schultzberg]: Then I would be like, yeah, go for it. [Tim Wilson]: Or potentially, if the theory is, well, it was this kind of incidental thing that happened to be part of it, but it wasn't the core focus when we run an experiment where I've doubled down on that to say this should now [Tim Wilson]: I should now really detect a strong signal because it's backing that up.

[Mårten Schultzberg]: That sort of touches a little bit on the other blog post that you mentioned that has to do with what the intent with an experiment is. [Mårten Schultzberg]: I haven't really talked about it yet. [Tim Wilson]: Let's talk about that one. [Tim Wilson]: Boy, I got giddy on that one too. [Mårten Schultzberg]: Should you want me to give the TLDR on that one too? [Tim Wilson]: Yes, please do. [Mårten Schultzberg]: Yeah, so the idea with that one is I often have like it has sort of come from a lot of the conversations that we've had with people running experiments, talking about the learning framework.

[Mårten Schultzberg]: And then people are like, hey, we have a lot of neutral experiments here. [Mårten Schultzberg]: We run high quality experiments, but we don't find things. [Mårten Schultzberg]: And so one thing that I've sort of identified from working with teams that Spotify, but also externally other companies is that [Mårten Schultzberg]: People are often sort of starting to optimize the idea in their head before they've tried that the idea is at all something that will affect the user users.

[Mårten Schultzberg]: And so what I mean by that is that people are, you know, when they identify something that they think is like, this is important for our users, like, let's, let's use a stupid example, like a button color or something, you know, like we think it's important. [Mårten Schultzberg]: And then immediately, instead of saying, OK, we should first answer the question, is it important or not? [Mårten Schultzberg]: Do users care or not?

[Mårten Schultzberg]: Instead, they immediately start thinking about which color is the best. [Mårten Schultzberg]: And so they jump from, we have no idea if people care about this to having the conversation about which color is the best. [Mårten Schultzberg]: So sort of presuming that people care at all which color this has, besides having a high enough contrast so you can see it. [Mårten Schultzberg]: And so this blog post was me just trying to formulate that, like the distinction between identifying if an aspect of your user experience is something that you can optimize if it has sort of an effect on users in any way, people care about it on the one hand, and optimizing that once you have identified that it's something that people care about on the other hand.

[Mårten Schultzberg]: So sort of identifying something versus optimizing something. [Mårten Schultzberg]: And so I think that this thing that we talked about now is a little bit about maybe if you run an experiment where you thought something, you thought that it was important with some aspect, or you tried to optimize it, and then you find something [Mårten Schultzberg]: something new, some metric that you didn't anticipate to move. [Mårten Schultzberg]: That might cost the sort of idea in your head to be like, hey, maybe there is a mechanism here that people care about.

[Mårten Schultzberg]: Maybe people actually care about how many items we show on this screen. [Mårten Schultzberg]: I was thinking about the ranking, but as a side effect of that, we showed more things. [Mårten Schultzberg]: So we saw that, I don't know, lower down that the list clicks increased or something like that. [Mårten Schultzberg]: And maybe that's an indication that this is a mechanism that people care about. [Mårten Schultzberg]: I think this going in between the states of identifying something to optimize and optimize the thing you have identified and doing that explicitly and deliberately is something that a lot of product teams would benefit from.

[Mårten Schultzberg]: It's easy to fall in the trap of trying to do both at once, I think. [Tim Wilson]: Totally. [Tim Wilson]: Is a cousin to the optimizing [Tim Wilson]: I mean, the framing of say, which is kind of a, I think it might even be in the article, like the case for taking a bigger swing, take the big swing first, make sure that connects, even if it's a, you know, in a while, it's like, yes, there's something here, now we can tune it.

[Tim Wilson]: And that, I think of it from a, I mean, from a marketing analytics perspective, where companies will say, [Tim Wilson]: Let's just try it out and see what happens. [Tim Wilson]: It's kind of a death knell. [Tim Wilson]: It's going to be an underinvestment in a new channel or a new tactic where logically, it's going to be really hard to detect a signal because it winds up getting kind of tempered down to a pretty subtle change. [Tim Wilson]: The logic is, well, if this thing actually matters, then we can make [Tim Wilson]: a nominal investment and we'll see this outsized lift as opposed to saying, does this matter at all?

[Tim Wilson]: Double down on it for some period of time. [Tim Wilson]: Go hard. [Tim Wilson]: See if you actually see something and then say, okay, we definitely need to be in this channel or using this tactic or doing this to the user experience. [Tim Wilson]: Now we need to sort of figure out [Tim Wilson]: Did we actually spend twice as much as we needed to, we can get the same? [Tim Wilson]: Where are the diminishing returns? [Tim Wilson]: It does feel like culturally it's a tough, human nature is risk averse.

[Tim Wilson]: Saying, try something and know that you'll find that it is okay to find that it [Tim Wilson]: didn't work. [Tim Wilson]: A big swing with a neutral result feels like it has a lot more merit than a little small tap with a neutral result. [Tim Wilson]: That's the fun in that. [Mårten Schultzberg]: That's precisely it. [Mårten Schultzberg]: This actually what provoked me to write it was discussions about the neutral outcome in the learnings framework where people are like, [Mårten Schultzberg]: people are like, yeah, but neutral is no fun.

[Mårten Schultzberg]: I don't care if it was powered or not. [Mårten Schultzberg]: I don't want neutral. [Mårten Schultzberg]: And that got me thinking, well, if you don't like the neutral result, it means that the question you posed wasn't interesting enough. [Mårten Schultzberg]: Because I would be like, if I'm convinced as a product person that people care about this thing in our app, if I change this, people are going to care. [Mårten Schultzberg]: And then I make a drastic [Mårten Schultzberg]: change and nobody cares.

[Mårten Schultzberg]: I've run the experiment, I have high precision in my estimates and nobody cares. [Mårten Schultzberg]: If that's not the learning to be excited about, I don't know what is, to be honest. [Mårten Schultzberg]: That really shows that I'm 100% off with my understanding of what people care about, which is truly strong learning. [Mårten Schultzberg]: But on the other hand, this change that I made was like, yeah, I really think our users care about this aspect and I made a minuscule change to it and I didn't find anything.

[Mårten Schultzberg]: I might think for a long time about if this was the right change that I made, or if it was... You just get stuck in weird things. [Mårten Schultzberg]: But one way that I have sort of sold that, because I agree that people are risk-averses to run both. [Mårten Schultzberg]: If you run a neighbor test, people tend to want to be like... But I think I know what users like. [Mårten Schultzberg]: I want to go for the identify and optimize at the same time version of this thing, where...

[Mårten Schultzberg]: I try to choose the right value for my customers or my users. [Mårten Schultzberg]: But I also say, just also add then, if you haven't actually identified that this is something that people care about or that matters for your business or where it might be. [Mårten Schultzberg]: Add the [Mårten Schultzberg]: more sort of provocative version. [Mårten Schultzberg]: I call it maximum viable product, I think, because of course, this has to be reasonable.

[Mårten Schultzberg]: If you make some button larger than the screen, then of course, you're going to see some change. [Mårten Schultzberg]: So it has to be within the limits from what is still a usable function, but that is still extreme. [Mårten Schultzberg]: So the maximum change that you think is like, but this is still, this is not [Tim Wilson]: You're saying doing that within kind of a multivariate, say we've got our control, we've got what the optimized and identified at the same time version, and then we have an identify only version.

[Tim Wilson]: And it's okay if that identify version detects like the biggest effect, you can say, yeah, that was kind of hedging to make sure that [Tim Wilson]: that we got something out of it. [Tim Wilson]: And if that one that was identified and optimized simultaneously didn't, then we're probably still on a good track. [Tim Wilson]: It just turns out we're not so omniscient that we can come up with the perfect variant in one shot. [Mårten Schultzberg]: I think it's smart also from, I mean, a lot of companies at least, Spotify and other companies that I work with, they're all struggling with having big enough sample size, right?

[Mårten Schultzberg]: both because they have limited traffic, but also because they're interested in small effects, generally speaking. [Mårten Schultzberg]: But the nice thing about making a very drastic change is that it should have a large effect. [Mårten Schultzberg]: If you're making this maximum viable change, then that should cause a large effect. [Mårten Schultzberg]: So you should be able to say, yeah, but now I pull this lever as hard as it's possible to pull.

[Mårten Schultzberg]: So this should cause maybe 5% change, like whether it's good or bad. [Mårten Schultzberg]: And so you can maybe run smaller experiments. [Mårten Schultzberg]: If you're in a situation where it's hard for you to know what you should, like you have a hard time finding bandwidth essentially for optimizing things, then I think it's a smart idea to do these more drastic changes to identify what you should then spend larger experiments on optimizing.

[Mårten Schultzberg]: Because the truth is that when you start optimizing, even if it's a nice convex surface for this thing, button size or something, the closer you come to the optimum there, the larger samples you're going to need to be able to identify those steps. [Val Kroll]: It seems like the framing that I really liked in this article is the building the right thing versus building the thing right. [Val Kroll]: And it feels like the stakes couldn't be higher in everything you guys are just talking about in a product context because it's not just about changing a button color.

[Val Kroll]: In a lot of cases, this isn't about UX. [Val Kroll]: It's about adding additional features or different capabilities. [Val Kroll]: and you're hoping to impact things like customer lifetime value, not just did they get to the next screen, right? [Val Kroll]: So it's not just like checkout flows, right? [Val Kroll]: I think I was thinking about this. [Val Kroll]: I've actually spent more time than the average human should thinking about the changes that have been happening lately inside of my United app.

[Val Kroll]: So I'm United Loyal, I fly United and the app has been changing a ton lately. [Val Kroll]: And we went from, there was one place where I could change my seat to every single screen within this app. [Val Kroll]: I can change my, which I do appreciate. [Val Kroll]: I'm definitely someone who loves feeling a lot of control over changing my seat. [Val Kroll]: But I'm like, what were the conversations that happened internally that said, you know what?

[Val Kroll]: The user needs to be able to change their seat while they're checking in their bag, while they're checking to see what gate their flight is at. [Val Kroll]: Anyway, just to bring this back to an actual question, building the thing right, and maybe the feature is great, the new functionality that you're adding, but maybe you have gone about it the wrong way, which has impacted the ability for someone to understand [Val Kroll]: What exactly this is capable of?

[Val Kroll]: Maybe it was a micro copy issue, or maybe it was in the wrong place in the flow, which feels more like optimization. [Val Kroll]: Even though this framing and big swing versus small change, that sounds really objective. [Val Kroll]: If you put them side by side, that's clear. [Val Kroll]: I'm especially interested because now you are in a product role to get a little meta about it. [Val Kroll]: How do you think about [Val Kroll]: what is, when would you ever recycle a concept in a different context?

[Val Kroll]: Because it does feel like the optimization killed your ability to understand if it was viable. [Mårten Schultzberg]: The truth is here that this is difficult. [Mårten Schultzberg]: I think especially starting with that building the thing right versus building the right thing. [Mårten Schultzberg]: Some things you have to do quite a lot of building to even check if it's the right thing. [Mårten Schultzberg]: If you're building a new feature, there might be a lot of things that you have to get in place to even see if it's something that people cares about.

[Mårten Schultzberg]: once you've seen that they care about it, like maybe they don't like it. [Mårten Schultzberg]: And that's because you haven't built it right yet. [Mårten Schultzberg]: So like, I mean, it's this is a very stylized blog post, of course. [Mårten Schultzberg]: But the truth is, is much more muddy. [Mårten Schultzberg]: So yeah, I mean, in practice, [Mårten Schultzberg]: I think that one of the things that have been discussed a lot that Spotify and other places is like, okay, but with experimentation, where is the room for the product intuition and making bets on things and stuff like that?

[Mårten Schultzberg]: And I've always liked to say that these are completely [Mårten Schultzberg]: uh you know they're they're augmenting each other they're helping each other like it's no they you can make you can have this strong intuition still and you can make these bets what experimentation helps you with this actually validating that your bet was good and helping you change your direction if it wasn't good and so [Mårten Schultzberg]: What I'm trying to say is that, of course, sometimes, and maybe not even rarely, when we're building experimentation tooling, we have to build for quite some time before we can answer either of these questions.

[Mårten Schultzberg]: And it's hard to disentangle them even. [Mårten Schultzberg]: So I'd say that we build a completely new feature for experimentation, then some new methodology or something. [Mårten Schultzberg]: It's hard to even have [Mårten Schultzberg]: What's the dimension along here I can test if this is a lever worth pulling? [Mårten Schultzberg]: That's maybe a question more for market research or user research, all those kinds of things.

[Mårten Schultzberg]: Yeah, so that's the truth. [Mårten Schultzberg]: I think it's just a lot of, I think the teams that I'm writing this blog post for that I'm thinking about are the teams that sort of have a product already and they've been owning it for a while and they feel a bit stuck in terms of like they're not getting the sort of return of investment rate that they would like from their expectation. [Mårten Schultzberg]: They see that they have a lot of neutral results and they're wondering if they should run much longer experiments or what they should do about it.

[Mårten Schultzberg]: But yeah, I don't know. [Mårten Schultzberg]: Felt like partly cop out from your question there. [Val Kroll]: No, no, it's good. [Val Kroll]: It's, I mean, there's no clear question. [Tim Wilson]: Come on, Ben. [Tim Wilson]: I mean, it's, he basically said that it's like, it's like intuition with experimentation combines. [Tim Wilson]: It's kind of like you need to combine like the facts and the feelings. [Val Kroll]: I knew exactly where you were going with that when he said.

[Val Kroll]: Come together. [Tim Wilson]: Cheesy. [Tim Wilson]: So. [Val Kroll]: Okay, so. [Val Kroll]: Before I lose the thread because I last question, by the way, because we're, we're don't do that to me. [Val Kroll]: No, no, no, no. [Val Kroll]: I've got like three more, but I'll go fast. [Val Kroll]: I'll go, we'll go fire around. [Val Kroll]: Okay. [Val Kroll]: So you're talking about, um, uh, no one really likes the neutral results talking about some intuition with product.

[Val Kroll]: I'm going to talk about those outcomes. [Val Kroll]: So obviously if there's a win positive outcome, it ships. [Val Kroll]: If it hurt the experience, it doesn't ship. [Val Kroll]: If there was [Val Kroll]: an issue with the test set up, you hit an SRM or whatever, it doesn't ship neutral. [Val Kroll]: I want to talk about that. [Val Kroll]: Are there scenarios where the product intuition says, even though this was neutral, it makes sense for where the roadmap is going or some decisions we're making from branding, like maybe we're [Val Kroll]: This is building towards a bigger bet in the larger ecosystem to make things easier to share, more social.

[Val Kroll]: How do you think about the ship or no ship kind of action as it relates to those neutral results? [Mårten Schultzberg]: It's a great question. [Mårten Schultzberg]: My general recommendation there is that as long as you've decided before you run the experiment that you're going to ship if it's neutral, [Mårten Schultzberg]: I'm all good with it. [Mårten Schultzberg]: I think that there's a ton of situations where it makes sense to ship something if it didn't change anything, especially if you're building infrastructural type changes or if you're building towards something.

[Mårten Schultzberg]: We're building a lot of Spotify, building out AI features, as everyone else I suppose, but there's a lot of changes that we're making to our infrastructure just to be able to support features that we're planning to build. [Mårten Schultzberg]: And when we're making those changes, the idea is that we're hoping that nothing will change. [Mårten Schultzberg]: Maybe we're doing stuff to make things faster or something like that, but that's a bonus if it changes anything at any point.

[Mårten Schultzberg]: So there's a lot of changes that we are expecting won't make any difference. [Mårten Schultzberg]: So what we do then is that we essentially run what we call rollouts where we only have guardrail metrics, actually. [Mårten Schultzberg]: So we say, as long as [Mårten Schultzberg]: As long as we can prove that we didn't harm these metrics, we're going to ship it. [Mårten Schultzberg]: So then by using the rollout, you're sort of declaring your attempt from the beginning that like, hey, we're planning to ship this as long as it's not bad, which can sort of be a quite nice way to just make it explicit.

[Mårten Schultzberg]: That's completely fine. [Mårten Schultzberg]: But then again, I think that I just want to add a small caveat here that they also, I've heard a lot of product people at Spotify and other places talk about that like, [Mårten Schultzberg]: And this, even maybe if a metric doesn't look great or if it's neutral and stuff like that, there is this, I think, almost human fallacy to say, like, this is strategically imported, let's ship it anyway.

[Mårten Schultzberg]: And so I think it's, even though that's true, and I think that's why it's sort of an easy fallacy to fall into, or like it's an easy [Mårten Schultzberg]: trap. [Mårten Schultzberg]: That can be true, but I think everyone should think about how large proportions of the things we ship should be shipped from the argument this is strategically important. [Mårten Schultzberg]: Pretty small proportion is my general sense. [Val Kroll]: Everyone gets three here.

[Val Kroll]: Something like that. [Mårten Schultzberg]: I would love if I could give people a budget for those kinds of things. [Mårten Schultzberg]: I think it's all about trying to avoid the pitfalls of changing the objective when you see the results. [Mårten Schultzberg]: We do that all the time at Spotify. [Mårten Schultzberg]: We're shipping a ton of things that are neutral. [Mårten Schultzberg]: A lot of them are shipped with rollouts where we just explicitly say, we are planning to ship this thing for some reason.

[Mårten Schultzberg]: It might be business statistically or we have to improve our back end to scale for more traffic or whatever it might be. [Mårten Schultzberg]: We're going to ship it. [Mårten Schultzberg]: So we just want to know that we're not harming things. [Val Kroll]: I like that. [Val Kroll]: Okay, so Tim, I'm sorry. [Val Kroll]: I have to sneak in. [Val Kroll]: So what you're talking about here is a very nuanced [Val Kroll]: It feels like a nuanced analytical discussion.

[Val Kroll]: Should this be a rollout or how should this be exactly validated? [Val Kroll]: How do you think about the education? [Val Kroll]: Because you're not talking about an audience of 400 people who are deeply steeped in the analytics or the rationale for why you'd make some of those choices. [Val Kroll]: How do you think about the education piece to these different product teams? [Mårten Schultzberg]: Yeah, I mean, it's super important. [Mårten Schultzberg]: So I've spent, I wouldn't say majority, but a very big portion of my time at Spotify building educational material and mechanisms for this.

[Mårten Schultzberg]: I think that we have, I think, two strategies for this. [Mårten Schultzberg]: I think the first one is to keep the [Mårten Schultzberg]: the tool as simple as we possibly can, so have as few options as possible. [Mårten Schultzberg]: So we're talking about a lot of nuanced stuff here, but we also have removed a lot of stuff from our platform and simplified a lot of stuff and removed a lot of options, so made it quite opinionated.

[Mårten Schultzberg]: to minimize the things that people actually have to understand and know. [Mårten Schultzberg]: So that's one side. [Mårten Schultzberg]: On the other side is that we have very explicitly and deliberately built educational material and tooling for experimentation for many years. [Mårten Schultzberg]: So with confidence, we have this whole boot camp of self-serve courses. [Mårten Schultzberg]: We've also given a bunch of courses.

[Mårten Schultzberg]: We have something called Quick Starts, which is a very basic tutorial for like, this is how you run an experiment, this is how you run a rollout. [Mårten Schultzberg]: and those kinds of things. [Mårten Schultzberg]: I know it's a super important thing, but I think it has to come from two sides here. [Mårten Schultzberg]: You have to try to make the thing that people should learn as simple as possible because people don't have time.

[Mårten Schultzberg]: People have a lot of other things that they need to be good at and learn and understand, and then you have to create the material so that they can learn those things that they have to learn. [Mårten Schultzberg]: That's our solution to that. [Mårten Schultzberg]: I mean, we have thought a lot about that. [Mårten Schultzberg]: There's a lot of things that everyone that joins Spotify is onboarded to experimentation immediately, and they go through certain what's called golden paths at Spotify, which is like onboarding to certain things.

[Mårten Schultzberg]: And so if you're a mobile developer, then you learn how to work with our feature flags in mobile, and you run an AA test as part of your mobile engineer. [Mårten Schultzberg]: onboarding, for example. [Mårten Schultzberg]: So there's like, we have infiltrated the whole organization with experimentation onboarding and materials. [Mårten Schultzberg]: And that has helped. [Tim Wilson]: Wow. [Tim Wilson]: And Val, I'm going to have to put some duct tape.

[Tim Wilson]: I was like... And we're going to have to move to wrap. [Tim Wilson]: But I just have seven more. [Tim Wilson]: Val and the role of Moee Kiss on this episode. [Val Kroll]: Yeah, right? [Mårten Schultzberg]: No. [Mårten Schultzberg]: I have zero stress at least, so don't worry about me. [Tim Wilson]: Well, this great discussion, I love sort of the thinking about what are we doing, why are we doing it, and how can tooling and education and culture and framing all sort of come together.

[Tim Wilson]: So thanks for coming on for this discussion. [Tim Wilson]: But before we leave, the last thing we do on the show is go around the horn and we share a last call, something that might be of interest to our users. [Tim Wilson]: And Mårten, you're our guest. [Tim Wilson]: Do you have a last call you'd like to share? [Mårten Schultzberg]: Yes. [Mårten Schultzberg]: So one thing that I'm completely, like I have been for actually for many years, but now renewed is the YouTube channel Three Blue One Brown.

[Mårten Schultzberg]: If I'm not the first one, that just makes me happy because it's the best. [Mårten Schultzberg]: The thing that I'm particularly thinking about now is the videos on Transformers and LLMs. [Mårten Schultzberg]: This YouTube channel is essentially a channel that visualizes a bunch of math. [Mårten Schultzberg]: That sounds maybe not fun, but it is so insanely good. [Mårten Schultzberg]: They have a long series on linear algebra that I think if I would have actually seen it when I was taking linear algebra, it would have helped me a lot.

[Mårten Schultzberg]: But they also have a bunch of super, super nice things on LLMs and Transformers, which I think is like [Mårten Schultzberg]: If you are, like most people, like hearing that word many times and you have like, yeah, it's some kind of neural net. [Mårten Schultzberg]: Maybe I haven't used a neural net once or twice, but like you have no idea really how it works. [Mårten Schultzberg]: Those videos are so very, very good. [Mårten Schultzberg]: So I recommend them highly.

[Tim Wilson]: We have reached out to have, we had an exchange trying to get him to come on the show. [Tim Wilson]: I think it might have been around to talk about neural networks. [Tim Wilson]: He was in the process of like moving. [Tim Wilson]: So he's on our list to try to get him on. [Tim Wilson]: That's a- Good reminder. [Tim Wilson]: That's a good one, good reminder to go back because they are, they're like, I've sampled some of those and I'm like, this is so clear.

[Tim Wilson]: And how does a human being have the time to produce something like this? [Mårten Schultzberg]: Yeah, I mean, Grant Sanderson who has that channel. [Mårten Schultzberg]: I mean, he seems to be like one of the true geniuses alive. [Mårten Schultzberg]: Like, I mean, just a side note here is like he's doing this super nice like animations of math and you just built that library himself, the library he built. [Mårten Schultzberg]: It's just...

[Tim Wilson]: Come on. [Tim Wilson]: We're going to use this call out when this comes out to reach out to him again and say, hey, come chat with us. [Mårten Schultzberg]: I would listen 100%. [Tim Wilson]: Awesome. [Tim Wilson]: Val, what about you? [Tim Wilson]: Do you have a last call? [Val Kroll]: I do. [Val Kroll]: And it's actually related to today's episode. [Val Kroll]: So this is a medium article published on Unix Collective, article written by James Skinner.

[Val Kroll]: It's called Escaping the AI Sludge Why MVPs Should Be Delightful. [Val Kroll]: And there's a lot in here, one of the cases he makes is that like using AI is just like regurgitating like we're not going to get to that delight level if we're just, you know, using AI to help, you know, develop those different. [Val Kroll]: net new versions that are being tested within a product context. [Val Kroll]: But he talks about the MLPs. [Val Kroll]: I'm obsessed with MVP's, Mårten, I should tell you, just understanding different people's perspective.

[Val Kroll]: But the MLP is the minimum levelable product. [Val Kroll]: And he also referenced one, the minimum viable whatever, because there's so many acronyms related to this, with people trying to figure out exactly what that level of fidelity should be, what type of investment you should make before you experiment. [Val Kroll]: He does talk about experimentation at the end, which I do love, but there's a lot of really good examples. [Val Kroll]: And I love reading from that design product perspective.

[Val Kroll]: So, but it's a, it's a good read, about 10 minute read. [Val Kroll]: So it's a good one. [Val Kroll]: And Tim, how about you? [Val Kroll]: Do you have a last call for today? [Tim Wilson]: I've got a smidge of housekeeping and a last call. [Tim Wilson]: So we are now like into month number two of 2026, which means we're heading into a conference season. [Tim Wilson]: Actually, I am, [Tim Wilson]: sitting in Budapest, Hungary as you were listening to this, if you're listening to it when it came out.

[Tim Wilson]: A couple of analytics power hour conference attendee appearances coming up in Nashville. [Tim Wilson]: If you're in the States, there's the Datatune conference that Val and I will both be attending on March 6th and 7th. [Tim Wilson]: Some critical mass of the [Tim Wilson]: Analytics Power Hour crew, we will be recording a show with a live audience at the Marketing Analytics Summit in Santa Barbara, California on April 28th and 29th.

[Tim Wilson]: Those are PSAs more than last calls. [Tim Wilson]: My last call would be friend of the show, past guest, Katie Bauer, the wrong but useful sub-stack wrote a post called The Next Data Bottle Neck, which [Tim Wilson]: I thought it was a unique and really thought-provoking take on the whole drive towards conversational analytics and not the will it or won't it or the technical challenges of it, but when looking at what people are asking for and why they actually seem to be mundane [Tim Wilson]: requests that they seem to be kind of just simple data fetching requests, not these super nuanced things.

[Tim Wilson]: So she has a lot of musings that can be a little unsettling for the analyst, but then she actually kind of wraps by making the case that really it goes back to good analysts really thinking about the business deeply. [Tim Wilson]: So it's a worthwhile read. [Tim Wilson]: So I was a threefer, but I've labeled two of them as being a housekeeping writer in the last class. [Val Kroll]: Can I ask one more question then? [Mårten Schultzberg]: That's how you get airtime in this show, right?

[Tim Wilson]: I'm drunk on power. [Tim Wilson]: Is Michael as drunk on Tamiflu? [Tim Wilson]: Tamiflu? [Tim Wilson]: I don't know what the flu medications are. [Tim Wilson]: Yeah. [Tim Wilson]: By the time this comes out, he will be back to good health and he will vow to never get sick again and cede the mic to me. [Tim Wilson]: So this was great. [Tim Wilson]: Thanks again, Mårten, for coming on. [Tim Wilson]: This was a really fun discussion. [Mårten Schultzberg]: My pleasure.

[Mårten Schultzberg]: It's really nice. [Mårten Schultzberg]: Thank you so much for having me. [Tim Wilson]: Awesome. [Tim Wilson]: Everybody get your Spotify subscription up to speed. [Tim Wilson]: This is what's driving Spotify's next round of growth is the confidence podcast appearance. [Mårten Schultzberg]: Quarterly call coming, so like, please. [Val Kroll]: There you go. [Tim Wilson]: Perfect. [Tim Wilson]: If you are listening and you've enjoyed this show or other shows, we would always love a rating and review.

[Tim Wilson]: I'll do a little call on audible and read out this one from Apple Podcast that just T5272018 left. [Tim Wilson]: It was titled Smart and Funny. [Tim Wilson]: And it was love the insights and laughs I get from this podcast. [Tim Wilson]: You all have a high bar for analysts and the value they can add, which I so appreciate. [Tim Wilson]: And you share all of that perspective via hilarious and authentic banter. [Tim Wilson]: Keep it up.

[Tim Wilson]: Wait, let me check. [Tim Wilson]: That is our podcast. [Tim Wilson]: Yeah, that is this one. [Tim Wilson]: So that was kind of nice. [Tim Wilson]: We'll always love to get ratings and reviews. [Tim Wilson]: Theoretically, that is how we expand the reach of the show, that and recording video and putting them on YouTube. [Tim Wilson]: So we'll just double down on the ratings and reviews. [Tim Wilson]: If you're a fan of the show and would like to have a sticker for your laptop or water bottle or whatever, you can go to analyticshour.io and request a sticker.

[Tim Wilson]: We'll ship one over. [Tim Wilson]: If you have something to say, a thought for a topic, criticism, your own little witticism that you'd like to share, you can reach out to any of us or the show as a whole on LinkedIn. [Tim Wilson]: You can catch us on the measure slack or you can just send an email to contact at analyticshour.io. [Tim Wilson]: So, with that, for Val and for Michael in absentia from his sickbed, I'm Tim Wilson and no matter what your reason, whether you're identifying or you're optimizing or you're being just aggressively neutral in your findings, you should always keep analyzing.

[Announcer]: Thanks for listening. [Announcer]: Let's keep the conversation going with your comments, suggestions, and questions on Twitter at @analyticshour on the web at analyticshour.io, our LinkedIn group, and the Measure Chat Slack group. [Announcer]: Music for the podcast by Josh Crowhurst. [Announcer]: Those smart guys wanted to fit in, so they made up a term called analytics. [Announcer]: Analytics don't work. [Charles Barkley]: Do the analytics say go for it, no matter who's going for it?

[Charles Barkley]: So if you and I were on the field, the analytics say go for it. [Charles Barkley]: It's the stupidest, laziest, lamest thing I've ever heard for reasoning in competition. [Tim Wilson]: Yeah, we've sent, Australia is the one that's the real Australia. [Val Kroll]: Singapore. [Tim Wilson]: We'll take weeks. [Tim Wilson]: Singapore, one made it all the way to Singapore, came back to Ohio. [Tim Wilson]: Never came to me, turned around and went back to Singapore.

[Tim Wilson]: So it was like eight weeks. [Val Kroll]: The box was like smashed. [Val Kroll]: The gift wasn't ruined, but the box was in shambles. [Tim Wilson]: There is now more packing material. [Tim Wilson]: I did change after seeing that. [Tim Wilson]: It's a process update. [Mårten Schultzberg]: I guess I should save all of my comments about it for the actual recording. [Val Kroll]: Yeah, we'll get into it for sure. [Val Kroll]: I'm very excited.

[Mårten Schultzberg]: It wasn't that terrible. [Mårten Schultzberg]: The distortion wasn't that terrible. [Val Kroll]: So every time you do that while we actually record, because you'll definitely be doing that multiple times, I'm just kidding. [Mårten Schultzberg]: Yeah. [Val Kroll]: Yeah, it looks like. [Val Kroll]: Last for me. Okay. [Mårten Schultzberg]: Part of your signal yelling at you. [Val Kroll]: Your guests. [Tim Wilson]: All right, let's try it again.

[Val Kroll]: Rock flag and focus on those learnings.

Transcript supplied by the publisher with the episode.

The Analytics Power Hour

by Michael Helbling, Moe Kiss, Tim Wilson, Val Kroll, and Julie Hoyer · English · Business

Attend any conference for any topic and you will hear people saying after that the best and most informative discussions happened in the bar after the show. Ready any business magazine and you will find an article saying something along the lines of "Business Analytics is the hottest job category o

More from The Analytics Power Hour

  1. 17 Mar 2026 · 1 hr 6 min

    #293: Tool Selection and the Unhelpfulness of Feature Comparisons

    The one rule about the Analytics Power Hour is that we don't talk about specific tools. But that doesn't mean we won't talk about tool SELECTION! Jason Packer recently released the second edition of Google Analytics Alternatives , (also available on Amazon ) and his approach in the book is very much not an RFP-like "check which features your tool offers" system. And his rationale for that seems just as applicable (to us, at least!) for any data platform selection, be it a digital/product analytics platform, a BI tool, database or storage infrastructure, or, well, you name it! Ultimately, the…

  2. 3 Mar 2026 · 1 hr 4 min

    #292: AI Without Adult Supervision with Aubrey Blanche

    As Kevin McCallister once taught us: just because the house is still standing doesn't mean everything's under control. Everyone's racing to adopt AI, but has anyone actually read the fine print? For this year's International Women's Day episode, we are joined by Aubrey Blanche to unpack the hype, the hidden tradeoffs, and the quiet ways teams are giving up agency in the name of "productivity." We explore how data and tech teams are uniquely prepared and positioned to ask better questions, measure what really matters, and avoid letting the AI teenager run the house. Learn more about "phantom…

  3. 17 Feb 2026 · 1 hr 3 min

    #291: The Data Work that Lives in the Shadows

    We know what the work of the data practitioner is, right? It's everything from managing data ingestion to data governance to report development to experimental design to basic and advanced analytics. It's writing (or vibe-writing?) SQL or Python or R while also being adept at whatever data stack—no matter how modern—is at hand. Of course, it's a lot more, too! And that's the topic of this episode: the unofficial, often unheralded, but often quite important "shadow work" of the analyst—the myriad tasks required to effectively glue together all the data work that occurs out in broad daylight…

  4. 20 Jan 2026 · 1 hr 10 min

    #289: The Imperative of Developing Business Acumen

    That darn data. It's so complicated and fragmented and gap-filled and noisy that no amount of time is ever enough to truly get to the bottom of all of its complexity. As a result, it's pretty easy to fill all of our time handling as much of that underlying data messiness as possible. At what cost, though? It's easy for the analyst's connection to the business to suffer as they get mired (too) deeply in the data and lose sight of the broader business needs. In this episode, the gang had a chat about business acumen—what it is, how to develop it, and why it's a must-have for any data or…

  5. 6 Jan 2026 · 1 hr 1 min

    #288: Our LLM Suggested We Chat about MCP. Kinda' Meta, No?

    If there's one thing that we absolutely knew would be coming along with the increased interest and use of AI, it would be… more acronyms! And, along with the acronyms, we pretty much could predict that we see a lot of online flexing through casual dropping of said acronyms as though they're deeply understood by everyone who's anyone. We tackled one such acronym on this episode: MCP! That's "model context protocol" for those who like their acronyms written out, and Sam Redfern joined us to help us wrap our heads around the topic. You see, MCP is kinda' like some other more familiar acronyms…

  6. 23 Dec 2025 · 1 hr 1 min

    #287: 2025 Year in Review

    It's the most…won…derful…tiiiiime…of the year! And by that, we mean it's the time of the year when we sit back, look at each other, and ask, "Where did all the time go?!" We brought back a very special someone for this episode as we collectively reflected on the year—show highlights (and what about those shows have stuck with us), industry reflections, and a little shameless shilling for Tim's book (are you still short on a few stocking stuffers? Order now…!). This episode's Measurement Bite from show sponsor Recast is a brief explanation of Granger causality (and how it's NOT actually a…

  7. E307 · 29 Sep 2026 · 47 min

    #307: AI Does (and Does Not) Work for Data Storytelling

    There's a version of this episode where AI replaces the data visualization expert entirely, and there's a version where it's useless without one. Neither is quite right, according to Cole Nussbaumer Knaflic of Storytelling with Data , who joined Tim, Moe, and Julie to sort out where AI genuinely earns its keep in data storytelling and where it's just producing a shinier first draft of the same shitty slide. Cole-admittedly a skeptic-turned-convert who once joked she'd retire before having to deal with any of this-walks through why the humans who benefit most from AI are the ones who already…

  8. E306 · 15 Sep 2026 · 46 min

    #306: Decision Support Has Been the Point All Along

    Before there was BI, there were "decision support systems." Somewhere along the way, we seem to have quietly dropped the "decision support" part and just kept the systems. Zohar Strinka, founder of Analytics Strategies and creator of the Meta-Problem Method, joined us to put the point back where it belongs: if nothing is going to be done differently after the analysis, then the analysis had exactly zero effect on the world. But -- and this is the part that's easy to miss -- that does NOT mean marching up to a stakeholder and demanding, "What decision are you going to make?" People don't want…

  9. E305 · 1 Sep 2026 · 1 hr 4 min

    #305: Personal Interest + Analytics Chops = Career?

    Stephen Follows has spent 15 years turning film-industry curiosity into a career — asking questions like how much movies really earn, why poster colors have shifted over decades, and whether old studio domains are still up for grabs. In this episode, he joined us to talk about building that unusual path, the cautionary tale of TheNumbers.com, and his theory about Sandra Bullock. Go start your passion project. This episode is brought to you, in part, by our sponsors, Prism from Ask-Y and Stape . For complete show notes, including links to items mentioned in this episode and a transcript of…

  10. 18 Aug 2026 · 1 hr 8 min

    #304: I Can Haz AI?

    Everyone's posting their AI projects online like proud pet owners sharing videos of their cat doing something marginally impressive — cute, occasionally clever, sometimes a little cringe. And the Analytics Power Hour is no different! In this co-hosts-only episode, Tim, Michael, and Julie skip the thought leadership hot takes and just… compare notes. What have they actually built? What broke? What surprised them? From a custom GPT podcast librarian to a full-blown show production app wired up to Neon, Vercel, Resend, and about five other things Michael is only sort of sure he set up…

Every episode of The Analytics Power Hour →

Take it with you

The Melo app keeps playing with the screen off, works in the car and on your watch, wakes you to your station, and browses the whole catalogue offline. Free, no ads, no account.

Get it on Google Play