YouSaid · the spoken record

Ronny Kohavi

lines on the record
99
first
2023-07-27
most recent
2023-07-27
sittings or episodes
1
sources
podcast

Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections

  1. So, I think what you said is fair. I mean, I do want to allocate some percentage of resources to big vets. As you said, we've been optimizing this thing to hell. Could we completely redesign it? It's a very valid idea. You may be able to break out of a local minima. What I'm telling you is 80% of the time you will fail. So be ready for that, right? What people usually expect is my redesign is going to work. No. You're most likely going to fail, but if you do succeed, it's a breakthrough.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  2. I teach this in my class, but I think I've posted this on LinkedIn and answered some questions. I'm happy to put that in the notes.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  3. Right? There was this complete denial that it's possible that 50% of ideas that Microsoft is implementing in a three-year development cycle, by the way, this is how long it took Office to release. It was a classical every three years we release. And the data came about showing that, yeah, Bing was the first to truly implement experimentation at scale. And we shared with the rest of the companies the surprising results. And so when Office... And this was credit to Chilu and Sachya Nadella. They were ones that says, Ronnie, you try to get office-to-run experiments. We'll give you the air support. And it was hard, but we did it. It took a while, but office started to run experiments and they realized that many of their ideas are failing.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  4. But I think organizations that start to run experiments are humbled early on from the smaller changes. You're right. Nobody, I'll tell you a funny story. When I came from Amazon to Microsoft, I joined the group, and for one reason or another, that group disbanded a month after I joined. And so people came to me and said, look, you just joined the company. You're at partner level. You figure out how you can help Microsoft. And I said, I'm going to build an experimentation platform because nobody at Microsoft is running experiments. And 50% more than 50% of ideas at Amazon that we tried failed. The classical response was, we have better PMs here

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  5. And by the way, I'm not opposed to large redesigns. I try to give the team the data to say, look, here are lots of examples where big redesigns fail. Try to decompose your redesign if you can't decompose it to one factor at a time to a small set of factors at a time. And learn from these smaller changes what works and what doesn't. Now, it's also possible to do a complete redesign. I'm just, as you said yourself, they'd be ready to fail. Right? I mean, do you really want to work on something for six months or a year and then run the A-B test and realize that you've hurt revenues or other key metrics by several percentage points and a data-driven organization will not allow you to launch? What are you going to write in your annual review?

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  6. I mean, this is the sun cost fallacy, right? We invested so many years in it. Let's launch this, even though it's bad for the user. No, that's terrible. Yeah, yeah. So this is the other advantage of recognizing this humble reality that most ideas fail. If you believe Do them in smaller increments, learn from OFAT one factor at a time. Do one factor learn from it and adjust of the 17, maybe you have four good ideas. Those are the ones that will launch and be positive.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  7. Absolutely. Yeah. In fact, I've published some of these in LinkedIn posts showing a large set of big launches that redesigns that dramatically failed. And it happens very often. So the right way to do this is to say, yes, we want to do a redesign, but let's do it in steps and test on the way and adjust. So you don't need to take 17 new changes. That many of them are going to fail, start to move incrementally in a direction that you believe is beneficial. Just on the way

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  8. We all love them I mean, this is the humbling reality, and people talk about the fact that A-B testing sometimes leads you to incremental. I actually think that many of these small insights lead to fundamental insights about which areas to go, some strategies we should take, some things we should develop helps a lot.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  9. Lifetime value through email when they unsubscribe, let's offer them by default to unsubscribe from this campaign. So when you get an email, you know, there's a new book by the author, the default to unsubscribe would be unsubscribe me from author emails. And so now the negative, the countervailing metric is much smaller. And so again, this was a breakthrough in our ability to send more emails and understand based on what users were unsubscribing from, which ones are really beneficial.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  10. So we did some data science study on the side, and we said, What is the value that we're losing from an unsubscribe? And we came up with a number of a few dollars. But the point was now we have this countervailing metric. We say, here's the. Money that we generate from the emails. Here's the money that we're losing on long-term value. What's the trade-off? And then when we started to incorporate those formulas, more than half the campaigns that were being sent were negative So it was a huge insight at Amazon about how to send the right campaigns. And this led, and this is what I like about these discoveries, this fact that we integrated the unsubscribed led us to a new feature. Say, well, let's not lose their future.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  11. How do we give credit to that team? And the initial version was well, whenever a user comes from the email and purchases something on Amazon, we're going to give that email credit. Well, it turned out this had no countervailing metric. The more emails you send, the more money you're going to credit the team. And so that led to spam. Literally a really interesting problem. The team just ramped up the number of emails that they were sending out and claimed to make more money and their fitness function improved. And then so then we backed up and then we said, okay, we can either phrase this as a constraint satisfaction problem. You're allowed to send user an email every X days. War, which is what we ended up doing, is let's model the cost of spamming the users. What's that cost? Well, when they unsubscribe, we can't mail them.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  12. So there's two ways that I like to think about it. One is you can run long term experiments for the goal of learning something. So I mentioned that at Bing. We did run these experiments where we increased the ads and decreased the ads so that we will understand what happens to key metrics. The other thing is, you can just build models that use some of our background knowledge or use some data science to look at historical. I'll give you another good example of this. When I came to Amazon, one of the teams that reported to me was the email team that it was not the transactional emails when you buy something, you get an email, but it was the team that sent these recommendations. Here's a book by an author that you bought. Here's a product that we recommend. And the question is

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  13. To me, the key here, the key word is lifetime value. Is you have to define the OEC such that it is causally predictive of the lifetime value of the user. And that's what causes you to think about things properly, which is am I doing something that just helps me short term or am I doing something that will help me in the long term? Once you put that model of lifetime value, people say, okay, what about retention rates? We can measure that. What about the time to achieve a task? We can measure that. And those are these countervailing metrics that make it make the OEC useful.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  14. You can say I want to improve conversion rate, but you can be smarter about it and say it's not just enough to convert a user to buy or to pay for a listing. I want them to be happy with it several months down the road when they actually stay there. So that could be part of your OEC to say what is the rating that they will give to that. Listing when they actually stay there. And that causes an interesting problem because you don't have this data now. You're going to have it three months from now when they actually stay. So you have to build the training set that allows you to make a prediction about whether this user, whether Lenny, is going to be happy at this cheap place or whether, no, I should offer him something more expensive because Lenny likes to stay at nicer places where the water actually is hot and comes out of the faucet.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  15. Then you need to insert these other criteria. And what am I doing to the user experience one way around it is to put this constraint? Another one is just to have these other metrics, again, something that we did to look at the user experience. How long does it take the user to reach a successful click? What percentage of sessions are successful? These are key metrics that were part of the overall evaluation criterion that we've used. I can give you another example, by the way, from the hotel industry or Airbnb that we both worked at.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  16. Can have wider, bigger ads. I'm just going to count the pixels that you take, the vertical pixels. And I will give you some budget. And if you can under the same budget, make more money, you're good to go. So that to me turns the problem from a badly defined, let's just make more money, right? Any page can start plastering more ads and make more money short term, but that's not the goal. The goal is long term.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  17. The question is what happens to the user experience and how is that going to impact you in the long term? So we've run those experiments and we were able to map out this number of ads causes this much increase to churn. This number of ads causes this much increase to the time that users take to find a successful result. And we came up with an OEC that is based on these metrics that allows you to say, okay, I'm willing to take this additional money if I'm not hurting the user experience by more than this much, right? So there's a trade-off there. One of the nice ways to phrase this as a constrained optimization problem, I want you to increase revenue, but I'm going to give you a fixed amount of average real estate that you can use. So you can, for one query, you can have zero ads. For another query, you can have three ads. For a third query.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  18. The OEC, or the overall evaluation criterion, is something that I think many people that start to dabble in AB testing miss. And the question is, what are you optimizing for? And it's a much harder question than people think because it's very easy to say we're going to optimize for money, revenue. That's the wrong question because you can do a lot of bad things that will improve revenue. So there has to be some countervailing metric that tells you how do I improve revenue without hurting the user experience. Okay, so let's take a good example with search. You can put more ads on the page and you will make more money. There's no doubt about it. You will make more money in the short term.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  19. Math, the statistics just don't work out for most of the metrics that you're interested in. In fact, I gave an actual practical number of a retail site with some conversion rate trying to detect changes that are at least 5% beneficial, which is something that startups should focus on. They shouldn't focus on the 1%. They should focus on the 5% and 10%. Then you need something like 200,000 users. So start experimenting when you're in the tens of thousands of users. You'll be only be able to detect large effects. And then once you get to 200,000 users, then the magic starts happening. Then you can start testing a lot more. Then you have the ability to test everything and make sure that you're not degrading and getting value address fermentation. So you ask for rule of thumb. 200,000 users, you're magical. Below that, start building the culture, start building the platform, start integrating.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  20. A million dollar question that everybody asked. So I actually will put this in the notes, but I gave a talk last year, what I called it is practical defaults. And one of the things I show there is that unless you have at least tens of thousands of users,

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  21. After a while, the cost of running experiments was so low that nobody was questioning the idea that everything should be experimented with. Now, I don't think we were there at Airbnb, for example. The platform at Airbnb was much less mature and required a lot more analysts in order to interpret the results and to find issues with it. So I do think there's this trade-off. You're willing to invest in a platform. It is possible to get the marginal cost to be close to zero. But when you're not there, it's still expensive and there are may be reasons why not to run ABTs.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  22. First of all, there are some necessary ingredients to A-B testing, and I'll just say outright not every domain is amenable to AB testing. You can't A-B test mergers and acquisitions, right? It's something that happens once. You either acquire or you don't acquire. So, you do have to have some necessary ingredient. You need to have enough units, mostly users, in order for the statistics to work out. So yeah, if you're too small. Maybe too early to AB test. But what I find is that in software, it is so easy to run A-B testing and it is so easy to build a platform. I don't say it's easy to build a platform, but once you build a platform, the incremental cost of running an experiment should approach zero. And we got to that at Microsoft where

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  23. Yeah, this is hard. This is hard, but that's again, that's the value of. Experiments, which are this oracle that gives you the data, you may be excited about things, you may believe it's a good idea, but ultimately the arbiter, the oracle is the controlled experiment. It tells you whether users are actually benefiting from it, whether you and the users, the company, and the users.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  24. You don't see it anymore. It existed for about a year and a half, and all the experiments were just negative to flat. And it was an attempt. It was fair to try it. I think it took us a little long to fail to decide that this is a failure. But at least we had the data. We had hundreds of experiments that we tried. None of them were a breakthrough. And I remember sort of mailing Chilu with some statistics showing that, you know, it's time to abort, it's time to fail on this. And, you know, he decided to continue more. And it's a million dollar question. Do you continue? And then maybe the breakthrough will come next month or do you abort? And a few months later, we abort it.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  25. Time to say, hey, if you go for something big, try it out, but be ready to fail 80% of the time. One true example that, again, I'm able to talk about because we put it in my book is we were at Bing. Trying to change the landscape of search. And one of the ideas, the big ideas, was we're going. So we hooked into the Twitter fire hose feed, and we hooked into Facebook. And we spent 100 person years. On this idea And it failed

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  26. Because some people say, well, if we only tested 17 things around this and you have to think about it's not just, it's like in stock, you need a portfolio, you need some experiments that are incremental that move you in the direction that you know you're going to be successful over time if you just try enough. But some experiments you have to allocate sometimes to these high risk, high reward ideas. We're going to try something that's most likely to fail. But if it does win, it's going to be a home run. And so you have to allocate some efforts to that. And you have to be ready to understand and agree that most will fail. Most of these high, and it's amazing how many times I've seen people come up with new designs or a radical new idea and they believe in it and that's okay. I'm just cautioning them all the time.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  27. Very clear that I'm a big fan of test everything, which is any code change that you make. Any feature that you introduce has to be in some experiment because, again, I've observed this sort of. Surprising result that even small bug fixes, even small changes can sometimes have surprising unexpected impact. And so I don't think it's possible to experiment too much. I think it is possible to focus on incremental changes.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  28. Big numbers. So when you're running in a group like Bing, which is running thousands and thousands of experiments, you want to be able to ask, has anybody did an experiment on this or this or this? And so that searching capability is in the platform. But more than that, I think just doing the quarterly meeting of the most successful, most interesting, sorry, not just successful, most interesting experiments is very key. And that also helps the flywheel of experimentation.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  29. Document that right, we had a large deck internally of these successes and failures. And we encourage people to look at them. The other thing that's very beneficial is just to have your whole history of experiments and do some ability to search by keywords, right? So I have an idea. Type a few keywords and see if the Thousands of experiments that ran. And by the way, these are very reasonable numbers at Microsoft, just to let you know, when I left in 2019, we were on a rate of about 20 to 25,000 experiments every year. So every working day, we were starting something like 100 new treatments.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  30. It's that negative that gives you insight. I'm just coming up with one example of that that I should mention. We were running this experiment at Microsoft to improve the Windows indexer. And the team was able to show on offline test that it does much better at indexing and they showed some relevance is higher and all these good things. And then they ran it as an experiment. And you know what happened? Surprising result. Indexing the relevance was actually high, but it killed a battery life. So, here's something that comes from left field that you didn't expect, it was consuming a lot more CPU on laptops, it was killing the laptops, and therefore, okay, we learned something. Let's document that. Let's remember this so that, you know, we now take this other factor into account as we design the next iteration.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  31. And the actual result differs by a lot. So that absolute value of the difference is large. Now you can expect something to be great and it's flat. Well, you learn something. But if you expect something to be small and it turns out to be great, like that ad title promotion, then you've learned a lot. Or conversely, if you expect that something will be small and it's very negative, you can learn a lot by understanding why this was so negative. And that's interesting. So we focus not just on the winners, but also surprising losers, things that people thought would be a no-brainer to run. And then for some reason, it was very negative. And sometimes...

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  32. In general, by the way, this is all about a lot of that is institutional memory, right? Which is can you document things well enough so that the organization remembers the successes and failures and learns from them. I think one of the mistakes that some company makes is they launch a lot of experiments and never go back and summarize the learnings. So I've actually put a lot of effort in this idea of institutional learning, of doing the quarterly meeting of the most surprising experiments. By the way, surprising is another question that people often are not clear about. What is a surprising experiment? To me, a surprising experiment is one where the estimated resolve beforehand

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  33. And goodui.org is exactly the site that tries to do what you're saying at scale. So guy's name is Jacob Blanovsky. He asks people to send them results of experiments and he derives, he puts them into patterns. There's probably like 140 patterns, I think, at this point. And then for each pattern, he says, well, who has that helped? How many times and by how much? So you have an idea of, you know, this worked three out of five times and it was a huge win. In fact, you can find that opening your window in there

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  34. I give you two resources. One of them is a paper that we wrote. Called Rules of Thumb. And what we tried to do at that time at Microsoft was to just look at thousands of experiments that run and extract some patterns. And so that's one paper that we can then put in the notes But there's another more accurate, I would say, resource that's useful that I recommend to people. And it's a site called GoodUI.org.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  35. Always start off by thinking that somehow they're different and their successor is going to be much, much higher, and they're all humbled.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  36. It is important to realize that when you have a platform, it's easy to get this number. You look at how many experiments were run and how many of them launched, not every experiment maps to an idea. So it's possible that when you have an idea, your first implementation, you start an experiment, boom, it's egregiously bad because you have a bug. In fact, 10% of experiments tend to be aborted on the first date. Those are usually not that the idea is bad, but that there is an implementation issue or something we haven't thought about that forces on a board. You may iterate and pivot again. And ultimately, if you do two or three or four pivots or bug fixes, you may get to a successful launch. But those numbers of 80 to 92% failure rate are of experiments. Very humbling. I know that every group that starts to run experiments,

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  37. Well, first of all, I published three different numbers for my career. So overall at Microsoft, about 66%, two-thirds of ideas fail, right? And don't take the 66 as accurate. Like, you know, it's about two-thirds. And Bing, which is a much more optimized domain after we've been optimizing it for a while, the failure rate was around 85%. So it's harder to improve something that you've been optimizing for a while. And then at Airbnb, this 92% number is the highest failure rate that I've observed. Now, I've quoted other sources that, you know, it's not that I worked at groups that were particularly bad, booking Google ads, other companies published numbers that are around 80 to 90% failure rate of ideas. This is where it's important.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  38. You have a small gain. In fact, again, there's another number I'm allowed to say of these experiments, 92% failed to improve the metric that we were trying to move. So only 8% of our ideas actually were successful at moving the key metrics.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  39. Hundreds of people all working to improve being relevant, they have a metric, we'll talk about OECD, the overall evaluation criterion, but they have a metric that their goal is to improve it by 2% every year. It's a small amount. And that 2% you can see, here's 0.1, here's a 0.15, here's a 0.2, and then they add up to around 2% every year, which is amazing. Another example that I am allowed to speak about from Airbnb is the fact that we ran some 250 experiments in my tenure there in search relevance. And again, small improvements added up. So this became overall a 6% improvement to revenue. So when you think about 6%, it's a big number, but it came out not of one idea, but many, many smaller ideas that each gave you.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  40. Legal reasons or other things we you know there were some concern that we were not marking the ads properly so you have to suddenly do something that you know is going to hurt revenue but yes i think most results are inch by inch you improve small amounts lots of them i think the best example that i can say is a couple of them that i can speak about one is at bing the relevance team

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  41. Yeah, so again, this is a topic that's near and dear to my heart. Everybody wants these amazing results. And, you know, I show them in chapter one in my book multiple of these small effort, huge gain. But as you said, they're very rare. I think most of the time the winnings are made sort of this inch by inch. And there's a graph that I show in my book, a real graph of how Bing Ads has managed to improve the revenue per thousand searches over time. And every month you can see a small improvement and a small improvement, sometimes a degradation because of

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  42. That we learn so much from it. So we first did this. It was done in the UK for opening hotmail. And then we moved it to MSN so it would open search in Utah. And all the set of experiments were highly, highly beneficial. We published this. And I have to tell you, when I came to Airbnb, I talked to our joint friend Ricardo about this. And it was sort of darn. It was very beneficial. And then it was semi-forgotten, which is one of the things you learned about institutional memories. When you have winners, make sure to address them and remember them. So it was an Airbnb done for a long time before I joined that listings open in a new tab. But other things that were designed in the future were not done. And I reintroduced this to the team. And we saw big improvements.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  43. Yeah, and by the way, I don't know if you know the history of this, but I tell about this in class. We did this experiment. Way back around 2008, I think. And so this predates Airbnb. And I remember it was heavily debated why would you open something in Utah? The users didn't ask for it. It was a lot of pushback from the designers. And we ran that experiment. And again, it was one of these highly surprising results that made it.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  44. And this was just moving the second line to the first line. Now, you then go and run a lot of experiments to understand what happened here. Is it the fact that the title line has a bigger font, sometimes different color? So we ran a whole bunch of experiments. And this is what usually happens. We have a breakthrough. You start to understand more about what can we do. And there was suddenly a shift towards, okay, what are other things we could do that would allow us to improve revenue? We came up with a lot of follow-on ideas that helped a lot. But to me, this was an example of a tiny change that was the best revenue generating idea in Bing's history. And we didn't rate it properly, right? Nobody gave this the priority that in hindsight it deserves. And that's something that happens often. I mean, we are often humbled by how bad.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  45. And we did, and we looked for several times and we replicated the experiment several times, and there was nothing wrong with it. This thing was worth $100 million at the time when Bing was a lot smaller. And the key thing is it didn't hurt the user metric. So it's very easy to increase revenue by doing theatrics that, you know, displaying more ads is a trivial way to raise revenue. But it hurts the user experience. And we've done the experiments to show that. In this case, This was just a home run that improved revenue didn't significantly hurt the guardrail metrics. And so we were like in awe of, you know, what a trivial change that was the biggest revenue impact to Bing in all its history.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  46. The common thing at Bing, he launched the experiment. And a funny thing happened. We had an alarm, big escalation. Something is wrong with the revenue metric Now, this alarm fired several times in the past when there were real mistakes where somebody would log revenue twice or there's some data problem. But in this case, there was no bug. That simple idea increased revenue by about 12%. And this is something that just doesn't happen. We can talk later about Tuyman's law, but that was the first reaction, which is this is too good to be true. Let's find a bug.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  47. And when you think about that, and there's, you know, if you're going to look in my book or in the class, there's an actual diagram of what happened, the screenshots. But if you think about it, just realistically, it looks like a meh idea. Like, why would this be such a reasonable, interesting thing to do? And indeed, when we went back to the backlog, it was on the backlog for months and languished there. many things were rated higher but The point about this is it's trivial to implement. So if you think about return on investment, we could get the data by having some engineer spend a couple of hours implementing it. And that's exactly what happened. Somebody, Ed Bing, who kept seeing this in the backlog and said, my God, we're spending too much time discussing it. I could just implement it. He did. He spent a couple of days implementing it as it is.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  48. Yeah, so I think the opening example that I use in my book and in my class is the most surprising public example we can talk about. And this is kind of an interesting experiment. Somebody proposed to change the way that ads were displayed on being the search engine. And he basically said, let's take the second line and move it, promote it to the first line so that the title line becomes larger.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source

  49. I'm very clear that I'm a big fan of test everything, which is any code change that you make, any feature that you introduce has to be in some experiment because, again, I've observed this sort of Surprising result that even small bug fixes, even small changes can sometimes have surprising unexpected impact. And so I don't think it's possible to experiment too much. You have to allocate sometimes to these high risk, high reward ideas. We're going to try something that's most likely to fail, but if it does win, it's going to be a home run. And you have to be ready to understand and agree that most will fail. And it's amazing how many times I've seen people come up with new designs or a radical new idea and they believe in it and that's okay. I'm just cautioning them all the time to say, hey, if you go for something big, try it out, but be ready to fail 80% of the time.

    2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source