YouSaid · the spoken record
Ronny Kohavi
- lines on the record
- 99
- first
- 2023-07-27
- most recent
- 2023-07-27
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Finding me online is easy. It's LinkedIn. And what can people do for me? You know, understand the idea of control experiments as a mechanism to make the right data-driven decisions use science. Learn more by reading my book if you want. Again, all proceeds go to charity. And if you want to learn more, there's a class that I teach every quarter on Maven. We'll put in the notes how to find it and some discount for people who manage to stay all the way to the end of this podcast.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“All these observational studies that people made that were published, and then somehow a control experiment was run later on and proved that it was directionally incorrect. So I think there's a lot to learn about this idea of the hierarchy of evidence and share it with our family and kids and friends. And I think there's a book that's based on this. It's like how to read a book.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So there aren't enough units. Remember, I said unique 10,000 or something? Run true AB test. I do, I will say a couple of things. One is I try to emphasize to my family and friends and everybody this idea called the hierarchy of evidence. When you read something. There's a hierarchy of trust levels. If something is anecdotal, don't trust it. If there was an experiment, it was observational. Give it some bit of trust as you get more up and up to a natural experiment and a controlled experiment and multiple control experiments, your trust levels should go up. So I think that's a very important thing that a lot of people miss when they see something in the news is where does it come from? I have a talk that I've shared.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“The team presented instead of a PowerPoint, you start off with a structured document that tells you what you need, the questions you need to answer for your idea, and then review them as a team. And Amazon, these were like paper-based. Now it's all based on Word or Google Docs where people comment. And I think the impact of that was amazing. I think the ability to give people honest feedback and have them appreciate and have it stay after the meeting was in these notes on the document just amazing.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“This is something that I learned at Amazon, which is a structured narrative. So Amazon has some variants of this, which sometimes they go by the name of a six pager or something. But when I was at Amazon, I still remember that email from Jeff, which is no more PowerPoint. I'm going to force you to write a narrative. I took that to heart and many of the features that”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“He came in under a hole in the fence that was about this high. I cannot, I have a video of this thing just squishing underneath. We never would have assumed that it came from there, from the neighbor. But yeah, these things have just changed. And when you're away on a trip, it's always nice to be able to say, you know, I can see my house, everything's okay. At one point, we had a false alarm and the cops came in and had this amazing video of how they're entering the house and pulling the guns out.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Blink cameras So at Blink Camera is this small camera you stick in two AA batteries and it lasts for about six months. They claim up to two years. My experience is usually about six months. But it was just amazing to me how you can throw these things around in the yard and see things that you would never know otherwise. You know, some animals that go by. We had a skunk that we couldn't figure out how it was entering. So I threw five cameras out and I saw where he came in.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So it depends on the interview, but I'll give you when I do a technical interview, which I do less of, but one question that I love is amazing how many people it throws away for languages like C++ is. Tell me what the static qualifier does. And for multiple, you can do it for a variable, you can do it for a function. And it is just amazing that I would say more than 50% of people that interview for engineering job cannot get this and get it awfully wrong.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So I recently saw a short series called Chernobyl on the disaster. I thought it was amazingly well done Highly recommend it. Based on true events. As usual, there's some freedom for the artistic movie. It was kind of interesting at the end. They say this woman in the movie wasn't really a movie. It was a bunch of 30 data scientists, not data scientists, 30 scientists. That in real life presented all the data to the leadership of what to do.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“There's a fun book called Calling Bullshit. Which is despite the sort of name, which is a little extreme, I think, for the title, it actually has a lot of amazing insights that I love. And it sort of embodies, in my opinion, a lot of the Twimen's law of showing that things that are too extreme, your bullshit meter should go up and say, hey, I don't believe that. So that's my number one recommendation. There's a slightly older book that I love called Hard Facts, Dangerous Half-Truths, and Total Nonsense by the Stanford professors from the Graduate School of Business. Very interesting to see many of the things that we grew up with as sort of well understood turn out to have no justification. And then somewhat stranger book, which I love sort of on the verge of psychology, is called Mistakes Were Made, but not by...”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Tens of nights. They're like an agency or something, hundreds of nights. You may say, okay, let's just cap this. It's unlikely that people book more than 30 days in a given month. So that varies reduction technique will allow you to get statistically significant results faster. And a third technique is called Cupid, which is an article that we published. Again, I can give it in the notes, which uses the pre-experiment data to adjust the result. And we can show that you get the result as unbiased, but with lower variance in hence, it requires fewer users.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So I'll say a couple of things. One is if your platform is good, then when the experiment finishes, you should have a scorecard soon after. Maybe fix a day, but it shouldn't be that you have to wait a week for the data scientists. To me, this is the number one way to speed up things. Now, in terms of using the data efficiently, there are mechanisms out there under the title of variance reduction that help you reduce the variance of metrics so that you need less users, so that you can get results faster. Some examples that you might think about are capping metrics. So if your revenue metric is very skewed, maybe you say, well, if somebody purchased over $1,000, let's make that $1,000. In Airbnb, one of the key metrics, for example, is Knightsbook. Well, it turns out that some people book.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“We published a paper again. I'll give it in the notes with this sort of nice matrix of six axes and how you move from crawl to walk to run to fly and what you need to build on those axes. So one of the things that I do sometimes when I consult is I go into the Oregon and say, where do you think you are on these six axes? And that should be the guidance for what are the things you need to do next”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Build a platform that can quickly allow you to set up and run an experiment and then analyze it. I think one of the things that I will say at Airbnb is the analysis was relatively weak. And so lots of data scientists were hired to be able to compensate for the fact that the platform didn't do enough. And so, and this happens in other organizations too, where there's this trade-off, if you're building a good platform, invest in it so that more and more automation will allow people to look at the analysis without the need to involve a data scientist.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So, I think the motivation is to bring the marginal cost of experiments down to zero. So the more you self-service, right, go to a website, set up your experiment, define your targets, define the metrics that you want, right? People don't appreciate that the number of metrics starts to grow really fast. If you're doing things right at being, you could define 10,000 metrics that you wanted to be in your scorecard. Numbers, so it was so big, and people said it was computationally inefficient. We broke them into templates so that if you were launching a UI experiment, you would get this set of 2000. If you're doing a revenue experiment, you would get this set of 2000. If you're doing, so the point was.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“We decided that time on site is our OEC. And I said, wait a minute, some of your main goals as a support site is people spending more time on a support site a good thing or a bad thing. And then half the room thought that more time is better and half the room thought that more time is worse. And OEC is bad if directional you can't agree on it.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“There are some groups where you can come up with a good OEC. Some groups are harder. I remember one funny example was the Microsoft.com website, which this is not MSN, this is Microsoft.com, has multiple different constituencies that are trying to determine this is a support site and this is the ability to sell software through this site and warn you about safety and updates. It has so many goals. I remember when the team said we want to run experiments and I brought the group in and some of the managers and I said, do you know what you're optimizing for? It was very funny because they surprised me. They said, hey, Ronnie, we read some of your papers. We know there's this term called OEC.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think if you're starting out, find a place, find a team where experimentation is easy to run, and by that I mean they're launching often, right? Don't go with the team that launches every six months or office used to launch every three years. Go with the team that launches frequently they're running on sprints. They launch every week or two. Sometimes they launch daily. I mean, being used to launch multiple times a day. Make sure that you understand the question of the OEC. Is it clear what they're optimizing for?”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“My general experiences with Microsoft, where We went with this beachhead of Bing. We were running a few experiments and then we were asked to focus on Bing. And we scaled experimentation and built a platform at scale at Bing. Once Bing was successful and we were able to share all these surprising results, I think many, many more people in the company were amenable. And it was also the case that helped a lot that there's the usual cross-pollination people from Bing move out to other groups. And that helped these other groups say, hey, there's a better way to build software.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“You build 10% or do you build a 90%? I think for people starting the third party products that are available today are pretty good. This wasn't the case when I started working. So when I started building running experiments at Amazon, we were building the platform because nothing existed. Same at Microsoft. I think today there's enough vendors that provide good experimentation platforms that are trustworthy that I would say not a good way to consider using one of those.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So, if they have somebody in the org that has previously been involved experiment, that's a good way to consult internally. I think the key decision is whether you want to build or buy. There's a whole series of eight sessions that I posted on LinkedIn where I invited guest speakers to talk about this problem. So if people are interested, they can look at how what the vendors say and what agencies said about build versus buy question. And it's usually not a zero one. It's usually a both. You build some and you buy some and it's a question of.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“By the way, it's true on balance. You're probably better than 50 50, but people don't appreciate how much. 26% that I mentioned is high. And the reason that I want to be sure is that I think it leads to this idea of the learning, the institutional knowledge, which you want to be able to say is share with the org a success. And so you want to be really sure that you're successful. So by lowering the p-value, by forcing teams to work with the p-value maybe below 0.01 and do replication on hires, then you can be much more successful. And the false positive rate will be much, much lower.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“It's not 5%, it's 26%. So that's the number that you should have in your mind. And that's why when I worked at Airbnb, one of the things we did is we said, okay, if you're less than 0.05, but above 0.01, rerun, replicate. When you replicate, you can combine the two experiments and get a combined p-value using something called Fisher's method or Stauffer's method. And that gives you the joint probability, and that's usually much, much lower. So if you get 2.05s or something like that, then the probability that you've got them is much, much lower.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Of the time, or 80% of the time, and we can apply that number and compute that. We've done that in a paper that I will give in the notes so that you can assess the number that you really want, that's what's called the false positive risk. So I think that's something for people to internalize that what you really want to look at is this false positive risk, which tends to be much, much higher than the 5% that people think. So if you're, I think the classical example in the Airbnb where the failure rate was very, very high is that when you get a statistically significant result, let me actually pull the note so that I know how the actual number, if you're at Airbnb with a success rate of Airbnb search, where the success rate is only 8%, if you get a statistically significant result with a p-value less than 0.05, there is a 26% chance that this is a false positive result.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So rather than defining p values, I want to caution everybody that the most common interpretation is incorrect. P-value assumes it's a conditional probability or assumed probability. It assumes that the null hypothesis is true. And we're computing the probability that the data we're seeing matches the hypothesis. There's null hypotheses. In order to get the probability that most people want, We need to apply Bayes' rule and invert the probability from the probability of the data given the hypothesis to the probability that hypothesis given the data. For that, we need an additional number, which is the probability, the prior probability that the hypothesis that you're testing is successful or not, that's an unknown. What we do is we can take historical data and say, look, people fail two-thirds.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“I don't know if this is the right forum for explaining p-values because the definition of a p-value is simple. What it hides is very complicated. So I'll say one thing, which is Many people assign one minus p value as the probability that your treatment is better than control. So you ran an experiment, you got a p-value of 0.02. They think there's a 98% probability that the treatment is better than the control. That is wrong.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“That dinner investigate C because there's a large probability that something is wrong with the result. And I will say that 9 out of 10, when we call out Weima's law, it is the case that we find some flaw in the experiment. Now, there are obviously outliers, right, that first experiment that I shared where we promoted and made long-ad titles, that was successful, but that was replicated multiple times and double and triple checked and everything was good about it. Many other results that were so big turn out to be false. I'm a big, big fan of Flymas Law. There's a deck I can also give this in a note where I shared some real examples of Toy's law.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So Foyman's law, the general statement is if any figure that looks interesting or different is usually wrong. It was first said by this person in the UK who worked on radio media. But I'm a big fan of it. And, you know, my main... Claim to people is if the result looks too good to be true, if you suddenly moved your, you know, your normal movement of an experiment is under 1% and you suddenly have a 10% movement, hold the celebratory dinner. It was just your first reaction, right? Let's take everybody to a fancy dinner because we just improved revenue by millions of dollars.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“This is intuition. People just say, well, And then we have to be very conscientious and fight that bias and say when something looks too good to be true, investigate.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“And we noticed that people ignored it. They were starting to present results that had this banner. And so we blanked out the scorecard. We put a big red can't see this result. You have a sample ratio mismatch. Click OK to expose the result and why do we need that okay? We need that okay button because Want to be able to debug the reasons, and sometimes the metrics help you understand why you have a sample ratio mismatch. So we blanked out the scorecard. We had this button, and then we started to see that people press the button. It still presented the results of experiments with sample ratio mismatch method. So we ended up with an amazing compromise, which is every number in the scorecard was highlighted with a red line so that if you took a screenshot, other people could tell you how the sample ratio mismatched.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“When you say most common, I think the most common is bots. Somehow they hit the controller, the treatment in different proportions because you change the website, the bot may fail to parse the page and try to hit it more often. That's a classical example. Another one is just the data pipeline. We've had cases where we were trying to remove bad traffic under certain conditions and it was skewed because of the control and treatment. I've seen people that start an experiment in the middle of the site on some page, but they don't realize that some campaign is pushing people from the side. So there's multiple reasons. It is surprising how often this happens. And I'll tell you a funny story, which is When we first added this test to the platform, we just put a banner saying, you have a sample ratio, a mismatch, do not trust these results.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Amazing to see all these third party companies implement sample ratio mismatches. And all of them were reporting, oh my God, you know, 6%, 8%, 10%. So yeah, they were, it's sometimes fun to go back and say how many of your results were in the past were invalid before you had this sample ration mismatch test.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“And there's a paper that was published in 2018 where we share that at Microsoft, even though we'd be running experiments for a while, is around 8% of experiments that suffered from the sample regime mismatch. And it's a big number. Think about this. You're running 20,000 experiments a year. So many of them, 8% of them are invalid. And somebody has to go down and understand what happened here. We know that we can't trust the results, but why? So over time, you begin to understand there's something wrong with the data pipeline. There's something that happens with bots. Bots are a very common factor for causing a sample ratio mismatch. So there's a whole paper that was published by my team talks about how to diagnose sample ratio mismatches. In the last probably year and a half.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So people say, well, I don't know. It's not going to be exactly the same as 50.2 reasonable or not. Well, there's a formula that you can plug in. I have a spreadsheet available for those that are interested. And you can tell here's how many users are in control. Here's how many users have in treatment. My design was 50-50. And it tells you the probability that this could have happened by chance. Now, in a case like this, you plug in the numbers, it might tell you that this should happen one in half a million experiments. Well, unless you've run half a million experiment, very unlikely that you would get a 50.2 versus 49.8 split and therefore something is wrong with the experiment, right? Now people, I remember when we first implemented this check, we were surprised to see how many experiments suffered from this.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“There's a whole chapter of that in my book, but I'll say maybe one of the things that is the most common occurrence by far, which is a sample ratio mismatch. Now, what is a sample ratio mismatch? If you design the experiment to send 50% of users to control and 50% of users to treatment, it's supposed to be a random number or a hash function. If you get something off from 50% It's a red flag. So let's take a real example. Let's say you're running an experiment and it's large, it's got a million users, and you've got 50.2.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“between the A and the B and optimized he told them that it was statistically significant too many times optimized he learned optimizedly you know several people pointed I pointed this out in my Amazon review of the book that the optimized the authors wrote early on I said hey you're not doing this statistics correctly other you know Ramesh Jihari at Stanford pointed this out became a consultant to the company and then they fixed it but to me that's a Very good example of how to lose trust. A lot of trust in the market. They lost all this trust because they built something that had very much inflated error rate.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“People that started using Optimizedly thought that the platform was telling them they're very successful. But when they actually started to see Walatoldas, this is positive revenue, but I don't see this over time, like by now, we should have made double the money. So their questions started to come up around the trust in the platform. There's a very famous post that somebody wrote about how optimized they almost got me fired by a person who basically said, look, I came to the Oregon. I said, we have all these successes. But then I said something is wrong. And he tells of how he ran an AA test when there is no difference.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Can stop an experiment when the p value is statistically significant. That is a big mistake that inflates your, what's called type 1 error or the false positive rate materially. So if you think you've got a 5% type one error or you aim for that p-value less than 0.05. Using real time sort of p-value monitoring to optimize the offered, you would probably have a 30% error rate. So what this lend is that”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“To me, it is very important that when you present this and say, this is science, this is a control experiment, this is the result, you better believe that this is trustworthy. And so I focus on that a lot. I think it allowed us to gain the organizational trust that this is really, and the nice thing is when we built all these checks to make sure that the experiment is correct. If there were something wrong with it, we would stop and say, hey, something is wrong with the experiment. And I think that's something that some of the early implementations in other places did not do, and it was a big mistake. I mentioned this in my book, so I can mention this here, optimizely in its early days, we're very statistically naive. They sort of said, hey, we're real time. We can compute your p-values in real time.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“To me, the experimentation platform is the safety net and it's an oracle. So it serves really two purposes. The safety net means that if you launch something bad, you should be able to abort quickly, right? Safe deployments, safe velocity, there are some names for this. But this is one key value that the platform can give you. The other one, which is the more standard one, is at the end of the two-week experiment, we will tell you What happened to your key metric and do many of the other surrogate and debugging and guardrail metrics? Trust builds up. It's easy to lose. And so”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“By the way, all proceeds from the book are donated to charity, so if I'm pitching the book here, there is no financial gain for me from having more copies sold. I think we made that decision, which was a good decision. All proceeds go with the charity.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“And they were saying, so you'll be able to sell a few thousand copies and help the world. And I found my co-authors, which are great. And we wrote a book that we thought is not statistically oriented, has fewer formulas than you normally see, and focuses on the practical aspects and on trust, which is the key. The book, as I said, was more successful. It sold over 20,000 copies in English. It was translated to Chinese, Korean, Japanese, and Russian. And so it's great to see that we help the world become more data-driven with experimentation. And I'm happy because of that. I was pleasantly surprised.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“I was pleasantly surprised that It sold more than what we thought and more than what Cambridge predicted So when first we were approached by Cambridge after a tutorial that we did to write a book, I was like, I don't know. This is too small of a niche area.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Bookings went down materially, the company should suddenly not be data driven and do things differently. I think if Airbnb stayed the course, did nothing, the revenue would have gone up in the same way. In fact, if you look at one investment, one big investment that was done at the time was online experiences. And the initial data wasn't very promising. And I think today it's a footnote.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“So I think actually in a state like that it's even more important to run A B tests, right? Because what you want to be able to see is if we're making this change, is it actually helping in the current environment? You know, there's this idea of external generalizability. Is it going to work out now during COVID? Is it going to generalize later on? These are things that you can really answer with the controlled experiments. And sometimes it means that you might have to replicate them six months down when COVID, say, is not as impactful as it is. Saying that you have to make decisions quickly to me, I'll point you to the success rate. Like if in peacetime you're wrong two-thirds to 80% of the time, why would you be suddenly right in wartime? So I don't believe in the... Idea that because”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“It's a one off experiment where it's hard to assign value to some of the things that Airbnb is doing. I personally believe it could have been a lot bigger. And a lot more successful if it had run more controlled experiments. But I can't speak about some of those that I ran and that showed that some of the things that were initially untested were actually negative and could be better.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“As you know, and I'm restricted from talking about Airbnb, I will say a few things that I am allowed to say. One is in my team in Search Relevance, Everything was a B testant. So while Brian can focus on some of the design aspects, the people who are actually doing the neural networks and the search, everything was A-B tested to hell. So nothing was launching without an A-B test. We had targets around improving certain metrics and everything was done AB test. Now, other teams, some did, some did not. I will say that when you say things are going well, I think we don't know the counterfactual. I believe that had Airbnb kept people like Greg Greeley, which was pushing for a lot more data driven, and had Airbnb run more experiments, it would have been in a better state than today. But it's the counterfactual we don't know”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Unless it's a sort of a legal requirement, right? When legal comes along and says you have to do X or Y, you have to ship on flat or even negative. And that's understandable. But again, I think that's something that a lot of people make the mistake of saying legal told us we have to do this. Therefore, we're going to take the hits. No. Legal gave you a framework that you have to work under. Try three different things and ship the one that hurts the least.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“Well, it's not hurting. It's flat too negative. Some of them are flat. And by the way, flat to me, if something is not Statsig, that's a no ship. Because you've just introduced more code. There is a maintenance overhead to shipping your stuff. I've heard people say, look, we already spent all this time. The team will be demotivated if we don't ship it. And no, that's wrong, guys. Make sure that we understand that shipping this project has no value is complicating the code base. Maintenance costs will go up. You don't ship on flat.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source
“That's my rule of thumb. And you had, you know, I've heard people say it's 70% or 80%, but it's in that area where I think, you know, when you talk about how much to invest in the known versus the high risk, high reward, that's usually the right percentage that most organizations end up doing this allocation. You interviewed Treyas. I think he mentioned that Google is like 70% the search and ads and it's a 20% for some of the apps and new stuff. And then it's the 10% for infrastructure.”
2023-07-27 · Lenny's Podcast · The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon) · IDENTIFIED FROM THE TRANSCRIPT · source