7 basic things data scientists wish their PMs knew about A/B Testing

· Medium ·

6 min read Original article ↗

Gbenga Awodokun

7 basic things data scientists wish their PMs knew about A/B Testing — Credit: xkcd comics — https://xkcd.com/882/

In my experience overseeing more than 5,000 experiments at YouTube and Roku, there are 7 things every product manager should know about their A/B tests. However, after reading several tutorials and watching countless videos, I realized that much of the material is jargon-heavy and not suitable for someone without a strong statistical background, e.g., a product manager. What I share here may not be the best definitions of the concepts, but rather a more intuitive way to understand them. For something more “stats-correct”, I will recommend Kohavi’s HiPPO book. Now, let’s jump right in:

Causality and Correlation: Are the changes in metrics caused by changes introduced in the treatment? Not just happening at the same time. Correlation is not causation. That’s why your experiment must have a well-defined hypothesis and overall evaluation criteria (OEC), i.e., what you’re optimizing for. Note that many of your A/B tests (feature ideas) won’t move your metrics! By randomly and independently selecting users into groups, we prove causality in our A/B tests.

P-value: The probability of seeing a result or something more extreme given the null hypothesis is true (if your experiences A and B were the same). Sounds familiar, but doesn’t mean much to everyday people. The p-value is actually the most commonly misinterpreted concept in A/B testing. An intuitive way to think about it, though slightly less accurate, is that a low p-value means your data or results would be surprising if there were no real difference — so we reject ‘no difference’ as the explanation. A p-value < 0.05 means your result is statistically significant (the difference between your control and treatment is unlikely to be due to chance).

Confidence Interval (CI): If we repeat your A/B test many times, 95% of the intervals we calculate would contain the true parameter (target metric). Yeah, head spinning? A simpler alternative is that the CI tells you the range of effect sizes that are compatible with your data. If it doesn’t include zero, you are confident the effect isn’t nothing. Suppose an A/B test primary metric is visitors’ sessions with result: -0.03%, [−0.15%, +0.10%] is the confidence interval. No effect remains a plausible explanation since it straddles 0% — could be positive (a lift) or negative (a regression)

Confidence Interval

Sample Size: Let’s first discuss Alpha (typically α=0.05), Beta (typically β=0.20 or 20%), and Power or Sensitivity(typically 80%) in an A/B test. Alpha (α) is about false alarms or false positives (type I errors) when we reject the null, i.e., we claim there’s a change when there is none. Beta ( β) is about a miss or a false negative (type II error), i.e., there’s a change, but we fail to detect it. Therefore, power (1−β) is an 80% chance of detecting a real effect of the size you designed for. So, sample size determines whether you can achieve your chosen alpha and power for a given effect size. Let’s consider a simple example to illustrate this: a 5% lift in conversion rate with α=0.05, Power = 80% (β=0.20)

The question becomes: How many users do I need so that random noise is small enough to reliably see a 5% lift? That answer is the sample size calculation ~ 121,600 users per variant — abtestguide.com/abtestsize/

Primary and Guardrail metrics: The confusion often stems from how you define guardrail metrics. Guardrail metrics are a small set of metrics that, together with a small set of primary metrics, form the key metrics and define the hypothesis/decision condition of an experiment. For an A/B test to succeed, primary metrics improve, and none of the guardrail KPIs decrease. For a no-harm experiment, none of the guardrail metrics can decrease, and there is no need for a primary metric. Primary metrics are what you’re optimizing without hurting your guardrails, e.g, for an e-commerce website, primary: revenue per visitor, guardrails: latency.

Minimum Detectable Effect MDE and Segments:

Minimum Detectable Effect (MDE) is the change you are trying to detect. MDE is a business decision, not a statistical one. The PM must decide “what’s the smallest improvement worth shipping?” and that drives the sample size. The smaller your MDE, the larger the sample size you need; e.g., for a website with a 5% conversion rate, detecting a 5% relative change in conversion requires ~200K users (use the sample size link above).

Segments are groups within your experiment groups. For example, you can slice users in your experiment group by those visiting on Chrome, Safari / iOS, or Android, and so on. The more you slice your data, the more you increase the chance for results to be statistically significant when they are not (false positive) — in fact, 1 in 20 of your estimates will be statistically significant due to noise. Also, segments mean confidence intervals will be wider, making it harder to detect “real effects.”

Common Errors — Peeking, Novelty/Primacy, and Simple Ratio Mismatch

Peeking: Checking your test results before reaching the predetermined sample size, inflating the Type I error (false positive rate) beyond the intended significance level (remember α=0.05). Peeking is useful if you do not launch based on it (don’t announce success yet!), but you can abort if things go badly. Sequential testing exists as a proper solution when you need to peek.

Novelty/primacy error: Launching your features based on the initial lifts that may fade.

Sample Ratio Mismatch: It’s about the observed ratio deviating significantly from the expected ratio in an A/B test. System issues, like caching, can cause an imbalance between buckets. Run an SRM check (chi-squared test) on every experiment to detect systematic assignment bugs.

Everything above uses frequentist statistics, but many modern experimentation platforms now favor Bayesian methods. So instead of a p-value, you’ll get something like a “92%” probability that the test bucket beats control, which is intuitive and handles “peeking” more gracefully.

I hope you find these useful and intuitive. I would like to share more, so please leave a comment or reshare with others so they can find them as well. Grateful to Ross Raychev (my data scientist) for reviewing.