B · Statistics Cheat Cards
These cards hold the formulas behind Chapters 2, 4 and 16–23, one method per card, for the day you need the exact version. Find the card in the list below, check the line that says what it assumes, and follow the chapter link when you need the whole story.
Each card has the same parts: the question it answers (in italics), the formula, In words, Use it when, Assumes, Common mistake, and the chapter that teaches it. A few cards add one worked number from Steep. The grey names at the foot of a card are the papers behind the method, listed under Sources.
Describing data: Mean, median and percentiles · Variance and standard deviation · Weighted averages and Simpson’s paradox · Regression to the mean
Uncertainty: Standard error · Confidence interval · The bootstrap
Testing: p-value, alpha and the two errors · Two proportions · Two means (Welch) · Relative lift · Ratio metrics · Sample ratio mismatch · Multiple comparisons
Planning a test: Sample size per group · Power and the MDE
Faster tests: CUPED · Sequential testing · Bayesian A/B
Without a coin: Difference-in-differences · Regression discontinuity · Propensity scores
| Symbol | What it means, unless a card says otherwise |
|---|---|
| \(n\), \(n_A\), \(n_B\) | Number of units (usually users): in all, in group A, in group B |
| \(\bar x\), \(s\), \(s^2\) | Average (mean), sample standard deviation, sample variance |
| \(\hat p\) | A share measured in the data. A hat means “estimated from the data” |
| \(\operatorname{SE}\) | Standard error: the standard deviation of an estimate |
| \(\alpha\), \(\beta\) | False-alarm rate and miss rate. Power is \(1 - \beta\) |
| \(z_q\) | The point below which a share \(q\) of the standard normal curve lies: \(z_{0.975} = 1.96\), \(z_{0.95} = 1.645\), \(z_{0.80} = 0.84\) |
| \(\Phi(z)\) | The share of the standard normal curve below \(z\) |
| \(\log\), \(\ln\) | The natural logarithm (base \(e\)) |
| \(\lfloor h \rfloor\) | \(h\) rounded down to a whole number |
| \(E[T \mid X]\) | The average of \(T\) among units with this value of \(X\) |
| \(\operatorname{Var}\), \(\operatorname{Cov}\) | Variance; covariance (how two numbers move together) |
| \(\propto\) | “Proportional to”: grows in step with |
| \(\sim \operatorname{Beta}(a, b)\) | “Follows the Beta distribution with numbers \(a\) and \(b\)” |
| \(\max_{k \le K}\) | The largest value over \(k = 1, \dots, K\) |
| \(\Lambda_m\) | A likelihood ratio: how many times more likely the data are if there is an effect than if there is none |
Other letters, such as \(T\), \(X\) and \(Y\), are defined on the card where they appear, and can mean different things on different cards: \(T\) is a true value on one card, the treated group on another.
Describing data
Mean, median and percentiles · 均值、中位数与百分位数
What is a typical value, and how are the values spread?
\[\bar x = \frac{1}{n}\sum_{i=1}^{n} x_i\]
Sort the values, \(x_{(1)} \le \dots \le x_{(n)}\). The median is \(x_{((n+1)/2)}\) when \(n\) is odd, and the mean of the two middle values when \(n\) is even. Percentile \(q\) (straight-line rule), with \(h = (n-1)\,q + 1\):
\[Q(q) = x_{(\lfloor h \rfloor)} + (h - \lfloor h \rfloor)\big(x_{(\lfloor h \rfloor + 1)} - x_{(\lfloor h \rfloor)}\big)\]
In words The mean shares the total out equally. The median is the middle value. Percentile \(q\) has a share \(q\) of the values below it.
Use it when you need a total (mean × \(n\) = total: use the mean) or a typical customer (use the median and a few percentiles, above all for money, which is skewed).
Assumes nothing: it is arithmetic. But one wild value can move the mean as far as it likes; the median needs half the values to be wild.
Common mistake Reporting only the mean of skewed values. Also, tools differ when a percentile falls between two values: pandas, NumPy and DuckDB’s quantile_cont use the straight-line rule above, while quantile_disc returns a value from the data.
Variance and standard deviation · 方差与标准差
How far is a single value from the average, typically?
\[s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar x)^2, \qquad s = \sqrt{s^2}\]
For a yes/no value (1 = yes) with share \(\hat p\): \(s^2 \approx \hat p\,(1 - \hat p)\).
In words The standard deviation (SD) is the typical distance of one value from the average, in the values’ own unit.
Use it when you describe the spread of single values. It is also the input to every standard error and every sample size.
Assumes nothing for the number itself. Dividing by \(n - 1\), not \(n\), corrects a small bias: values sit a little closer to their own average than to the true mean.
Common mistake Mixing up SD and SE: the SD describes single values and does not shrink with more data; the SE describes an estimate and does. In code, NumPy’s np.std divides by \(n\) unless you pass ddof=1; pandas divides by \(n - 1\).
Weighted averages and Simpson’s paradox · 加权平均与辛普森悖论
Why can a total go one way while every segment goes the other?
For a group with a share \(w_c\) of its units in segment \(c\) and a rate \(r_c\) there:
\[R = \sum_c w_c\, r_c, \qquad \sum_c w_c = 1\]
The gap between two groups \(M\) and \(K\) splits in two:
\[\begin{aligned} R_M - R_K = {} & \textstyle\sum_c w_{M,c}\,(r_{M,c} - r_{K,c}) \\ & + \textstyle\sum_c (w_{M,c} - w_{K,c})\, r_{K,c} \end{aligned}\]
The first line compares the groups inside each segment; the second is the mix. A standardised rate gives both groups the same mix \(w^*\): \(R^{\text{std}} = \sum_c w^*_c\, r_c\).
In words A total rate is the segments’ rates, weighted by the group’s mix. If the mixes differ, the totals can disagree with every segment.
Use it when two groups spread differently across segments (cities, platforms) whose rates differ.
Assumes that the segments existed before the thing you compare, and were not changed by it.
Common mistake Splitting by something the treatment itself changes, such as “users who reached checkout” in a checkout test: the split can hide an effect or invent one.
Regression to the mean · 均值回归
Why is a record week usually followed by a quieter one?
If a measurement is a true value plus noise, \(X = T + \varepsilon\), and \(\mu\) is the long-run average, the best guess for the true value is
\[E[T \mid X] = \mu + \rho\,(X - \mu)\]
\[\rho = \frac{\operatorname{Var}(T)}{\operatorname{Var}(T) + \operatorname{Var}(\varepsilon)}\]
In words An extreme value is partly luck. The next one is expected closer to the average, and the pull is stronger when more of the variation is noise (smaller \(\rho\)).
Use it when you picked units because they were extreme (the best stores, the worst week, the top spenders) and you look at them again.
Assumes \(T\) and \(\varepsilon\) independent and normal, with \(\varepsilon\) averaging 0. The formula is exact then, and a good guide otherwise.
Common mistake Giving the credit for the fall (or the rise) to whatever you did in between. Compare with units picked the same way that got nothing. Matching on one noisy month has the same trap.
Uncertainty
Standard error · 标准误
How much would this estimate move if you drew a new sample?
\[\operatorname{SE}(\bar x) = \frac{s}{\sqrt n}, \qquad \operatorname{SE}(\hat p) = \sqrt{\frac{\hat p\,(1 - \hat p)}{n}}\]
For two independent estimates \(\hat a\) and \(\hat b\):
\[\operatorname{SE}(\hat a - \hat b) = \sqrt{\operatorname{SE}(\hat a)^2 + \operatorname{SE}(\hat b)^2}\]
In words The SE is the standard deviation of an estimate across all the samples you might have drawn. It shrinks with the square root of \(n\): four times the data, half the SE.
Use it when you report any estimate. It is the input to intervals, tests and power.
Assumes independent units: knowing one value tells you nothing about another. For the difference, the two estimates are independent too.
Common mistake Using \(s / \sqrt n\) when units are linked, such as days in a wet spell or orders from one user: the real SE is usually larger. Use the unit you randomised, or a block bootstrap.
Confidence interval for a mean or a share · 置信区间
Which values of the true number fit the data?
\[\bar x \pm 1.96\,\frac{s}{\sqrt n}, \qquad \hat p \pm 1.96\sqrt{\frac{\hat p\,(1 - \hat p)}{n}}\]
Here \(1.96 = z_{0.975}\). For the mean of a small sample, use \(t_{n-1,\,0.975}\) in its place (2.18 for \(n = 13\); close to 1.96 by \(n = 100\)).
In words The estimate, plus or minus about two standard errors. The method catches the true value in about 95 of 100 samples, if its assumptions hold.
Use it when you report an estimate. Put the interval next to zero, and next to the smallest effect that matters.
Assumes independent units, and enough of them for the average to be close to normal (the central limit theorem). For a share, a common guide is at least 10 “yes” and 10 “no”.
Common mistake Saying “a 95% chance that the true value is in this interval”. The 95% belongs to the method, not to one interval (a Bayesian credible interval makes that statement, with a prior). Near 0% or 100%, or with few “yes”, this interval is too narrow and can pass 0% or 100%; use a Wilson interval instead, a formula that stays between 0% and 100% (most statistics tools have it).
The bootstrap · 自助法
How uncertain is an estimate that has no simple formula?
Draw \(n\) units from the sample with replacement, compute the estimate \(\hat\theta^*\) on them, and repeat \(B\) times (a few thousand):
\[\widehat{\operatorname{SE}} = \operatorname{sd}\big(\hat\theta^*_1, \dots, \hat\theta^*_B\big)\]
\[\text{95\% interval} = \big[\,q_{0.025}(\hat\theta^*),\; q_{0.975}(\hat\theta^*)\,\big]\]
where \(q\) is a percentile of the \(B\) values (a percentile interval). When neighbours are linked, the block bootstrap draws whole blocks of \(\ell\) neighbouring units: overlapping blocks (Künsch) or non-overlapping blocks such as calendar weeks, \(\ell = 7\) days (Carlstein).
In words Treat the sample as the population, draw new samples from it, and watch how much the estimate moves.
Use it when the estimate comes from a model, or is a ratio, a median or anything else without an SE formula.
Assumes the resampled units (or blocks) are independent; there are not too few of them; the estimate is a smooth function of the data.
Common mistake Resampling the wrong unit (orders, when users were randomised). Trusting it for a dozen values (too narrow) or for extremes such as the largest value.
Chapter 17 Chapter 21 Efron (1979); Carlstein (1986); Künsch (1989)
Testing
p-value, alpha and the two errors · p 值、显著性水平与两类错误
If nothing had changed, how surprising would these data be?
For a test statistic \(z\) that is standard normal when nothing changed (\(H_0\), the null hypothesis), two-sided, and one-sided for the direction “up”:
\[p_{\text{two-sided}} = 2\,\big(1 - \Phi(|z|)\big)\]
\[p_{\text{one-sided}} = 1 - \Phi(z)\]
\[\begin{aligned} \alpha &= P(\text{reject } H_0 \mid H_0 \text{ true}) \\ \beta &= P(\text{do not reject } H_0 \mid \text{a real effect}) \end{aligned}\]
In words The p-value is the chance of data at least this extreme if nothing changed. Reject \(H_0\) when \(p < \alpha\). A Type I error is a false alarm (rate \(\alpha\), chosen by you). A Type II error is a miss (rate \(\beta\), set by the effect size, the noise and \(n\)).
Use it when you need a yes-or-no decision with a known false-alarm rate. Two-sided by default. One-sided only if the direction was fixed before the data, and a result the other way would not change what you do.
Assumes alpha, the side and the test were chosen before you looked, and the picture of “ordinary” (the null distribution) is right.
Common mistake Reading \(p\) as “the chance that nothing changed”, or “not significant” as “no effect”. Check the interval and the power.
Chapter 18 Wasserstein & Lazar (2016)
Difference of two proportions · 两个比例之差
Is B’s conversion different from A’s, and by how many points?
With \(\hat p_A = x_A / n_A\), \(\hat p_B = x_B / n_B\) (\(x\) counts the “yes”) and the pooled rate \(\hat p = (x_A + x_B) / (n_A + n_B)\):
\[z = \frac{\hat p_B - \hat p_A}{\sqrt{\hat p\,(1 - \hat p)\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}\]
and \(p = 2\,\big(1 - \Phi(|z|)\big)\). The 95% interval is \((\hat p_B - \hat p_A) \pm 1.96\,\operatorname{SE}\), with each group’s own variance:
\[\operatorname{SE} = \sqrt{\frac{\hat p_A(1 - \hat p_A)}{n_A} + \frac{\hat p_B(1 - \hat p_B)}{n_B}}\]
In words The gap in points, divided by the gap that chance alone would give.
Use it when the metric is yes-or-no per randomised unit: converted, came back.
Assumes independent units, analysed at the level you randomised; groups large enough for the normal curve.
Common mistake Mixing points and percent. Report the difference in points and the relative lift, and label each. (The test pools the groups and the interval does not, so at the edge of 0.05 the two can disagree.)
At Steep In checkout_v2 (final statuses, see Chapter 1), 66.10% of group A and 67.14% of group B ordered: +1.04 points, p = 0.004.
Difference of two means (Welch) · 两个均值之差(Welch t 检验)
Is B’s average per user (orders, revenue) different from A’s?
Write \(v_A = s_A^2 / n_A\) and \(v_B = s_B^2 / n_B\) (each group’s squared standard error). Then
\[t = \frac{\bar x_B - \bar x_A}{\sqrt{v_A + v_B}}\]
Compare \(t\) with a \(t\) distribution with Welch’s degrees of freedom, \(\nu\):
\[\nu = \frac{(v_A + v_B)^2}{v_A^2 / (n_A - 1) + v_B^2 / (n_B - 1)}\]
The 95% interval is \((\bar x_B - \bar x_A) \pm t_{\nu,\,0.975}\,\operatorname{SE}\), where \(\operatorname{SE}\) is the denominator (the bottom part) of \(t\). With thousands of users, \(t_{\nu,\,0.975} \approx 1.96\).
In words The gap in averages, divided by its standard error, with each group keeping its own spread.
Use it when the metric is a number per user. The book’s steep.metrics.compare_means runs this test.
Assumes independent units; group averages close to normal. Metrics with many very large values, such as revenue, need more users.
Common mistake Student’s pooled-variance \(t\)-test when the group sizes and spreads both differ (with equal sizes it gives the same \(t\)); or a per-order analysis when users were randomised.
Chapter 20 Chapter 21 Welch (1947)
Relative lift and its interval · 相对提升及其区间
By what percent is B higher than A, and how sure are we?
\[\widehat{\text{lift}} = \frac{\bar B}{\bar A} - 1\]
\[\operatorname{Var}\!\left(\frac{\bar B}{\bar A}\right) \approx \frac{s_B^2}{n_B\,\bar A^2} + \frac{\bar B^2\, s_A^2}{n_A\,\bar A^4}\]
95% interval: \(\widehat{\text{lift}} \pm 1.96 \sqrt{\operatorname{Var}}\). For a share, \(\bar A = \hat p_A\) and \(s_A^2 = \hat p_A (1 - \hat p_A)\).
In words A ratio has two sources of noise: B’s own, and A’s, because the lift is measured against A. This formula is the delta method: an approximation that treats the ratio as a straight line near the averages.
Use it when you report a change in percent, as most A/B readouts do.
Assumes independent groups of independent users; \(\bar A\) well away from 0; large groups.
Common mistake Comparing a relative lift with an absolute MDE, or the reverse: a 2% lift on a 66% conversion is 1.3 points, not 2.
At Steep checkout_v2 conversion: +1.57%, 95% interval +0.50% to +2.65%, as in Chapter 20.
Chapter 20 Chapter 21 Deng, Knoblich & Lu (2018)
Ratio metrics per user · 比率指标(Delta 方法)
How uncertain is a ratio such as revenue per order, when users were randomised?
In one group, user \(i\) has totals \(Y_i\) (say, revenue) and \(X_i\) (say, orders), for \(n\) users:
\[R = \frac{\bar Y}{\bar X}\]
\[\operatorname{Var}(R) \approx \frac{s_Y^2 - 2R\,s_{XY} + R^2 s_X^2}{n\,\bar X^2}\]
where \(s_Y^2\), \(s_X^2\) and \(s_{XY}\) are the variances and the covariance of the per-user totals. For B minus A, add the two groups’ variances.
In words Count each user, with all their orders, as one independent unit.
Use it when a metric divides two sums, and the unit you divide by (orders, sessions) is smaller than the unit you randomised (users).
Assumes independent users; large \(n\); \(\bar X\) well away from 0.
Common mistake Treating every order as independent, with variance \(s^2_{\text{order}} / \sum X_i\). Orders from one person are linked, so that SE is too small and false alarms rise. A user-level bootstrap is the check.
Chapter 21 Deng, Knoblich & Lu (2018)
Sample ratio mismatch (SRM) test · 样本比例失衡检验
Did the coin split the users as planned?
With \(N\) users and a planned share \(s\) in B, the expected counts are \(E_A = N(1 - s)\) and \(E_B = Ns\). With the observed counts \(O_A\) and \(O_B\):
\[\chi^2 = \frac{(O_A - E_A)^2}{E_A} + \frac{(O_B - E_B)^2}{E_B}\]
With two groups it has one degree of freedom: \(p = P(\chi^2_1 > \chi^2)\).
In words How far the counts are from the plan, compared with what a fair coin does. Many teams call a split broken only when \(p < 0.001\): the check runs on every test, and real SRM bugs give far smaller p-values.
Use it when you are about to read any result: on every test, and by day and by platform.
Assumes each user is assigned once, independently, with the planned share.
Common mistake Reading the result anyway, or “fixing” it by weighting the groups back to 50/50: weights cannot bring back missing users. Find out where they went.
At Steep checkout_v2 put 50.01% of 69,045 users in B (p = 0.973): no sign of SRM.
Chapter 21 Chapter 20 Fabijan et al. (2019)
Multiple comparisons · 多重比较
How do I keep false alarms in check when I read many tests?
With \(m\) independent tests and no real effects, the chance of at least one false alarm is \(1 - (1 - \alpha)^m\): 0.40 for \(m = 10\) and \(\alpha = 0.05\). Sort the p-values, \(p_{(1)} \le \dots \le p_{(m)}\).
- Bonferroni: reject test \(i\) if \(p_i \le \alpha / m\).
- Holm: reject in order while \(p_{(j)} \le \alpha / (m - j + 1)\); stop at the first that fails.
- Benjamini–Hochberg: find the largest \(k\) with \(p_{(k)} \le \frac{k}{m}\,\alpha\); reject the \(k\) smallest.
- Benjamini–Yekutieli: as Benjamini–Hochberg, with \(\alpha\) replaced by \(\alpha / \sum_{i=1}^{m} 1/i\) (about \(\alpha / 2.9\) for \(m = 10\)).
In words Bonferroni and Holm keep the family-wise error rate (FWER), the chance of any false alarm, at or below \(\alpha\). Benjamini–Hochberg keeps the false discovery rate (FDR), the expected share of false alarms among the results you call significant, at or below \(\alpha\): with \(V\) false alarms among \(R\) results called significant, \(\text{FDR} = E[V / \max(R, 1)]\).
Use it when one test is read on many metrics, segments or variants. Holm for decisions (it never rejects less than Bonferroni); Benjamini–Hochberg for a first screen of leads.
Assumes you count every test you ran. Bonferroni and Holm work for any dependence. Benjamini–Hochberg needs tests that are independent or that tend to move in the same direction together (for example, related metrics of one test); Benjamini–Yekutieli works for any dependence.
Common mistake Correcting only the tests you report; trusting the size of the top result.
Chapter 21 Holm (1979); Benjamini & Hochberg (1995); Benjamini & Yekutieli (2001)
Planning a test
Sample size per group · 每组样本量
How many users does each group need?
The MDE (minimum detectable effect) is the smallest true effect the test should catch with the planned power. Here it is relative: a share of the baseline.
For a share, from the baseline \(p_1\) to \(p_2 = p_1 (1 + \text{MDE})\), split 50/50, with \(V = p_1(1 - p_1) + p_2(1 - p_2)\):
\[n = \frac{(z_{1-\alpha/2} + z_{1-\beta})^2\, V}{(p_2 - p_1)^2}\]
For alpha 0.05 (two-sided) and power 80%: \(z_{0.975} + z_{0.80} = 1.96 + 0.84 = 2.8\).
For a mean with standard deviation \(\sigma\) and a change \(\delta\):
\[n = \frac{2\,(z_{1-\alpha/2} + z_{1-\beta})^2\, \sigma^2}{\delta^2} \approx \frac{16\, \sigma^2}{\delta^2}\]
The 16 is \(2 \times 2.8^2 \approx 16\), for the same alpha and power. A share \(s\) in B instead of half needs about \(1 / \big(4 s (1 - s)\big)\) times as many users in total when the two variances are close: 2.8 times for 90/10.
In words Users needed grow with (noise ÷ effect)². Halve the effect, and you need four times the users.
Use it when you plan, before the test starts. Then turn users into days with real traffic, and round up to whole weeks.
Assumes independent users; baseline and noise from recent whole weeks; the normal curve. Some calculators pool the variance in the first term, which gives a slightly different \(n\).
Common mistake Sizing for the effect you hope for, not the smallest one that matters.
Example A 10% baseline and a 5% relative MDE: \(p_1 = 0.10\) and \(p_2 = 0.105\), a change of 0.5 points. With alpha 0.05 and power 80%, each group needs about 57,800 users.
Power and the MDE · 统计功效与最小可检测效应
How likely is the test to catch an effect of a given size?
For a true difference \(\Delta\) whose estimate has standard error \(\operatorname{SE}\):
\[\begin{aligned} 1 - \beta = {} & \Phi\!\left(\frac{|\Delta|}{\operatorname{SE}} - z_{1-\alpha/2}\right) \\ & + \Phi\!\left(-\frac{|\Delta|}{\operatorname{SE}} - z_{1-\alpha/2}\right) \end{aligned}\]
The second term (a significant result in the wrong direction) is tiny. Turned around:
\[\text{MDE} \approx (z_{1-\alpha/2} + z_{1-\beta})\,\operatorname{SE} \approx 2.8\,\operatorname{SE}\]
for alpha 0.05 and power 80%, in the units of \(\Delta\). As \(\operatorname{SE} \propto 1/\sqrt n\), the MDE falls with the square root of \(n\).
In words Power is the chance of a significant result when the effect is really there. The minimum detectable effect (MDE) is the smallest effect the test finds with the planned power.
Use it when you plan a test, and when you read a “not significant” result: ask about the power for an effect that matters.
Assumes an estimate close to normal, with an SE known from planning data.
Common mistake Power computed for the effect this same test measured: it adds nothing (it is the p-value in another form). And “twice the users, half the MDE”: no, about 30% smaller. A low-power test that wins usually wins too big (the winner’s curse).
Chapter 19 Chapter 18 Gelman & Carlin (2014)
Faster tests
CUPED · CUPED 方差缩减
Can data from before the test make the interval narrower?
With the outcome \(Y\) and a covariate \(X\) measured before each user’s exposure:
\[Y^{\text{cuped}} = Y - \theta\,(X - \bar X)\]
\[\theta = \frac{\operatorname{Cov}(Y, X)}{\operatorname{Var}(X)}\]
with \(\theta\) and \(\bar X\) from both groups together. Then
\[\operatorname{Var}(Y^{\text{cuped}}) = \operatorname{Var}(Y)\,(1 - \rho^2)\]
where \(\rho\) is the correlation of \(Y\) and \(X\). The test is as precise as a plain test with \(1 / (1 - \rho^2)\) times as many users.
In words Compare what users did with what their own past predicted. The past cannot be changed by the treatment, so the comparison stays fair, with less noise.
Use it when a good pre-test covariate exists; often it is the same metric, before the test. Decide it in the plan.
Assumes \(X\) is fixed before exposure, so on average \(\bar X_B - \bar X_A = 0\) and no bias is added.
Common mistake A covariate measured after exposure (it can remove the real effect); expecting help for users with no history; choosing CUPED after seeing the plain result.
Chapter 22 Deng et al. (2013)
Sequential testing · 序贯检验
How can I look early without raising false alarms?
If you look \(K\) times and use the line 1.96 each time, the chance of a false alarm is more than \(\alpha\):
\[P\big(\max_{k \le K} |Z_k| > 1.96\big) > \alpha\]
Alpha spending moves the line. The information fraction \(t\), between 0 and 1, says how precise the estimate is now, compared with the end. The two classic spending functions (Lan and DeMets, 1983) are of O’Brien–Fleming type and Pocock type:
\[\alpha_{\text{OBF}}(t) = 2 - 2\,\Phi\!\left(\frac{z_{1-\alpha/2}}{\sqrt t}\right)\]
\[\alpha_{\text{Pocock}}(t) = \alpha \ln\big(1 + (e - 1)\,t\big)\]
At look \(k\), the line \(c_k\) is set so that, with no effect,
\[\begin{aligned} &P(\text{first crossing at look } k) \\ &\quad = \alpha(t_k) - \alpha(t_{k-1}) \end{aligned}\]
They approximate the original designs, which for equal looks use one constant line (Pocock, 1977) and \(c\sqrt{K/k}\) (O’Brien and Fleming, 1979). The mSPRT gives an always-valid p-value. It uses \(\Lambda_m\), a likelihood ratio after \(m\) users: how many times more likely the data are if there is an effect than if there is none, averaged over true effects of typical size \(\tau\):
\[p_n = \min\big(1,\ \min_{m \le n} 1/\Lambda_m\big)\]
With no effect, the chance that \(p_n\) ever falls below \(\alpha\) is at most \(\alpha\).
In words Alpha is a budget, and every look spends some. O’Brien–Fleming saves most of it for the end; Pocock uses the same line at every look, which spends more of it early.
Use it when you must be able to stop early for a clear win or clear harm: O’Brien–Fleming-type lines for the main decision, always-valid p-values for guardrails watched daily.
Assumes the spending function (for the mSPRT, its \(\tau\)) was fixed in advance, and when you look does not depend on the results so far; each outcome is complete when counted.
Common mistake Stopping at the first \(p < 0.05\); reporting the lift of the stopping day as the effect (early stops overstate it).
Chapter 21 Chapter 22 Pocock (1977); O’Brien & Fleming (1979); Lan & DeMets (1983); Johari et al. (2017)
Bayesian A/B for a share · 贝叶斯 A/B(比例)
What is the chance that B is better, and what do I lose if I pick B and it is not?
Give each group’s rate a \(\operatorname{Beta}(a, b)\) prior (flat: \(a = b = 1\)). After \(x_A\) “yes” among \(n_A\) users, the posterior is
\[p_A \mid \text{data} \sim \operatorname{Beta}(a + x_A,\; b + n_A - x_A)\]
and the same for B. From many draws of the two posteriors:
\[P(\text{B better}) = P(p_B > p_A \mid \text{data})\]
\[L_B = E\big[\max(p_A - p_B,\ 0) \mid \text{data}\big]\]
\(L_B\) is the expected loss of choosing B: the conversion you give up, counting zero when B is better.
A 95% credible interval for the lift: the 2.5th and 97.5th percentiles of \(p_B / p_A - 1\).
In words Start from a belief, update it with the data, and read direct answers from the result.
Use it when the decision maker asks “how likely is B better?”, or a sceptical prior should pull small, noisy effects toward zero.
Assumes independent users and a prior chosen before the data. Every answer holds under that prior: with little data, a different prior gives a different \(P(\text{B better})\), so report the prior. With a flat prior and much data, \(P(\text{B better}) \approx \Phi(z)\): one minus the one-sided p-value.
Common mistake “Bayesian, so I can peek.” Shipping when \(P(\text{B better}) > 95\%\), checked every day, gives many false wins. Check the stopping rule on A/A tests.
Without a coin
Difference-in-differences · 双重差分
What did a change do to one group, with another group as the “what if”?
The 2×2 version, on the log scale so that changes are percent. \(T\) is the treated group, \(C\) the comparison group, \(Y\) a total before or after:
\[\hat\delta = \log\frac{Y_{T,\text{after}}}{Y_{T,\text{before}}} - \log\frac{Y_{C,\text{after}}}{Y_{C,\text{before}}}\]
The effect, where \(g\) is each group’s change, is
\[e^{\hat\delta} - 1 = \frac{1 + g_T}{1 + g_C} - 1\]
The regression version, for city \(c\) on day \(t\):
\[\begin{aligned} \log(\text{orders}_{ct}) = {} & a_c + \lambda_t + \gamma\,\text{rain}_{ct} \\ & + \delta\, D_{ct} + \varepsilon_{ct} \end{aligned}\]
with \(a_c\) one level per city, \(\lambda_t\) one level per day, and \(D_{ct} = 1\) for the treated city from the start date, 0 otherwise.
In words The treated group’s change, minus the comparison group’s change.
Use it when a change hit one group at a known time, and similar groups were not touched.
Assumes parallel trends: without the change, both would have moved by the same percent; and nothing else hit only the treated group then. Check the trends before the start; try fake start dates (placebo tests).
Common mistake Standard errors that ignore links between days are usually too narrow. With one treated unit, the usual repair (standard errors that treat each city’s days as one linked group, called clustering) cannot work; use the fake dates.
Chapter 23 Bertrand, Duflo & Mullainathan (2004)
Regression discontinuity · 断点回归
What does crossing a sharp line in a rule do?
With \(x_i\) the running variable minus the cutoff, \(G_i = 1\) when \(x_i \ge 0\), and only the units with \(|x_i| \le h\) (the bandwidth):
\[y_i = a + b\,x_i + \tau\,G_i + c\,x_i G_i + e_i\]
\(\tau\) is the jump at the cutoff (a sharp RD). When the line changes only part of the treatment (a fuzzy RD), divide the two jumps:
\[\hat\tau_{\text{fuzzy}} = \frac{\text{jump in the outcome}}{\text{jump in the treatment}}\]
In words Compare units right above and right below the line, each side with its own slope.
Use it when a rule gives the treatment at a cutoff of a number: spending, a score, a date.
Assumes units cannot place themselves exactly on one side, and nothing else changes at the line. Check the counts on each side, and traits the treatment cannot change. A narrow \(h\) compares more similar units, but fewer: report several. Many tools also give points near the cutoff more weight.
Common mistake Reading \(\tau\) as the effect for everyone: it describes units near the cutoff.
Chapter 23 Imbens & Lemieux (2008)
Propensity scores · 倾向得分(匹配与加权)
How can I compare treated and untreated people who differ for other reasons?
\[e(x) = P(\text{treated} \mid X = x)\]
Matching (the effect on the treated): compare inside cells \(s\) of look-alike units, with \(n_{1s}\) treated units in cell \(s\):
\[\hat\tau = \sum_s \frac{n_{1s}}{n_1}\,(\bar y_{1s} - \bar y_{0s})\]
Weighting (the average effect on everyone, ATE), with \(T_i = 1\) for a treated unit:
\[\hat\tau_{\text{ATE}} = \frac{1}{n}\sum_i \left[\frac{T_i\, y_i}{\hat e(x_i)} - \frac{(1 - T_i)\, y_i}{1 - \hat e(x_i)}\right]\]
In practice, scale the weights in each group to sum to 1.
In words Make the groups alike on what you measured, then compare.
Use it when you cannot randomise, and the reasons for treatment are measured.
Assumes no hidden confounders; overlap (\(0 < e(x) < 1\) for every unit compared); \(X\) measured before the treatment.
Common mistake Matching on things the treatment changed; trusting huge weights (scores near 0 or 1); skipping the balance check. A coin makes groups alike in everything; matching, only in what you measured.
Chapter 23 Rosenbaum & Rubin (1983)
Sources
The papers named on the cards. Each was checked against its publisher record (Crossref or OpenAlex).
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B, 57(1), 289–300. DOI
- Benjamini, Y., & Yekutieli, D. (2001). The Control of the False Discovery Rate in Multiple Testing under Dependency. The Annals of Statistics, 29(4), 1165–1188. DOI
- Bertrand, M., Duflo, E., & Mullainathan, S. (2004). How Much Should We Trust Differences-in-Differences Estimates? The Quarterly Journal of Economics, 119(1), 249–275. DOI
- Carlstein, E. (1986). The Use of Subseries Values for Estimating the Variance of a General Statistic from a Stationary Sequence. The Annals of Statistics, 14(3). DOI
- Deng, A., Knoblich, U., & Lu, J. (2018). Applying the Delta Method in Metric Analytics. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 233–242. DOI
- Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data. Proceedings of the Sixth ACM International Conference on Web Search and Data Mining (WSDM), 123–132. DOI
- Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics, 7(1), 1–26. DOI
- Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L., & Dmitriev, P. (2019). Diagnosing Sample Ratio Mismatch in Online Controlled Experiments. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2156–2164. DOI
- Gelman, A., & Carlin, J. (2014). Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspectives on Psychological Science, 9(6), 641–651. DOI
- Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
- Imbens, G. W., & Lemieux, T. (2008). Regression Discontinuity Designs: A Guide to Practice. Journal of Econometrics, 142(2), 615–635. DOI
- Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2017). Peeking at A/B Tests: Why It Matters, and What to Do About It. Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1517–1525. DOI
- Künsch, H. R. (1989). The Jackknife and the Bootstrap for General Stationary Observations. The Annals of Statistics, 17(3), 1217–1241. DOI
- Lan, K. K. G., & DeMets, D. L. (1983). Discrete Sequential Boundaries for Clinical Trials. Biometrika, 70(3), 659–663. DOI
- O’Brien, P. C., & Fleming, T. R. (1979). A Multiple Testing Procedure for Clinical Trials. Biometrics, 35(3), 549–556. DOI
- Pocock, S. J. (1977). Group Sequential Methods in the Design and Analysis of Clinical Trials. Biometrika, 64(2), 191–199. DOI
- Rosenbaum, P. R., & Rubin, D. B. (1983). The Central Role of the Propensity Score in Observational Studies for Causal Effects. Biometrika, 70(1), 41–55. DOI
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129–133. DOI
- Welch, B. L. (1947). The Generalization of “Student’s” Problem when Several Different Population Variances are Involved. Biometrika, 34(1–2), 28–35. DOI