Trustworthy Online Controlled Experiments¶
Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.
Cited by¶
8 citations across 8 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Comparative Method
- Mapping Comparative Method into engineering A/B testing and benchmarking:
This sourceOperational treatment of A/B testing as the engineering realization of comparative-method logic; covers selection bias, statistical power, and unit-of-analysis problems — supports both the Knowledge-Transfer mapping and the Applied/Industry example.
- Mapping Comparative Method into engineering A/B testing and benchmarking:
- Control Sample
- And in software it is shadow deployments and canary releases holding most traffic on the baseline.
This sourceHoldout/control arms, randomized assignment, and concurrent measurement in online experimentation; canary and shadow deployments.
- And in software it is shadow deployments and canary releases holding most traffic on the baseline.
- Factorial Design
This sourceTier C (bibliography only): the modern reference on online controlled experiments / A/B testing, including multivariate (factorial) testing on digital platforms.
- Intervention
- The recognition that randomized assignment severs all incoming dependencies transfers verbatim from clinical RCTs to software A/B testing to chaos engineering: the same prime explains why A/B-test design requires independence of assignment from pre-treatment characteristics, why fault injection rather than passive monitoring is required for reliability claims, and why these techniques identify causal mechanisms that no correlational analysis of production logs can.
This sourceEstablishes that randomized assignment in A/B tests and canary deployments severs treatment from pre-treatment characteristics, identifying causal effects production logs cannot.
- The recognition that randomized assignment severs all incoming dependencies transfers verbatim from clinical RCTs to software A/B testing to chaos engineering: the same prime explains why A/B-test design requires independence of assignment from pre-treatment characteristics, why fault injection rather than passive monitoring is required for reliability claims, and why these techniques identify causal mechanisms that no correlational analysis of production logs can.
- Proxy–Target Fidelity
- The prime's management response is the industry's: pair the optimized proxy with guardrail metrics the optimizer is not allowed to game (long-horizon retention, survey-based satisfaction, complaint rates), hold out distributions to detect collapse, and treat a rising proxy with a flat or falling guardrail as evidence of fidelity erosion rather than success.
This sourceDefines guardrail (counter-) metrics paired with the optimized OEC to detect harm the primary proxy cannot see, e.g., protecting retention/quality while optimizing an engagement metric.
- The prime's management response is the industry's: pair the optimized proxy with guardrail metrics the optimizer is not allowed to game (long-horizon retention, survey-based satisfaction, complaint rates), hold out distributions to detect collapse, and treat a rising proxy with a flat or falling guardrail as evidence of fidelity erosion rather than success.
- Randomization
- However, in pragmatic real-world implementations (open-label trials, mobile-app experiments with visible assignment), maintaining full concealment is impossible or counterproductive
This sourceKohavi A/B testing randomized controlled experiments online platforms.
- However, in pragmatic real-world implementations (open-label trials, mobile-app experiments with visible assignment), maintaining full concealment is impossible or counterproductive
- Variance Bounds Selection Response
- A consumer-products company runs a mature A/B-testing programme and watches its headline conversion metric improve at three percent per quarter for eight quarters, then decay to half a percent at unchanged variant volume and unchanged statistical rigour.
This sourceA/B-testing improvement rate set by the variance of the variant population; low-variance programmes plateau.
- A consumer-products company runs a mature A/B-testing programme and watches its headline conversion metric improve at three percent per quarter for eight quarters, then decay to half a percent at unchanged variant volume and unchanged statistical rigour.
- Washout Failure
- The identical structure governs online A/B testing: re-exposing the same user pool to a new variant before behavioral adaptation to the old variant has decayed contaminates the new variant's lift estimate, and the fix is a washout gap matched to the adaptation-decay timescale or a between-users split so no user sees both.
This sourceCovers carryover/residual and novelty/primacy effects in online experiments and remedies such as washout periods and between-users assignment so no user sees both variants.
- The identical structure governs online A/B testing: re-exposing the same user pool to a new variant before behavioral adaptation to the old variant has decayed contaminates the new variant's lift estimate, and the fix is a washout gap matched to the adaptation-decay timescale or a between-users split so no user sees both.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 2 other ways.
- https://doi.org/10.1017/9781108653985 ×5
- https://www.cambridge.org/core/books/trustworthy-online-controlled-experiments/D97B26382EB0CB2DC2DB4ED3141769F1 ×1
Registry ID ref:ad40f7c11e4f · see in the full table