ZHENESJAKOTHVIRUFRAR

Hypothesis Testing

One-Line Definition

Hypothesis testing is a formal statistical procedure for deciding whether a claim about a population — such as "our new checkout flow converts better" — is supported by sample data, or whether the observed effect is more likely just random noise.


Real-Life Analogy: The Courtroom Trial

Think of hypothesis testing as a criminal trial. The defendant starts with a presumption of innocence — this is the null hypothesis (H₀), the default position that "nothing is happening, there is no effect." The prosecution presents evidence — this is your sample data. The jury does not prove guilt with absolute certainty; instead, it asks: *if the defendant were truly innocent, how likely is it that we would see evidence this strong?* If that probability is sufficiently small, the jury rejects the presumption of innocence and convicts.

Critically, a verdict of "not guilty" is not the same as "innocent." In statistics, failing to reject H₀ never proves H₀ is true — it only means the data did not provide strong enough evidence against it. And just as a jury can convict an innocent person (Type I error) or acquit a guilty one (Type II error), hypothesis tests carry two well-defined error risks that we control with thresholds like α = 0.05.


Core Formula

The workhorse of hypothesis testing is the test statistic, which measures how far your observed result sits from what the null hypothesis predicts, in units of standard error:

$$z = \frac{\bar{x} - \mu_0}{\sigma / \sqrt{n}}$$

Where:

- x̄ = sample mean (your observed result)

- μ₀ = the value claimed under the null hypothesis

- σ = population standard deviation (or sample standard deviation *s* for a t-test)

- n = sample size

The denominator σ/√n is the standard error. Notice the implication: as sample size grows, the standard error shrinks, so even a small effect becomes statistically detectable. This is why a 0.2% lift can be "significant" with 5 million visitors but invisible with 5,000.

You then convert the z-score into a p-value — the probability of observing a result at least this extreme if H₀ were true. If p < α (commonly 0.05), you reject H₀.


Comparison with Related Terms

TermWhat it answersKey outputTypical DTC use
**Hypothesis Testing**Is this effect real, or noise?p-value, reject/fail to reject H₀"Did the new landing page lift conversion?"
**Confidence Interval**What range plausibly contains the true effect?Interval, e.g., [+0.3%, +1.8%]"How much did conversion actually improve?"
**A/B Test**Which of two variants performs better?Winner + significanceTesting two ad creatives head-to-head
**Statistical Power**Could we have detected a real effect?1 − β, typically 0.80Sizing a test before launch
**Bayesian Inference**What's the probability H₁ is true given data?Posterior distributionContinuous monitoring of test results

Note that hypothesis testing and confidence intervals are two sides of the same coin: a 95% CI that excludes zero is equivalent to rejecting H₀ at α = 0.05. A/B testing is simply hypothesis testing applied to a business experiment.


Use Cases in DTC & Cross-Border E-commerce

1. Conversion rate optimization. You test a new product page against the control. Control converts at 3.2%; variant converts at 3.6% across 40,000 sessions each. A two-proportion z-test returns p = 0.002, so you reject H₀ and ship the variant.

2. Pricing and promotion validation. Before rolling out a 15% discount site-wide, you test it on a sample. If the lift in units sold is not statistically significant, you avoid eroding margin for no reason.

3. Ad creative and channel testing. Comparing Meta vs. TikTok ROAS across markets — hypothesis testing tells you whether a 0.4 difference in ROAS reflects a genuine channel advantage or normal variance.

4. Supplier and fulfillment quality. Testing whether a new 3PL partner's defect rate (1.1%) is genuinely lower than the incumbent's (1.6%), or within noise.

5. Email and retention campaigns. Whether a win-back sequence genuinely raises 30-day repeat purchase rate, or the uplift is seasonal.


Common Misconceptions

"p = 0.03 means there's a 3% chance the null hypothesis is true." Wrong. The p-value is P(data this extreme | H₀ true), not P(H₀ | data). Confusing these is the single most common error in e-commerce analytics.

"Statistically significant = business significant." A 0.05% conversion lift can be statistically significant with 10 million sessions yet worth almost nothing in revenue. Always pair p-values with effect size and confidence intervals.

"Failing to reject H₀ proves there's no effect." Absence of evidence is not evidence of absence. Your test may simply be underpowered — with n = 500, you'd miss most realistic effects.

"You can peek at results and stop when significant." Repeatedly checking a fixed-horizon test inflates the false positive rate dramatically — peeking daily at α = 0.05 can push the real Type I error rate above 20%. Use sequential testing or pre-register the sample size.

"More tests = more reliable conclusions." Running 20 variants against one control at α = 0.05 means roughly a 64% chance at least one shows a false "win" (1 − 0.95²⁰). This is the multiple comparisons problem; correct with Bonferroni or Benjamini-Hochberg.

"Statistical significance tells you which variant is better." It tells you the difference is unlikely to be zero — not that the effect is large, durable, or generalizes to other markets.


Related Terms

- Null Hypothesis (H₀) — the default "no effect" claim being tested

- Alternative Hypothesis (H₁) — the effect you're trying to detect

- p-value — probability of data at least this extreme under H₀

- Significance Level (α) — your false-positive tolerance, usually 0.05

- Type I Error — rejecting a true H₀ (false positive)

- Type II Error (β) — failing to reject a false H₀ (false negative)

- Statistical Power — 1 − β, the chance of detecting a true effect

- Effect Size — magnitude of the difference, independent of sample size

- Confidence Interval — range of plausible values for the true effect

- A/B Test — the most common applied form of hypothesis testing in e-commerce

- Sample Size Calculation — pre-test planning to achieve target power

Mastering hypothesis testing is what separates DTC operators who scale on evidence from those who scale on hunches. Used well, it is a disciplined filter for the endless stream of "this might work" ideas — and a guardrail against shipping changes that only looked good by luck.