The logic of a hypothesis test
A hypothesis test starts from a skeptical position. You assume nothing special is going on, then ask how surprising your sample would be if that were true. If the sample would be very surprising, you reject the assumption. If it would not, you do not.
The assumption you start with is the null hypothesis (H0). It usually says there is no difference, no effect or no change. The alternative hypothesis (H1 or Ha) says there is. You never prove H0. You either reject it, because the evidence against it is strong enough, or fail to reject it, because it is not.
| Term | Meaning | Example |
|---|---|---|
| Null hypothesis (H0) | The default claim of no effect or no difference | The average delivery time is 3.0 days |
| Alternative hypothesis (H1) | What you are looking for evidence of | The average delivery time is not 3.0 days |
| Significance level (alpha) | The risk of wrongly rejecting a true H0 that you will accept, often 0.05 | 5 percent |
| Test statistic | A number that measures how far the sample is from H0 in standard error units | t = 2.00 |
| p-value | The probability of a result at least this extreme if H0 were true | p = 0.053 |
| Decision rule | Reject H0 if p is less than alpha, or if the test statistic is beyond the critical value | Reject if p < 0.05 |
Whether H1 is two-tailed (not equal) or one-tailed (greater than, or less than) depends on the question you were asked, and you must decide before looking at the data.
Type I and Type II errors
| H0 is actually true | H0 is actually false | |
|---|---|---|
| Reject H0 | Type I error (false positive), probability alpha | Correct decision, the probability is called power |
| Fail to reject H0 | Correct decision | Type II error (false negative), probability beta |
Setting alpha lower, say 0.01, makes Type I errors rarer but Type II errors more likely. Larger samples reduce both kinds of error. In business terms, a Type I error might mean launching a new product based on a result that was just luck, and a Type II error might mean abandoning a good idea because the test failed to detect its benefit. The costs of each should guide the alpha you choose.
The six steps
- State the hypotheses in words and symbols, and say whether the test is one-tailed or two-tailed.
- Choose the significance level (alpha), unless it is given.
- Choose the test that fits the data and the question, and check its assumptions.
- Calculate the test statistic and the p-value or critical value.
- Make the decision: reject or fail to reject H0.
- Write a conclusion in context, in plain language about the business question.
Always finish with step 6. A correct calculation without a business conclusion earns only part of the marks.
Choosing the right test
| Question | Data | Test | Key assumptions |
|---|---|---|---|
| Is a population mean equal to a value? (standard deviation known or large sample) | One numerical sample | One-sample z-test | Known standard deviation or large n |
| Is a population mean equal to a value? (standard deviation unknown) | One numerical sample | One-sample t-test | Roughly normal data, or large n |
| Do two independent groups have different means? | Numerical, two separate samples | Two-sample t-test (pooled or Welch) | Independent samples, roughly normal |
| Did a measure change for the same units, before and after? | Numerical, paired | Paired t-test | Differences roughly normal |
| Is a proportion equal to a value? | One categorical sample | One-sample z-test for a proportion | n x p and n x (1 - p) at least 10 |
| Do two proportions differ? | Two categorical samples | Two-proportion z-test | Large enough samples |
| Are two categorical variables related? | Counts in a table | Chi-square test of independence | Expected counts at least 5 |
| Do three or more means differ? | Numerical, several groups | One-way ANOVA | Independent groups, similar variances |
If the assumptions are doubtful, say so in your answer and state what you would use instead, for example a non-parametric test.
Worked example 1: one-sample t-test
Delivery time (hypothetical)
A company claims the average delivery time is 3.0 days. A sample of 36 orders has a mean of 3.4 days and a standard deviation of 1.2 days. Test at the 5 percent level whether the true mean differs from 3.0.
- Hypotheses: H0: mean = 3.0. H1: mean is not 3.0 (two-tailed).
- Test statistic: t = (3.4 - 3.0) / (1.2 / square root of 36) = 0.4 / 0.2 = 2.00, with 35 degrees of freedom.
- Critical value: for a two-tailed test at 0.05 with 35 degrees of freedom, about 2.03.
- Decision: 2.00 is just inside the critical value, so we fail to reject H0. The p-value is about 0.053, slightly above 0.05.
- Conclusion: the sample gives some evidence that deliveries are slower than claimed, but at the 5 percent level it is not strong enough to conclude that the true average differs from 3.0 days. A larger sample might settle it.
This example is deliberately borderline. A p-value of 0.053 does not mean the claim is true. It means the evidence is not conclusive at the chosen level.
Working on this assignment now? Get a price for help with your paper.
Get an instant quoteWorked example 2: two-sample t-test
Store layouts (hypothetical)
A retailer tests a new store layout. 40 customers visiting stores with the new layout spent a mean of $48.20 with a standard deviation of $6.00. 45 customers in stores with the old layout spent a mean of $45.00 with a standard deviation of $7.00. Is there a difference at 5 percent?
- Hypotheses: H0: the two means are equal. H1: they are not (two-tailed).
- Standard error: square root of (36 / 40 + 49 / 45) = square root of (0.900 + 1.089) = 1.410.
- Test statistic: t = (48.20 - 45.00) / 1.410 = 2.27, with about 83 degrees of freedom (Welch).
- Decision: the critical value is about 1.99 and the p-value is about 0.026. Because p is below 0.05, reject H0.
- Conclusion: average spending is higher with the new layout, and the difference of $3.20 is unlikely to be due to chance alone.
Whether the extra $3.20 matters commercially is a separate question, which leads to the next point.
Worked examples 3 and 4: a proportion and chi-square
Proportion test (hypothetical)
A manager believes 30 percent of customers prefer a new product. In a sample of 200, 74 prefer it (37 percent).
H0: p = 0.30. H1: p is not 0.30. Standard error = square root of (0.30 x 0.70 / 200) = 0.0324. z = (0.37 - 0.30) / 0.0324 = 2.16. The two-tailed p-value is about 0.031, below 0.05, so reject H0. The data suggest that the true share is not 30 percent, and the estimate is higher.
Chi-square test of independence (hypothetical)
Does purchase depend on channel? Of 200 online visitors, 60 bought and 140 did not. Of 200 store visitors, 90 bought and 110 did not.
Total buyers are 150 of 400. Expected buyers in each channel are 200 x 150 / 400 = 75, and expected non-buyers are 125. Chi-square = (60-75)^2/75 + (140-125)^2/125 + (90-75)^2/75 + (110-125)^2/125 = 3 + 1.8 + 3 + 1.8 = 9.6, with 1 degree of freedom. The critical value at 0.05 is 3.84 and the p-value is about 0.002, so reject H0.
Conclusion: purchase and channel are related. Stores convert 45 percent of visitors and online 30 percent in this sample.
Statistical significance is not practical significance
With a very large sample, tiny differences become statistically significant. A change in average order value of 8 cents might be significant with 500,000 orders and still be irrelevant to the business. In a business setting, always add a sentence on size and relevance, and consider a confidence interval, which gives a range of plausible values for the true difference. For the layout example, a 95 percent interval for the extra spend would be about $3.20 plus or minus 1.99 x 1.41, which is from about $0.40 to $6.00. Whether that is worth the cost of the new layout is the real decision.
A confidence interval and a two-tailed test at the same level agree: if the interval for a difference excludes zero, the test rejects H0.
Doing it in Excel
| Task | Excel approach |
|---|---|
| Two-tailed p-value for a t statistic | =T.DIST.2T(ABS(t), df) |
| Two-sample t-test directly from data | =T.TEST(range1, range2, 2, 3) for two tails with unequal variances |
| Normal tail areas for z tests | =NORM.S.DIST(z, TRUE) gives the area to the left of z |
| Chi-square p-value from observed and expected tables | =CHISQ.TEST(observed_range, expected_range) |
| All-in-one output tables | Data Analysis ToolPak: t-Test, z-Test, ANOVA, Descriptive Statistics |
Software gives the p-value, but you still need to state hypotheses, check assumptions and write the conclusion. Report results in the form your course uses, for example t(35) = 2.00, p = .053.
Common mistakes
- Saying accept H0 Say fail to reject. Not finding evidence is not proof that nothing is going on.
- Misreading the p-value It is not the probability that H0 is true. It is the probability of data this extreme if H0 were true.
- Choosing one tail after seeing the data Decide the direction from the question before you look at the results.
- Skipping assumptions State and check them, especially for small samples.
- Stopping at the number Always write a conclusion in context and say whether the difference matters.
- Confusing independent and paired samples Before and after on the same people is paired. Two separate groups are independent.
If you want help with a statistics problem set or a write-up, you can get a quote for statistics homework help.