Types of data
The type of data decides which summaries and charts make sense, so classify it first.
| Type | Description | Examples | Suitable summaries |
|---|---|---|---|
| Nominal | Categories with no order | Region, product type, payment method | Counts, percentages, mode |
| Ordinal | Ordered categories | Satisfaction ratings, job grade | Median, percentages by rank |
| Interval | Numbers with equal gaps but no true zero | Temperature in Celsius, some index scores | Mean, standard deviation |
| Ratio | Numbers with equal gaps and a true zero | Sales, cost, age, time | All summaries, including ratios and percent changes |
Data may also be a sample (a subset) or a population (everything), which affects formulas such as the standard deviation.
Measures of center and spread: a worked example
Weekly sales for nine stores (in thousands of dollars): 12, 15, 15, 18, 20, 22, 25, 30, 41.
| Measure | Calculation | Result |
|---|---|---|
| Mean | (12 + 15 + 15 + 18 + 20 + 22 + 25 + 30 + 41) / 9 = 198 / 9 | 22.0 |
| Median | The middle (fifth) value of the ordered list | 20 |
| Mode | The most frequent value | 15 |
| Range | 41 - 12 | 29 |
| Sample variance | Sum of squared deviations (652) divided by n - 1 (8) | 81.5 |
| Sample standard deviation | Square root of 81.5 | 9.03 |
| Coefficient of variation | Standard deviation divided by mean = 9.03 / 22 | 41 percent |
The deviations from the mean are -10, -7, -7, -4, -2, 0, 3, 8 and 19, and their squares sum to 652. For a population you would divide by 9, not 8, giving a variance of 72.4 and a standard deviation of 8.51. Use n - 1 for samples, because a sample tends to understate the spread of the population.
The mean (22.0) is above the median (20), which signals a right-skewed distribution: the high value of 41 pulls the mean up. When data are skewed or have outliers, the median is a better description of a typical value.
Quartiles, the IQR and outliers
Quartiles split ordered data into four parts. Q1 is the 25th percentile and Q3 the 75th. The interquartile range (IQR = Q3 - Q1) measures the spread of the middle half and is not affected by extremes.
For the nine values, the lower half (12, 15, 15, 18) gives Q1 = 15, and the upper half (22, 25, 30, 41) gives Q3 = 27.5. IQR = 12.5. A common rule flags values beyond 1.5 times the IQR outside the quartiles as outliers: the fences are 15 - 18.75 = -3.75 and 27.5 + 18.75 = 46.25. The value 41 is inside the fences, so it is high but not a formal outlier.
Different methods give different quartiles
Textbooks and software compute quartiles slightly differently. Excel's QUARTILE.INC gives Q3 = 25 for this data, while QUARTILE.EXC gives 27.5. State which method you use, and follow your course.
A box plot displays the minimum, Q1, median, Q3 and maximum, and marks outliers. It is the best quick picture of spread and skew.
Probability rules
| Rule | Formula | When to use |
|---|---|---|
| Complement | P(not A) = 1 - P(A) | When it is easier to find the opposite |
| Addition | P(A or B) = P(A) + P(B) - P(A and B) | Either event happening |
| Multiplication (independent) | P(A and B) = P(A) x P(B) | When one event does not affect the other |
| Conditional | P(A given B) = P(A and B) / P(B) | When one outcome is known |
| Multiplication (general) | P(A and B) = P(B) x P(A given B) | When events are dependent |
A contingency table (hypothetical)
200 customers: 120 shop online and 80 in store. Of the online customers, 60 are loyalty members, and of the store customers, 20 are.
- P(member) = 80 / 200 = 0.40. P(online) = 120 / 200 = 0.60.
- P(online and member) = 60 / 200 = 0.30.
- P(member given online) = 60 / 120 = 0.50.
- Independent? If they were, P(online and member) would be 0.60 x 0.40 = 0.24. It is 0.30, so shopping channel and membership are not independent: online customers are more likely to be members.
Working on this assignment now? Get a price for help with your paper.
Get an instant quoteBayes' rule and total probability
Which machine made the defect? (hypothetical)
Machine A makes 60 percent of output and 2 percent of its items are defective. Machine B makes 40 percent, with 5 percent defective. An item is chosen at random.
- Total probability of a defect: 0.60 x 0.02 + 0.40 x 0.05 = 0.012 + 0.020 = 0.032.
- Given it is defective, the probability it came from A: 0.012 / 0.032 = 0.375, or 37.5 percent.
Even though A makes most items, B is responsible for most defects (62.5 percent), because its defect rate is higher.
A tree diagram helps: draw branches for the machine, then for defective or not, and multiply along each branch. Sum the branches that end in a defect for the total, then divide the branch of interest by that total.
Expected value for decisions
The expected value of an outcome is the sum of each value times its probability. It gives the long-run average and is the basis for many business decisions.
A project decision (hypothetical)
A project has a 60 percent chance of earning $50,000 and a 40 percent chance of losing $20,000.
- Expected value: 0.60 x 50,000 + 0.40 x (-20,000) = 30,000 - 8,000 = $22,000.
- Standard deviation of outcomes: square root of [0.6 x (50,000 - 22,000)^2 + 0.4 x (-20,000 - 22,000)^2] = square root of 1,176,000,000, about $34,300.
The project is attractive on average, but the spread shows it is risky. Two projects with the same expected value can have very different risk, which is why both numbers matter.
Binomial and normal distributions
The binomial distribution counts successes in a fixed number of independent trials with a constant probability. The probability of exactly k successes in n trials is C(n, k) x p^k x (1 - p)^(n - k).
Binomial
Each customer buys with probability 0.4. What is the probability exactly 3 of 5 customers buy? C(5, 3) = 10, so the probability is 10 x 0.4^3 x 0.6^2 = 10 x 0.064 x 0.36 = 0.2304. The mean number of buyers is n x p = 2 and the standard deviation is the square root of n x p x (1 - p), about 1.10.
The normal distribution is the bell-shaped curve used to model many measurements. Convert a value to a z-score, z = (x - mean) / standard deviation, then use a table or software to find the area.
Normal
Delivery times have a mean of 3.0 days and a standard deviation of 0.5 days. What proportion take longer than 3.8 days? z = (3.8 - 3.0) / 0.5 = 1.6, and the area above z = 1.6 is about 0.055, so roughly 5.5 percent of deliveries exceed 3.8 days. For the slowest 5 percent, the cutoff is mean plus 1.645 standard deviations: 3.0 + 1.645 x 0.5 = 3.82 days.
In Excel, =NORM.S.DIST(1.6,TRUE) returns the area to the left, and =BINOM.DIST(3,5,0.4,FALSE) returns 0.2304. See our guide to hypothesis testing for how these ideas lead to tests.
Common mistakes
- Using the mean for skewed data Use the median, or report both and comment on the skew.
- Dividing by n instead of n - 1 Use n - 1 for sample variance and standard deviation.
- Treating dependent events as independent Multiply probabilities directly only when one event does not change the other.
- Mixing up P(A given B) and P(B given A) These are different, as the defect example shows.
- Forgetting to state the units Give units for means and standard deviations.
- Unlabeled charts Every chart needs a title, labeled axes and units.
If you want help with a statistics problem set, you can order statistics homework help.