Explore data, compare distributions, calculate normal probabilities, and fit regression lines, with worked examples and practice.
From questions to data
How long does a trip to school usually take? How much do travel times vary? Does leaving earlier help? Statistics gives us tools to organize observations, describe patterns, and assess claims.
Statistics: describing data and drawing conclusions
Statistics is the field concerned with collecting, classifying, organizing, summarizing, analyzing, and interpreting data. Descriptive statistics organizes, summarizes, and describes the observed data. Inferential statistics uses a sample to draw conclusions about characteristics of its population, with uncertainty.
Units, variables, and data values
A variable is a characteristic that can vary across individuals or objects, or over time. An experimental unit is the individual or object on which a variable is measured. A data value is the result of measuring or recording that variable on one unit.
Population and sample
A population is the collection of all measurements of interest; a sample is a selected subcollection of those measurements.
Example · Heights in a school
Suppose we want to study the heights of all Grade 11 students in a school at the start of the school year. The population consists of all their height measurements. If we select 30 students and measure their heights, those 30 measurements form a sample.
Repeated measurements still count
If two students both measure 160 cm, keep both values: each measurement belongs to a different student.
Parameter and statistic
A parameter is a numerical characteristic of the population; a statistic is a numerical summary calculated from the sample.
Summarizing the measurements in a sample is descriptive. Using that summary to estimate a population characteristic is inferential. The commute study below shows the distinction.
Example · A commute study
To describe today’s journeys among students who travelled to school, you ask 40 of those students to record their travel times in minutes.
Experimental unit: one student.
Variable: travel time, in minutes.
Sample: the 40 recorded travel-time measurements.
Population of measurements: today’s travel times for all students who travelled to this school. Those students are the population of units.
The sample mean describes those 40 trips. Generalizing to the entire school depends on how the sample was selected.
How many characteristics do we record for each unit? This determines whether we study one variable or relationships among several.
Number of variables
Univariate data record one variable per unit; bivariate data record two. Here, multivariate data record more than two variables per unit.
Count variables, not people
Five students with one height each give univariate data. One hundred students with height and travel time each give bivariate data.
How many variables?
Example
Useful question
Univariate
Travel time for each student
What is a typical travel time?
Bivariate
Departure time and travel time
Are earlier departures associated with shorter trips?
Multivariate
Departure time, distance, transport mode, travel time
How do several characteristics relate to travel time?
Example · Read a student record
To explore students’ academic profiles, a researcher records five characteristics for each of five selected undergraduates. Each row belongs to one student; the first column identifies that student rather than adding a measured characteristic. Read across a row to connect one unit to its data values, then down a column to follow one variable.
Student ID
GPA
Gender (recorded code)
Year
Major
Enrolled units
1
2.0
F
First
Psychology
16
2
2.3
F
Second
Mathematics
15
3
2.9
M
Second
English
17
4
2.7
M
First
English
15
5
2.6
F
Third
Business
14
For student 3, the student is the experimental unit, GPA is the variable, and 2.9 is its observed data value. The full five-variable record is multivariate. Studying GPA alone gives univariate data: the five GPAs form the sample, while all undergraduates’ GPAs at the specified university and time form the population of measurements.
Try it · Units, variables, and an inference
Using the student-record table, identify the sample for enrolled units, count its observations, and classify a study using only GPA and enrolled units. Does the mean GPA of these five selected students necessarily equal the university’s mean GPA?
Show solution
The sample is 16, 15, 17, 15, 14 enrolled-unit measurements. Its size is five; the two 15s belong to different students and both count. GPA together with enrolled units gives two variables per student, so the data are bivariate. The sample mean GPA is (2.0 + 2.3 + 2.9 + 2.7 + 2.6)/5 = 2.5, a statistic. It need not equal the population mean, a parameter. The selection method is unspecified, so representativeness cannot be assumed.
Think about the collection process
A large sample can still be biased. Surveying only students who drive leaves out other transport modes. Check who was included, who was missed, the units, and how the question was asked.
Try it · Identify the claim
A survey of 60 volunteers reports a median study time of 90 minutes. A student concludes that every student in the school studies for 90 minutes. What is wrong?
Show solution
The median summarizes the volunteers; it does not describe every individual. Volunteers may not represent the school. Even a representative sample would support an estimate with uncertainty, not an exact claim about everyone.
Variables & measurement
Categorical and numerical variables
A qualitative (categorical) variable records a quality or category on each unit. A quantitative (numerical) variable records an amount.
Discrete and continuous variables
A discrete quantitative variable has a finite or countably infinite set of possible values. A continuous quantitative variable is modeled as capable of taking any real value in an interval.
Possible values versus observed values
Classify a variable by its possible values, not just the finite list collected in one study.
Example · What do the values represent?
A number used as an identifier, such as a student number, is still a label. Number of siblings is discrete; elapsed travel time is modeled as continuous. Rounding a time to the nearest minute makes its recorded values discrete without changing the underlying continuous quantity. Discrete values need not be integers: a score restricted to half-points is discrete.
Example · A proportion can be discrete
A recovery rate must be defined carefully before classifying it. For a fixed cohort of N patients, the observed proportion recovered is R/N with R = 0, 1, …, N, so it has only N + 1 possible values and is discrete. A continuous rate model requires an additional modeling assumption. Likewise, QPI/GPA is quantitative, but its discrete or continuous treatment depends on the permitted scores and recording precision.
Level
What comparisons mean
Example
Nominal
Same or different category; no natural ordering
Transport mode: walk, bus, train
Ordinal
An order exists; gaps are not necessarily equal
Satisfaction: low, medium, high
Interval
Equal differences are meaningful; zero is conventional
Temperature in °C
Ratio
Equal differences and ratios are meaningful; zero represents none
Elapsed time, distance, number of books
Common mistake · Numbers do not determine the measurement level
Twice a Celsius temperature does not mean twice as hot. Twice an elapsed time does mean twice the duration. Numeric codes for categories do not make their averages meaningful.
Try it · Classify the variable
Jersey number
Finishing rank in a race
Temperature in °C
Number of siblings
Time taken to run 100 metres
Show solution
Categorical, nominal: a label.
Categorical, ordinal: ordered ranks.
Numerical, interval; treated as continuous.
Numerical, discrete, ratio.
Numerical, continuous, ratio.
Classification depends on what the recorded value represents, rather than its appearance.
Before choosing an average or a graph, we need to know what the recorded values mean. A transport label, a finishing rank, and a journey time allow different comparisons. We will classify the variable and its measurement level so that our later calculations answer a meaningful question.
Mean, median & mode
Three ways to describe a typical value
Mean: add all observations and divide by the number of observations, x̄=Σxᵢn.
Median: sort the observations. For odd n, choose the middle value. For even n, average the two middle values.
Mode: the most frequent value or category.
Mode convention
Several values can tie for the highest frequency. Here, a dataset with no repeated values is described as having no mode.
Example · Start by sorting
Seven students score 60, 60, 70, 80, 60, 70, 50 on a 100-point assessment. A teacher wants a representative score for this group. Compare the mean, median, and mode.
Sorted: 50, 60, 60, 60, 70, 70, 80.
Mean: 450/7 ≈ 64.29. Median: the 4th value, 60. Mode: 60, appearing three times.
The mean balances all seven scores, the median locates the middle observation, and the mode identifies the most common score. They answer different questions, so they need not agree.
Which average best represents the data?
It depends on the question. Are we describing the middle observation, finding the most common value, or using a total to plan resources? The mean uses every observation, but unusually large or small values can pull it away from most of the data. Compare these two situations before choosing a summary.
Example · Which average describes the number of children?
Five families have 2,4,3,5,7 children. We want to describe how many children a family has, on average. Which average should we use?
The mean is (2+4+3+5+7)/5=21/5=4.2 children per family. It uses all five counts and expresses the total of 21 children per five families.
In order, the counts are 2,3,4,5,7. The median is 4 children: this is the middle count, with two counts below it and two above it. There is no mode, since every count occurs once.
Both the mean and median are reasonable summaries here. The values have no strongly isolated extreme, and the two summaries are close. Use 4.2 to express the total number of children per family, or 4 to describe the middle of the ordered counts. The general question alone does not make one uniquely best.
A fractional mean is not a mistake
No individual family can have 4.2 children, but a mean need not be a possible individual observation. A decimal result alone does not make the mean inappropriate. In the salary example below, the two highest values create a much larger difference between the mean and median.
Example · What does a typical staff member earn?
Ten staff members earn monthly salaries of ₱15,000, ₱18,000, ₱16,000, ₱14,000, ₱15,000, ₱15,000, ₱12,000, ₱17,000, ₱90,000, and ₱95,000. A job applicant wants to understand what staff usually earn.
Measure
Monthly salary
What it describes
Mean
₱30,700
Total payroll divided by 10 staff members
Median
₱15,500
The midpoint of the two middle salaries
Mode
₱15,000
The most common salary
Sorted in thousands of pesos, the salaries are 12,14,15,15,15,16,17,18,90,95. The median is (15+16)/2=15.5, and 15 occurs three times. The total is ₱307,000, giving a mean of ₱30,700.
Eight of the ten salaries are ₱18,000 or less. The two highest salaries pull the mean upward. The median is suitable for describing the middle salary; the mode answers which salary is most common. The mean remains useful for payroll: 10 × 30,700 = 307,000 pesos.
Another approach · Report the separated groups
To make the large gap visible, we can report the eight lower salaries and the two highest salaries separately, while also giving the overall summary:
Eight staff members earn ₱12,000–₱18,000, with a mean of ₱15,250. The other two earn ₱90,000 and ₱95,000. Across all ten staff members, the mean is ₱30,700.
The eight lower salaries total ₱122,000, so their mean is 122,000/8=15,250. This is the mean of that subgroup, not the mean of all staff.
Keep the full picture
Do not silently discard unusually high or low observations, or assume they are errors. State which values were separated and why. If there are genuinely different job groups, reporting those groups separately may be more informative; salaries alone do not establish those roles. A formal rule for flagging potential outliers appears with box plots later in this lesson.
Explore · What happens when one value changes?
Edit the dataset, compare presets, or add an extreme value. Predict which summaries will change the most.
See the sorted data and box plot
Changes stay in this page until you leave or reload it. Sample calculations need at least two values.
Try it · Even-sized data
Find the mean, median, and mode(s) of 101,101,106,108,108,106.
Show solution
Sorted: 101,101,106,106,108,108. Mean: 630/6=105. Median: (106+106)/2=106. All three values 101, 106, and 108 share the highest frequency of 2, so they are tied modes.
We have decided what kind of data we have. Now imagine explaining a whole list of observations to someone who has not seen it: which value would you call typical? A very large observation pulls the mean toward it, while the median depends on position. Keep that contrast in mind as we compare the three summaries.
Combine summaries without losing the counts
A pooled mean summarizes all observations from disjoint groups. If their sizes are \(n_1,n_2\) and means are \(\bar x_1,\bar x_2\), then \(\bar x=(n_1\bar x_1+n_2\bar x_2)/(n_1+n_2)\). The group sizes are weights because each individual observation must count once.
Compare two commute groups
Ten walkers have a mean travel time of \(12\) minutes; thirty bus riders have a mean of \(28\) minutes. Find the overall mean and explain why \((12+28)/2\) is misleading.
Show worked solution
The group totals are \(120\) and \(840\) minutes. Thus \(\bar x=960/40=24\) minutes. The unweighted result \(20\) gives ten walkers as much influence as thirty riders. The correct mean lies between the two means and closer to the larger group's mean, which is a useful reasonableness check.
Combining means needs counts. Combining variances needs still more information: groups may also differ in their centers. Do not average summary statistics without checking what they summarize.
Predict the new mean
Five students have a mean score of 12. Every score increases by 3. What is the new mean?
Variance & standard deviation
Two groups can have the same mean and very different consistency. Compare {35,35,35,35,35} with {10,25,35,45,60}: both have mean 35, but only one has no spread.
Measuring variation
Range: maximum − minimum.
Population variance:σ²=Σ(xᵢ−μ)²N. Population SD:σ=√σ².
Sample variance:s²=Σ(xᵢ−x̄)²n−1. Sample SD:s=√s².
The sample formula requires n > 1; the population formula requires N > 0.
Why square the deviations?
Positive and negative deviations from the mean cancel when added. Squaring prevents this cancellation. Taking the square root returns the result to the original units; standard deviation is not the average absolute distance.
Example · Calories in ten burgers
A survey of ten fast-food restaurants recorded the calorie content of a midsized hamburger at each restaurant. Treat the ten calorie measurements as a sample and calculate the mean and sample standard deviation, rounding to two decimal places.
514, 507, 502, 498, 496, 506, 458, 478, 463, 514
Sum the values: 4,936. Mean: x̄=493.6 calories.
Subtract 493.6 from each value, square the differences, and add them: Σ(xᵢ−x̄)²=3916.4.
For this sample, divide by 9: s²≈435.16 calories².
Take the square root: s≈20.86 calories.
The range is 514−458=56 calories. Use the Burger calories preset above to compare the sample and population formulas.
Common mistake · Keep the units straight
Variance has squared units. Standard deviation has the original units. Use n−1 for a sample variance used to estimate population variation, and N when describing the complete population. Do not round intermediate deviations.
Try it · Same center, different spread
Find the sample standard deviation of {25,35,35,35,45}. Compare it with {35,35,35,35,35}.
Show solution
The first mean is 35. Squared deviations total 100+0+0+0+100=200. Thus s²=200/4=50 and s=√50≈7.07. The second dataset has all deviations equal to 0, so its standard deviation is 0.
A typical value is a useful beginning, but it does not tell the whole story. Two classes can have the same average and very different results. We will now measure how far the observations tend to lie from their center, so that our summary describes consistency as well as level.
Z-scores & percentiles
Position relative to the group
A z-score expresses a value in standard deviation units:
z=x−μσ for population summaries, or z=x−x̄s for sample summaries.
The standard deviation must be positive.
For this lesson, percentile rank uses the strictly-below convention: number of observations below xtotal number of observations × 100%.
Interpreting position
A positive z-score is above the mean and a negative one is below. Percentile rank instead counts observations below a value. Other percentile conventions handle ties differently; use the stated method.
Example · Ages compared with the mean
An employer summarizes staff ages with a mean of 36 years and standard deviation of 8.3 years. Compare Tammy (23) and Joy (44): who is farther from the mean in standard deviation units?
Tammy, age 23: z=(23−36)/8.3≈−1.57. Joy, age 44: z=(44−36)/8.3≈0.96.
Tammy is farther from the mean in standard deviation units. These two numbers alone do not tell us how common either age is; that requires the distribution’s shape.
Try it · Forward and backward z-scores
A group receives a mean of 43 pieces of candy, with SD 2. Find the z-score for 20 pieces. If Zia’s z-score is 6, how many pieces did she receive?
Show solution
z=(20−43)/2=−11.5. To recover an observation, use x=μ+zσ. Zia received 43+6(2)=55 pieces. These are algebraic results; no normal distribution assumption is needed to compute them.
Try it · Interpret percentile rank
Brad scored higher than 7,255 of 9,430 students. Find his percentile rank to the nearest integer.
Harriet is at the 45th percentile among 2,480 students by height. Approximately how many are shorter?
Show solution
(7255/9430)×100≈76.94%: about the 77th percentile.
0.45×2480=1116 students, using the stated percentile convention.
A percentile rank of 77 does not mean a test score of 77%.
We can describe a center and a spread, but how should we compare observations from different distributions? A standardized score measures position relative to a mean in units of standard deviation. A percentile answers a different question about relative rank. Keep the measurement of distance separate from the proportion of observations below a value.
Quartiles, IQR & box plots
Find quartiles by splitting the sorted data
Sort the observations.
The median is Q₂.
Split the observations into a lower and upper half. If n is odd, exclude the overall median from both halves.
The lower-half median is Q₁; the upper-half median is Q₃.
This is the median-of-halves convention used throughout this lesson. Other software may interpolate differently.
Interquartile range
The interquartile range is IQR=Q₃−Q₁.
This measures the distance between the lower and upper quartiles. Unlike the range, it uses the middle portion of the ordered data rather than the two extremes.
Identify potential outliers
Lower fence: Q₁−1.5(IQR). Upper fence: Q₃+1.5(IQR). Values strictly outside these fences are potential outliers.
Construct the modified box plot
Draw a box from Q₁ to Q₃, a line at the median, and whiskers to the most extreme observed values at or within the fences. Plot potential outliers as separate points.
Example · A long commute
Travel times (minutes): 10,12,14,16,18,20,22,24,60.
IQR=10. Fences: −2 and 38. The value 60 is an outlier; whiskers end at 10 and 24.
An outlier is a question, not an automatic deletion
Check for a recording error, a different process, or an unusual but valid observation. Explain any exclusions. A fence is a calculated cutoff; a whisker ends at an actual observed value.
Try it · Build a five-number summary
Eight students record their waiting times for a school shuttle, in minutes: 2,4,5,7,8,10,12,14. Summarize the middle half of the waits and check for unusually long waits. Using the median-of-halves convention, find the minimum, Q₁, median, Q₃, maximum, and any potential outliers.
Show solution
Five-number summary: (2,4.5,7.5,11,14). IQR=6.5 minutes. Fences: −5.25 and 20.75 minutes. There are no potential outliers under this rule, so the whiskers reach 2 and 14 minutes. The box spans 4.5 to 11 minutes; it describes the middle half of these recorded waits, not a guarantee for future journeys.
What a box plot shows
Quartiles are cut points in ordered numerical data: the first quartile marks approximately the lowest quarter, the second is the median, and the third marks approximately the lowest three quarters. For finite data, their numerical values depend on the stated quartile convention.
Reading a box plot
A box plot displays these quartiles and two whisker endpoints on a common numerical scale. The box spans the first to third quartile, representing the middle half of the ordered data approximately; a line marks the median.
A longer box means greater spread in that middle portion, not more observations. Comparing plots on the same scale helps compare typical values, variability, and unusual observations between groups.
Read the whiskers before interpreting extremes
This lesson uses a modified box plot: whiskers end at the most extreme observed values within the outlier fences, and values beyond them appear separately.
Some box plots instead extend whiskers to the minimum and maximum, so check the convention. A box plot does not show every observation or reveal multiple clusters as clearly as a dot plot or histogram.
The mean and standard deviation describe center and spread, but an unusually long commute can pull them away from most students’ experience. We now want a summary based on positions in the ordered data: where is the middle, and where does the middle half lie? Quartiles mark those positions; the interquartile range measures the width of the middle half, and a box plot displays them alongside the tails and possible outliers. We will sort the observations, find the quartiles using one stated convention, and then turn that summary into a picture.
Frequency tables & histograms
Build a histogram in four stages: raw data → ordered array → frequency distribution table → histogram.
1 · Start with the raw dataset
These are the ten-year population growth rates (%) of 50 Philippine cities and municipalities from 2000 to 2010. They are historical teaching data.
Determine the number of classes. Choose the smallest integer \(k\) for which \(2^k>n\). Since \(2^5=32\le50\) and \(2^6=64>50\), use \(k=6\).
Determine the class width. Round up: \[w=\left\lceil\frac{H-L}{k}\right\rceil=\left\lceil\frac{54.3-(-3.2)}{6}\right\rceil=\lceil9.583\ldots\rceil=10.\]
Select class limits. Start at −4, a convenient value slightly below −3.2. Successive lower limits differ by 10: −4, 6, 16, 26, 36, 46. The corresponding upper limits are 5, 15, 25, 35, 45, 55.
Find boundaries and count. Place each boundary halfway between adjacent class limits: \((5+6)/2=5.5\). LCB means lower class boundary; UCB means upper class boundary. Count values using \(\mathrm{LCB}\le x<\mathrm{UCB}\). A value on a shared boundary belongs to the class for which it is the lower boundary.
Class labels such as “−4 to 5” name the class limits; the boundaries determine membership. For example, 5.3 belongs in the first class because \(-4.5\le5.3<5.5\). Do not round the observations.
Frequency distribution of the 50 growth rates
Class (%)
LCB
UCB
Frequency
-4.5
5.5
5
5.5
15.5
14
15.5
25.5
19
25.5
35.5
7
35.5
45.5
3
45.5
55.5
2
Total
50
Select a class to inspect the values counted in it.
4 · Transform the table into a histogram
The horizontal axis shows classes, with bar edges at their boundaries. The vertical axis shows frequencies. Each rectangle has width 10 and a height equal to its class frequency. Adjacent rectangles touch.
Point to, tap, or focus a bar or class button to connect the histogram with its frequency table. The numbers remain unchanged.
Try it · Check the boundary rule
Which class would contain 15.5? How many of the actual observations are in the class 16 to 25, and what percentage of the data is that?
Show solution
15.5 belongs to the class 16 to 25 because it is that class’s lower boundary. There are 19 actual observations in this class: \(19/50=38\%\).
A histogram displays a frequency distribution
The frequency of a category or class interval is the number of observations assigned to it. A frequency distribution records those frequencies across the categories or intervals.
Frequency distribution and histogram
A histogram represents a numerical frequency distribution by adjacent rectangles over class intervals.
A histogram places numerical intervals along a number line and draws adjoining bars to reveal where values cluster, how widely they spread, and whether there are gaps or long tails. Here the intervals have equal widths, so their bar heights can directly represent counts.
A long list makes it difficult to see where observations cluster. Grouping values into intervals gives us a frequency table; drawing those counts produces a histogram. We will keep the same dataset throughout, so you can trace how each observation contributes to a row and then to a bar.
When histogram widths differ
A frequency density is a class frequency divided by its width.
In a frequency-density histogram, rectangle area represents frequency. The equal-width histogram above can use frequency as height because every bar has the same width. Unequal widths require this extra care.
Compare unequal travel-time intervals
A survey has \(8\) travel times in \([0,10)\) minutes and \(12\) in \([10,30)\). Which interval has more observations, and which has the higher concentration per minute?
Show worked solution
Interval (minutes)
Frequency
Width
Frequency density
[0,10)
8
10
0.8
[10,30)
12
20
0.6
The second interval has more observations, but they are spread across twice the width. The first has higher density: \(8/10=0.8\) versus \(12/20=0.6\) observations per minute. Bar areas are \(10(0.8)=8\) and \(20(0.6)=12\). A taller bar does not necessarily contain more observations when widths differ.
When reading a histogram, check the vertical-axis label and the class widths together. Equal widths let you compare counts by height; unequal widths require comparing areas.
Predict the height when a class widens
Two classes each contain 12 observations. The second class is twice as wide. In a frequency-density histogram, how does its height compare?
Choose & question a graph
Look at the graph before revealing the explanation. Check which impression comes from the data and which comes from the way it is drawn.
1 · Omitting the baseline
How much larger does Group A look than the other groups?
Inspect this versionReveal the issue and compare
The bars begin at 50%, so their visible lengths are 12, 4, and 4. Group A looks three times as large, even though 62% is only about 1.15 times 54%. The gap is 8 percentage points.
More appropriate comparison
For a magnitude comparison using bar lengths, a zero baseline preserves the ratio. Changing the axis does not change 62%, 54%, or 54%.
2 · Compressing the vertical scale
Does this series look almost flat? Read the values as well as the slope.
Inspect this versionReveal the issue and compare
A 0–40 scale compresses the differences. A 0–15 scale uses the plotting area more effectively. Both show the same five observations; a flatter line does not mean there was no change.
More appropriate comparison
From 2019 to 2020 the value doubles from 6 to 12 in both versions. These values are a teaching reconstruction of the plotted points. There is no single mandatory range for a line chart: show clear ticks and enough context for the question.
3 · Selecting a favorable time period
If you saw only the first two years, would you expect the growth to continue?
Inspect this versionReveal the issue and compare
A selected interval can emphasize an increase while hiding a later reversal. Here the first two points are taken from the same complete five-year series.
More appropriate comparison
The wider view shows a peak followed by a decline. The data are an illustrative reconstruction of the five-year pattern; the selected view is a true subset, so no values change between the two graphs.
4 · Using the wrong graph
Can these three percentages be slices of one whole?
Inspect this versionReveal the issue and compare
The labels total 152%, not 100%. A pie chart falsely suggests mutually exclusive parts of a single whole. The original labels cannot be valid pie-slice percentages.
More appropriate comparison
A bar chart compares the three percentages without claiming they partition one whole. Keep the original percentages; silently normalizing them to 100% would change their meaning.
5 · Reversing a visual convention
Which region attracts your attention first? Does darker shading mean a higher density?
Inspect this versionReveal the issue and compare
The darkest areas represent the lowest densities in the first view. Readers often expect darker shading to mean more. A correct legend helps, but a reversed convention can still invite a quick misreading.
More appropriate comparison
Both schematic maps use the same six values. Only the shading direction changes. These simplified regions demonstrate the density-map concept; they are not geographic or current population estimates. Always read the legend and units.
A graph should help us understand the data, but its design can also create a misleading impression. We will compare displays of the same information and ask what the axes, scale, time window, and shading encourage us to believe. Read the actual values before accepting that first impression.
Normal distributions
Normal distribution
A normal distribution is a continuous probability distribution with density
The parameters μ and σ are its mean and standard deviation.
This formula produces a symmetric bell-shaped curve centered at μ; its mean, median, and mode coincide. Larger σ spreads the curve more widely. You do not need to integrate the formula by hand to use the probability explorer below.
Read area, not height
The total area under the density curve is 1. Area over an interval represents probability; height at one point does not. For a continuous normal variable, the probability of any one exact value is zero.
The empirical rule
Interval
Approximate proportion
μ−σ to μ+σ
68%
μ−2σ to μ+2σ
95%
μ−3σ to μ+3σ
99.7%
Example · Tomato weights
Assume tomato weights are normally distributed with μ=0.61 lb and σ=0.15 lb.
Below 0.76 lb is below μ+σ: approximately 50%+34%=84%.
Above 0.31 lb is above μ−2σ: approximately 97.5%. Out of 6,000 tomatoes, expect about 6000(0.975)=5850.
Between 0.31 and 0.91 lb is within two SDs: approximately 95%. Out of 4,500, expect about 4500(0.95)=4275.
These are empirical-rule approximations. A normal CDF gives slightly different, more precise values.
Check the model before using the rule
The 68–95–99.7 rule is for approximately normal data. Strong skew, multiple peaks, or extreme outliers can make it unsuitable. A symmetric appearance alone does not prove that a distribution is normal.
Try it · A grade interval
Assume grades are normal with mean 81 and SD 5. Approximately what proportion lies between 86 and 91? In a class of 40, what is the expected number in that interval?
Show solution
The bounds are μ+σ and μ+2σ. Using the empirical rule, the interval contains (95%−68%)/2=13.5%. Expected count: 40(0.135)=5.4, or about 5 students. An expected count is an average prediction, not a guaranteed whole-number outcome.
So far we have described the data in front of us. A probability model lets us ask what values might be expected in similar situations. The normal curve is one such model; we should check whether its shape is a sensible description before using its areas as probabilities.
Standard normal probabilities
Standardize, then find an area
If X is normal with mean μ and SD σ, then Z=(X−μ)/σ has mean 0 and SD 1. Let Φ(z)=P(Z≤z).
Below a: Φ((a−μ)/σ).
Above a: 1−Φ((a−μ)/σ).
Between a and b: Φ((b−μ)/σ)−Φ((a−μ)/σ).
Explore · Shade a probability
Change the mean, standard deviation, or bounds. The area updates after you choose Calculate.
Numerical CDF approximation; calculations use unrounded z-scores. The plot shows ±4 standard deviations, while probabilities include the full tails.
Example · Reading a standard normal table
For a left-tail table, Φ(1.50)≈0.9332. The area above 1.50 is 1−0.9332=0.0668. By symmetry, Φ(−1.50)≈0.0668.
Some tables report the area between 0 and z instead. Read the table heading before using its numbers.
Try it · Speeds and scores
Speeds are normal with μ=85 km/h and σ=10 km/h. Find the probability of a speed above 100 km/h.
Test scores are normal with μ=38 and σ=6. Find the proportion between 30 and 45.
With μ=81 and σ=5, predict how many of 40 students score at least 92.
Show solution
z=1.5, so P(X>100)≈0.0668=6.68%.
Bounds: z₁=−1.3333…, z₂=1.1667…. Probability ≈0.7871=78.71% using unrounded values. A table with rounded z-scores can differ slightly.
z=2.2, upper-tail probability ≈0.0139. Expected count 40(0.0139)≈0.56, roughly 1 student; this is not a guarantee.
The curve is a model; the area beneath a specified part of it gives the probability we want. Sketch the region before consulting a table or calculator. Decide whether you need a left tail, a right tail, or an interval, so that the number you obtain answers the original question.
Regression & residuals
A line for prediction
Linear regression models the relationship between an explanatory variable x and a response y using ŷ=ax+b. The hat marks a predicted value.
A residual is e=y−ŷ: observed minus predicted. Least squares chooses a and b to minimize Σe².
A positive residual means the observed point lies above the fitted line; a negative residual means it lies below. Squaring these vertical errors prevents positive and negative errors from cancelling.
The least-squares formulas
a=Σ(xᵢ−x̄)(yᵢ−ȳ)Σ(xᵢ−x̄)² and b=ȳ−ax̄.
At least two distinct x-values are needed. Always look at the scatter plot before treating a line as a useful model.
Explore · Fit a line and inspect its errors
The first dataset records machine hours (x) and sweets produced (y). Each line below contains one observed pair.
x: machine hours; y: sweets produced.
Inspect predicted values and residuals
Interpolation and extrapolation
Interpolation predicts within the observed x-range. Extrapolation predicts outside it, where the relationship may change. An intercept at x=0 is not automatically meaningful if zero is far outside the observed range.
Try it · Interpret a residual
A line predicts 284 sweets for a particular run. The observed output is 281. Find and interpret the residual.
Show solution
e=281−284=−3. The model overpredicted by 3 sweets. The point lies below the fitted line. Its contribution to the sum of squared errors is 9 sweets².
A prediction and its residual answer different questions
A regression prediction is the output of a model for a specified input; it is not a replacement for the measured observation. A residual is the observed value minus that prediction.
Its sign identifies which side of the fitted line the observation lies on, and its size measures the error in the response variable’s units.
We have mainly studied one variable at a time. Now suppose we want to understand how two measurements move together. A fitted line summarizes that relationship, while the residuals show what the line has failed to explain. Both deserve our attention.
Correlation & causation
Pearson’s linear correlation coefficient
For paired observations with neither variable constant, the Pearson correlation coefficient is r=Σ(xᵢ−x̄)(yᵢ−ȳ)√[Σ(xᵢ−x̄)² · Σ(yᵢ−ȳ)²].
Interpret the coefficient
−1≤r≤1. Its sign gives the direction of the linear association.
A magnitude close to 1 indicates a strong linear association.
r=0 means no linear association; a strong curved relationship may still exist.
r has no units. A constant variable makes the denominator zero, so the coefficient is undefined.
Try the Curved relationship preset in the regression explorer. It produces y=x² at symmetric inputs: the relationship is exact, yet the linear correlation is zero.
Example · Association without a causal conclusion
Ice cream sales and swimming incidents may both rise during hotter months. Their association alone does not show that buying ice cream causes an incident. Temperature, season, and the number of people swimming are possible common influences.
Likewise, a fitted line for family size and aggregate poverty incidence describes an association in that dataset. It cannot establish what changing one family’s size would cause.
What r² can tell you
For ordinary least-squares regression with an intercept, r² is the proportion of observed variation in y accounted for by the fitted linear relationship. It is not prediction accuracy, a causal percentage, or proof that the model will generalize.
Try it · Evaluate the conclusions
“r=−0.9 is weaker than r=0.4 because it is negative.”
“r=0 means the variables are unrelated.”
“A strong correlation proves that x causes y.”
Show solution
False. Compare magnitudes: 0.9 is stronger than 0.4. The minus sign indicates direction.
False. The relationship could be nonlinear.
False. Confounding, reverse causation, selection effects, and coincidence remain possible explanations.
The fitted line and correlation describe an association, but they do not tell us what caused it. Suppose two measurements rise together: could a third factor influence both? Before turning a numerical pattern into an explanation, ask how the data were collected and what other explanations remain possible.
Challenge: a conclusion can reverse when groups are combined
An aggregate rate combines subgroup rates using their sizes as weights. Changing the mixture can reverse a comparison, even when each subgroup comparison points the same way. This is called Simpson’s paradox. It is a reason to investigate how data were collected, not permission to select whichever comparison supports a preferred claim.
Audit a tutoring headline
Two programs report passes in a foundation test. Among beginners, A has 8 passes out of 10 and B has 63 out of 90. Among experienced students, A has 81 out of 90 and B has 10 out of 10. Compare overall rates and each starting level. Does either overall rate establish which program is better?
Show worked solution
Overall A has \(89/100=89\%\) and B has \(73/100=73\%\). Within beginners, A is \(80\%\) and B is \(70\%\); within experienced students, A is \(90\%\) and B is \(100\%\). Here the subgroup rankings conflict, so neither program dominates both groups. The overall figure mainly mixes different starting populations and cannot establish a causal effect.
Honors · Construct an actual reversal
Now A has 7/10 beginner passes and 81/90 experienced passes. B has 72/90 beginner passes and 10/10 experienced passes. Prove that B has higher rates at both levels but A has a higher overall rate. Explain the source of the reversal.
Show worked solution
For beginners, \(80\%>70\%\); for experienced students, \(100\%>90\%\): B wins each comparison. Yet A has \(88/100=88\%\) overall and B \(82/100=82\%\). A serves mostly experienced students, while B serves mostly beginners. The overall percentages use different weights; they are not a like-for-like comparison. Without a defensible study design, neither comparison by itself proves causation.
Practice & review
Build it yourself · From your data to two graphs
Enter measurements from a small study—such as travel times in minutes. You will calculate a frequency distribution table (FDT), then a histogram, and finally a modified box plot. Each correct answer unlocks the next step. The graphs appear only after you have checked their required calculations.
Write down your reasoning before opening each solution. Choose a graph as well as a numerical summary when the question calls for a description of data.
Review 1 · Describe a dataset
Daily reading times are 10,15,15,20,25,35. Find the mean, median, mode, range, sample variance, and sample SD.
Replace 35 by 95 in the previous dataset. Find the new mean and median. Which is more stable?
Show solution
New data: 10,15,15,20,25,95. Mean=180/6=30; median=(15+20)/2=17.5. The mean increased by 10 minutes while the median stayed unchanged.
Review 3 · Read the position
A student scores 84 on a test with mean 72 and SD 8. Find the z-score. Can you determine an exact percentile from these facts alone?
Show solution
z=(84−72)/8=1.5. The score is 1.5 SD above the mean. An exact percentile requires the distribution. Under a normal model, the percentile would be about 93.32.
Review 4 · Quartiles and outliers
Use the median-of-halves convention for 1,2,3,4,5,6,7,20. Find Q₁, median, Q₃, IQR, and the whisker endpoints.
Show solution
Q₁=2.5; median=4.5; Q₃=6.5; IQR=4. Fences: −3.5 and 12.5. The value 20 is a potential outlier. Whiskers: 1 and 7.
Review 5 · A normal probability
Assume delivery times are normal with mean 30 minutes and SD 4. Find the proportion taking 26 to 34 minutes, and the expected count among 200 deliveries.
Show solution
The interval is within one SD of the mean. Empirical rule: about 68%, or 136 deliveries. The normal CDF gives approximately 68.27%, or 136.54 expected deliveries. Different precision explains the small difference.
Review 6 · Fit, predict, and question
Fit a least-squares line to (1,2),(2,3),(3,5). Predict y at x=2.5. Find the residual for (2,3). Is predicting at x=10 interpolation?
Show solution
x̄=2, ȳ=10/3; centered cross-products total 3 and squared x-deviations total 2. Thus slope=1.5 and intercept=1/3.
ŷ=1.5x+1/3. At x=2.5, ŷ≈4.0833. At x=2, ŷ=3.3333…, so the residual is −1/3. Predicting at 10 is extrapolation beyond the observed range [1,3].
A complete statistical explanation
State the question and units. Describe how the data were collected. Choose summaries and displays that fit the variable. Interpret the result in context, and explain what the evidence cannot establish.
Additional practice: mixed exercises
Unknown observations
Use the mean equation and the positions in the sorted list.
For \(5,9,7,5,7,11,9,x\), find \(x\) so that (a) the unique mode is \(9\), (b) the median is \(7.5\), and (c) the mean is \(10\). Treat these as separate questions.
The data are \(11,16,19,11,7,10,a,b\), where \(a\le b\). If the mean is \(13\) and the median is \(12.5\), find \(a,b\).
Show worked solution
(a) \(5,7,9\) each occur twice, so \(x=9\) makes \(9\) the unique mode. (b) The seven known values sort as \(5,5,7,7,9,9,11\). The middle pair must be \(7\) and \(8\), so \(x=8\). (c) The known sum is \(53\). Thus \((53+x)/8=10\), giving \(x=27\).
The known sum is \(74\), so \(a+b=8(13)-74=30\). The known values sort as \(7,10,11,11,16,19\). If \(a\le11\), the fourth and fifth positions cannot sum to \(25\). Thus \(a>11\); since \(a\le b\) and their sum is \(30\), \(a\le15\). The middle pair is then \(11,a\), giving \((11+a)/2=12.5\). Hence \(a=14,b=16\). Check: \(7,10,11,11,14,16,16,19\).
Compare spread and quartiles
State the variance convention. For quartiles, use medians of halves, excluding the overall median when the size is odd.
For each dataset, find range, variance, and standard deviation: \(A=(35,35,35,35,35)\), \(B=(25,35,35,35,45)\), \(C=(10,25,35,45,60)\). Give both population and sample results.
A population of \(100\) observations has mean \(80\) and standard deviation \(5\). Add \(10\) to every observation. Find the new mean and standard deviation.
Find the quartiles of \(7,4,1,2,3,5,6\).
Find the quartiles of \(1,4,6,7,3,5,2,8\).
Show worked solution
All means are \(35\). Sums of squared deviations are \(0,200,1450\). Ranges are \(0,20,50\). Population variances divide by \(5\): \(0,40,290\); population SDs are \(0,\sqrt{40}\approx6.32,\sqrt{290}\approx17.03\). Sample variances divide by \(4\): \(0,50,362.5\); sample SDs are \(0,\sqrt{50}\approx7.07,\sqrt{362.5}\approx19.04\).
The mean becomes \(80+10=90\). Each deviation is unchanged: \((x+10)-(80+10)=x-80\). Therefore the standard deviation remains \(5\).
Sorted: \(1,2,3,4,5,6,7\). The median is \(Q_2=4\). Excluding it, the halves are \(1,2,3\) and \(5,6,7\), so \(Q_1=2,Q_3=6\).
Sorted: \(1,2,3,4,5,6,7,8\). Thus \(Q_2=(4+5)/2=4.5\), \(Q_1=(2+3)/2=2.5\), and \(Q_3=(6+7)/2=6.5\).
Relative position and normal distributions
Show the standardization step. Use the empirical rule where requested.
Candy counts have mean \(43\) and standard deviation \(2\). Find the z-score for \(20\) candies. How many candies correspond to \(z=6\)?
Brad scored higher than \(7255\) of the \(9430\) test takers. Find his percentile rank, rounded to an integer.
Tomato weights are normal with mean \(0.61\) lb and standard deviation \(0.15\) lb. Using the empirical rule, estimate the proportion below \(0.76\) lb.
For the same tomato distribution, estimate how many of \(6000\) weigh more than \(0.31\) lb and how many of \(4500\) weigh between \(0.31\) and \(0.91\) lb. Use the empirical rule.
Car speeds are normal with mean \(85\) km/h and standard deviation \(10\) km/h. Find the probability of a speed above \(100\) km/h.
Show worked solution
\(z=(20-43)/2=-11.5\). Reverse the formula: \(x=43+6(2)=55\).
Using the percentage strictly below, \(100(7255/9430)\approx76.94\). He is at approximately the \(77\)th percentile.
\(z=(0.76-0.61)/0.15=1\). Below one standard deviation above the mean: \(50\%+34\%=84\%\).
The limits have z-scores \(-2\) and \(2\). Above \(-2\): \(0.975(6000)=5850\). Between \(-2\) and \(2\): \(0.95(4500)=4275\).
\(z=(100-85)/10=1.5\). The upper-tail probability is \(1-\Phi(1.5)\approx0.0668\), or \(6.68\%\).
Fit and interpret the full dataset
Find the least-squares line and correlation coefficient before making predictions.
For \((x,y)=(1,4),(3,1),(4,3),(6,-1),(8,0)\), find the least-squares line, interpret the correlation coefficient, and interpolate \(y\) at \(x=5\).
A sweets machine records hours \(3.8,4.2,4.4,4.1,4.2,4.0\) and corresponding outputs \(275,287,291,281,286,278\). Find the least-squares line, interpret the correlation, and estimate the hours needed for \(300\) sweets using this line.
Show worked solution
The means are \(\bar x=4.4\), \(\bar y=1.4\). Centered sums are \(S_{xx}=29.2\), \(S_{xy}=-17.8\), \(S_{yy}=17.2\). Thus \(b=S_{xy}/S_{xx}=-89/146\) and \(a=\bar y-b\bar x=298/73\). The fitted line is \(\hat y\approx4.0822-0.6096x\). Correlation \(r=-17.8/\sqrt{29.2(17.2)}\approx-0.7943\) indicates a strong negative linear association. At \(x=5\), \(\hat y=151/146\approx1.0342\). This is interpolation because \(5\) is within \([1,8]\).
The means are \(\bar x=247/60\) and \(\bar y=283\). Centered sums are \(S_{xx}=5/24\), \(S_{xy}=6\), \(S_{yy}=182\). Thus \(b=28.8\), \(a=164.44\), and \(\hat y=28.8x+164.44\). Correlation \(r=6/\sqrt{(5/24)(182)}\approx0.9744\), a strong positive linear association. Solve \(300=28.8x+164.44\) to get \(x\approx4.71\) hours. This extrapolates beyond the observed maximum of \(4.4\) hours, so treat it cautiously.
Interpret displays and audit conclusions
A display compresses data. Before using it to make a decision, ask which information survives the compression and which has been lost. The following problems connect the frequency table, box plot and numerical summaries already developed in this lesson.
Interpretation · Same box plot, different data
Two groups report waiting times in minutes. Group A: \(0,2,2,4,6,8,8,10\). Group B: \(0,1,3,4,6,7,9,10\). Use the median-of-halves convention. Determine each five-number summary, decide whether their modified box plots coincide, and compare population variances. What cannot be concluded from equal box plots?
Hint
Both lists are sorted. Compute the medians of the four-value halves, then compare the squared deviations from the common mean.
Show worked solution
For both groups, the median is \((4+6)/2=5\), \(Q_1=2\), \(Q_3=8\), minimum \(0\), and maximum \(10\). The IQR is \(6\); fences are \(2-1.5(6)=-7\) and \(8+1.5(6)=17\). Every value lies inside the fences, so the whiskers are \(0\) and \(10\): the box plots coincide.
Both means are \(5\). Group A has squared-deviation sum \(25+9+9+1+1+9+9+25=88\), giving population variance \(88/8=11\). Group B has sum \(25+16+4+1+1+4+16+25=92\), giving \(92/8=11.5\). A box plot does not reveal every observation, the exact frequencies inside a quartile, or the variance. Equal box plots do not imply identical distributions.
Multi-step · Audit a frequency-table design
A class records \(32\) whole-number travel times, with minimum \(6\) and maximum \(36\). Choose the smallest \(k\) with \(2^k>32\), compute the initial width \(w=\lceil(H-L)/k\rceil\), then check whether classes starting at \(6\) cover the maximum. Give corrected limits and boundaries if needed. Can frequencies be recovered from this information alone?
Hint
The rule is strict. The last upper boundary must be above the maximum, not below it. The range does not tell you how observations are distributed.
Show worked solution
Because \(2^5=32\) and \(2^6=64>32\), take \(k=6\). The initial width is \(\lceil30/6\rceil=5\). Six classes with whole-number limits start at \(6,11,16,21,26,31\); the final upper limit is \(35\), with upper boundary \(35.5\). This excludes \(36\), so increase the width to \(6\).
The corrected limits are \(6\text{–}11,12\text{–}17,18\text{–}23,24\text{–}29,30\text{–}35,36\text{–}41\). Boundaries are \([5.5,11.5),[11.5,17.5),[17.5,23.5),[23.5,29.5),[29.5,35.5),[35.5,41.5)\). Now the maximum is included. Frequencies cannot be calculated without the actual observations; many datasets share the same size and extrema.
Advanced / Honors · What a grouped table cannot determine
A histogram uses classes with whole-number limits \(0\text{–}4,5\text{–}9,10\text{–}14\), each with frequency \(2\). All observations are integers. Determine the smallest and largest possible exact means. Find the midpoint estimate, and explain why it need not equal the exact mean.
Hint
To minimize a total, place each observation at its permitted lower limit. To maximize it, use the upper limits.
Show worked solution
The smallest possible total is \(2(0)+2(5)+2(10)=30\), giving mean \(5\); it is attained by \(0,0,5,5,10,10\). The largest is \(2(4)+2(9)+2(14)=54\), giving mean \(9\), attained by \(4,4,9,9,14,14\). The midpoint estimate is \([2(2)+2(7)+2(12)]/6=7\). It replaces unknown observations by class midpoints, so it is an estimate, not a uniquely determined exact mean. Both extreme datasets produce the same grouped frequencies.