NBDHE Review · Biostatistics (Community Health and Research Principles)
Biostatistics I: Central Tendency, Dispersion, and the Normal Distribution
On this page 6 sections
In 30 seconds
The NBDHE includes biostatistics questions that require you to select the appropriate measure of central tendency (mean, median, mode) for a given data distribution, interpret standard deviation in the context of normal distributions, and recognize when skewed data makes the mean misleading. You must also understand the properties of the normal distribution and apply the empirical rule (68-95-99.7) to estimate proportions of data within standard deviation intervals.
The college version
Core Review
Measures of Central Tendency
Central tendency describes the "center" of a data distribution — the typical or representative value.
Mean (Arithmetic Average)
Formula: Mean = (Sum of all values) ÷ (Number of values)
- The most commonly used measure of central tendency
- Highly sensitive to outliers (extreme values) — this is its key weakness
- Every value in the dataset contributes to the mean
- Appropriate for: symmetric distributions without extreme outliers (approximately normal data)
- Example: Five probing depths (mm): 2, 3, 2, 4, 3. Mean = (2+3+2+4+3)/5 = 14/5 = 2.8 mm
Median
The middle value when data are arranged in order. For an odd number of observations, it is the exact middle value. For an even number, it is the average of the two middle values.
- Resistant to outliers — this is its key strength
- Does not use all data values (only the middle position matters)
- Appropriate for: skewed distributions, data with outliers, ordinal data
- Example (same data): 2, 2, 3, 3, 4 → median = 3
- Example with outlier: 2, 2, 3, 3, 15 → median = 3 (unchanged), but mean = 25/5 = 5 (pulled up by the 15)
Mode
The most frequently occurring value in the dataset.
- Can have no mode, one mode (unimodal), or multiple modes (bimodal, multimodal)
- Appropriate for: categorical/nominal data
- Limited utility for continuous data
- Example: 2, 2, 3, 3, 4 → bimodal (2 and 3 each occur twice)
Choosing Between Mean and Median
The decision rule for the NBDHE:
- Mean: Use when data are SYMMETRIC and approximately normally distributed. In a perfectly symmetric distribution, mean = median = mode.
- Median: Use when data are SKEWED (asymmetric) or contain OUTLIERS. The median is the preferred measure for:
- Income data (very skewed; a few wealthy individuals pull the mean up)
- Waiting times (often right-skewed with long tails)
- DMFT scores in children (typically right-skewed; most have low scores, a few have very high scores)
- Hospital length of stay
- Any data where you see "the average is X, but most people are below X" — that's a sign the mean is being pulled by a right tail
Key test concept: In a RIGHT-skewed distribution (long tail to the right), mean > median > mode. In a LEFT-skewed distribution, mean < median < mode. The mean is always pulled toward the skew.
Measures of Dispersion
Dispersion describes how spread out the data are around the center.
Range: Maximum value − Minimum value
- Simple but extremely sensitive to outliers
- Uses only two data points, ignores all information about the distribution between them
- Example: DMFT scores from 0 to 16 → Range = 16
Variance: The average of the squared deviations from the mean
Formula: Variance (σ² for population, s² for sample) = Σ(x − mean)² / (n − 1)
- (n − 1) in the denominator for samples (Bessel's correction) — this makes the sample variance an unbiased estimator of the population variance
- Squaring the deviations makes variance difficult to interpret directly (units are squared)
- Example: Data: 2, 3, 2, 4, 3; Mean = 2.8; Deviations: −0.8, 0.2, −0.8, 1.2, 0.2; Squared: 0.64, 0.04, 0.64, 1.44, 0.04; Sum = 2.80; s² = 2.80/4 = 0.70
Standard Deviation (SD): The square root of the variance
Formula: SD = √(Variance)
- Returns the measure to the original units, making it directly interpretable
- Represents the "typical" distance of observations from the mean
- Larger SD = more spread; smaller SD = data clustered tightly around the mean
- Example: s = √0.70 ≈ 0.84 mm
The Normal Distribution
The normal distribution (Gaussian distribution, bell curve) is a symmetric, unimodal probability distribution that describes many natural and biological phenomena.
Key properties:
- Symmetric around the mean (mean = median = mode)
- Defined entirely by two parameters: mean (μ) and standard deviation (σ)
- Total area under the curve = 1 (100% of observations)
- Asymptotic: The tails approach but never touch the x-axis (extends to ±∞)
- Continuous (not discrete)
The Empirical Rule (68-95-99.7 Rule):
For data that follow a normal distribution:
- 68% of observations fall within ±1 SD of the mean
- 95% of observations fall within ±2 SD of the mean
- 99.7% of observations fall within ±3 SD of the mean
Example: If the mean probing depth in a population is 2.5 mm with SD = 0.8 mm, assuming normality:
- ~68% of sites have probing depths between 1.7 and 3.3 mm (±1 SD)
- ~95% of sites have probing depths between 0.9 and 4.1 mm (±2 SD)
- ~99.7% of sites have probing depths between 0.1 and 4.9 mm (±3 SD)
Clinical application: Reference ranges for laboratory values (e.g., "normal" blood pressure, "normal" HbA1c) are typically defined as the interval within ±2 SD of the mean in a healthy reference population, capturing approximately 95% of healthy individuals.
When the Mean is Misleading: Skewness
Consider a dental public health example: DMFT scores in a community screening of 100 third-graders:
- 75 children have DMFT = 0
- 15 have DMFT = 2
- 5 have DMFT = 5
- 3 have DMFT = 8
- 2 have DMFT = 12
Mean = (75×0 + 15×2 + 5×5 + 3×8 + 2×12)/100 = (0 + 30 + 25 + 24 + 24)/100 = 103/100 = 1.03
Median = 0 (more than half have zero DMFT)
The mean (1.03) suggests, misleadingly, that the "typical" child has about one DMFT. In reality, 75% have zero. The few children with very high DMFT values pull the mean upward. In this case, the median (0) is a far better description of the "typical" child's caries experience.
Reporting both mean and median, or the proportion with DMFT = 0 alongside mean DMFT, provides a more complete picture in skewed data.
Clinical/Board Application
Board-style question: "A dental public health researcher reports that the mean DMFT in a Head Start population is 4.2 (SD = 3.1). Assuming a normal distribution, approximately what percentage of children have DMFT scores between 1.1 and 7.3?"
Answer: 68%. The interval 1.1 to 7.3 is 4.2 ± 3.1 (mean ± 1 SD). By the empirical rule, ~68% of observations fall within ±1 SD of the mean.
Common Traps
- Trap: Automatically reporting the mean for all data. If the question describes skewed data or mentions outliers, the median is the correct answer.
- Trap: Confusing variance and standard deviation. SD = √(variance). If a question gives you variance = 9, the SD = 3.
- Trap: Applying the empirical rule to non-normal data. The 68-95-99.7 rule only works for approximately normal distributions. Highly skewed data (like DMFT) do not follow this rule.
- Trap: Forgetting that SD represents "typical" deviation, not the maximum possible deviation. Data can (and do) lie beyond ±2 SD — about 5% of them.

Eli explains
The same idea, in plain words
Explain it like I’m 10
The mean is just the average — add everything up and divide by how many. But averages can lie! If nine kids have zero cavities and one kid has 20 cavities, the "average" is 2 — but 90% of kids have none. The median (the middle kid, after lining everyone up) tells you that most kids have zero.
Standard deviation is a fancy way of saying "how spread out things are." If everyone's probing depths are between 2 and 3 mm, the SD is tiny. If some people have 1 mm and others have 8 mm, the SD is big.
The bell curve rule: About 68% of people are "within one SD of average," 95% within two SDs, and 99.7% within three. That's why "normal" lab values cover about 95% of healthy people — they're the mean ± 2 SD.
Key takeaways
- Mean is sensitive to outliers; median is resistant
- In right-skewed data: mean > median > mode
- Standard deviation: square root of variance; same units as the data
- 68% within ±1 SD; 95% within ±2 SD; 99.7% within ±3 SD
- (n−1) for sample variance; n for population variance
- Range: max − min; ignores distribution
- Q1: Income data in a community typically follows a right-skewed distribution. Which measure of central tendency is MOST appropriate for describing the "typical" income?
- A. Mean
- B. Median ✓ — In right-skewed data, the mean is pulled upward by high-income outliers and overstates the typical value. The median is resistant to outliers and provides a better measure of central tendency for skewed distributions.
- C. Mode
- D. Range
- Q2: In a normally distributed dataset of DMFT scores with mean = 3.0 and standard deviation = 1.5, approximately what proportion of observations fall between 1.5 and 4.5 DMFT?
- A. 50%
- B. 68% ✓ — The interval 1.5 to 4.5 is mean ± 1 SD (3.0 ± 1.5). By the empirical rule for normal distributions, approximately 68% of observations fall within one standard deviation of the mean.
- C. 95%
- D. 99.7%
- Q3: A dataset has a variance of 16 mm². What is the standard deviation?
- A. 2 mm
- B. 4 mm ✓ — Standard deviation = √(variance) = √16 = 4 mm. Remember to report SD in the original units (mm, not mm²).
- C. 8 mm
- D. 256 mm
Quick check
3 questions here. Answers stay hidden until you check.
In a normally distributed dataset of DMFT scores with mean = 3.0 and standard deviation = 1.5, approximately what proportion of observations fall between 1.5 and 4.5 DMFT?
A dataset has a variance of 16 mm². What is the standard deviation?
Study toolsYou’ll learn to
You’ll learn to
- Define and calculate mean, median, and mode; identify which is appropriate for different data distributions
- Explain range, variance, and standard deviation as measures of dispersion
- Describe the properties of the normal distribution (bell curve)
- Apply the empirical rule (68-95-99.7 rule)
- Interpret standard deviation in clinical and population health contexts
- Recognize when mean is misleading (skewed distributions) and when median is preferred
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
