NBDHE Review · Biostatistics (Community Health and Research Principles)
Biostatistics II: Sensitivity, Specificity, Correlation, and p-Value Interpretation
On this page 6 sections
In 30 seconds
This topic covers the most commonly misinterpreted statistical concepts on the NBDHE. You must be able to calculate sensitivity and specificity from 2×2 contingency tables, memorize the mnemonic devices SnNout and SpPin, correctly interpret correlation coefficients (including the critical maxim "correlation does not imply causation"), and avoid the most common p-value misinterpretation — that the p-value represents the probability that the null hypothesis is true (it doesn't). Distinguish statistical significance from clinical significance.
The college version
Core Review
The 2×2 Contingency Table
All diagnostic test evaluation begins with the 2×2 table, which cross-tabulates test results against the gold standard (truth):
| Disease Present | Disease Absent | |
|---|---|---|
| Test Positive | True Positive (TP) | False Positive (FP) |
| Test Negative | False Negative (FN) | True Negative (TN) |
Gold standard (reference standard): The definitive method for determining whether disease is truly present (e.g., histopathology for oral cancer, full-mouth periodontal charting with radiographic bone loss for periodontitis, clinical/radiographic exam for caries).
Sensitivity
Definition: The proportion of people WITH the disease who test positive. Sensitivity answers the question: "If the patient HAS the disease, how likely is the test to catch it?"
Formula: Sensitivity = TP / (TP + FN)
- A test with high sensitivity has FEW false negatives. It "catches" the disease reliably.
- SnNout: A highly Sensitive test, when Negative, rules out the disease. If sensitivity is very high (say, 99%), a negative result means it's very unlikely the patient has the disease — because the test almost never misses a true case.
Clinical example: A caries detection device with sensitivity of 95% means that 95% of truly carious surfaces will be identified. Only 5% of carious surfaces are missed (false negatives).
Specificity
Definition: The proportion of people WITHOUT the disease who test negative. Specificity answers: "If the patient does NOT have the disease, how likely is the test to correctly say so?"
Formula: Specificity = TN / (TN + FP)
- A test with high specificity has FEW false positives. It doesn't "cry wolf."
- SpPin: A highly Specific test, when Positive, rules in the disease. If specificity is very high (say, 98%), a positive result strongly suggests the patient truly has the disease — because the test almost never falsely flags healthy people.
Clinical example: A diagnostic test with specificity of 90% means that 90% of healthy (disease-free) surfaces test negative. 10% of healthy surfaces produce false alarms (false positives).
Worked 2×2 Example
A new oral cancer screening device is evaluated against biopsy (gold standard) in 1,000 patients:
| Biopsy Positive (Cancer) | Biopsy Negative (No Cancer) | Total | |
|---|---|---|---|
| Test Positive | 80 (TP) | 60 (FP) | 140 |
| Test Negative | 20 (FN) | 840 (TN) | 860 |
| Total | 100 | 900 | 1,000 |
Calculations:
- Sensitivity = 80 / (80 + 20) = 80/100 = 0.80 = 80% (the test catches 80% of cancers)
- Specificity = 840 / (840 + 60) = 840/900 = 0.933 = 93.3% (93.3% of non-cancer patients correctly test negative)
- Positive Predictive Value (PPV) = TP / (TP + FP) = 80 / 140 = 0.571 = 57.1% (If the test is positive, 57.1% probability the patient truly has cancer)
- Negative Predictive Value (NPV) = TN / (TN + FN) = 840 / 860 = 0.977 = 97.7% (If the test is negative, 97.7% probability the patient truly does NOT have cancer)
Critical distinction: Sensitivity and specificity are INHERENT PROPERTIES of the test and do not change with disease prevalence. PPV and NPV DO change with prevalence — as prevalence increases, PPV increases and NPV decreases. The same test performs differently in a high-risk specialty clinic (high prevalence → high PPV) versus a low-risk community screening (low prevalence → low PPV, more false positives).
SnNout and SpPin
| Mnemonic | Meaning | Clinical Application |
|---|---|---|
| SnNout | High Sensitivity, Negative result → rules OUT disease | A negative result on a highly sensitive test is strong evidence AGAINST disease |
| SpPin | High Specificity, Positive result → rules IN disease | A positive result on a highly specific test is strong evidence FOR disease |
This is tested by asking: "A test with 99% sensitivity is negative. What can you conclude?" Answer: Disease is very unlikely to be present (SnNout).
Correlation
Pearson correlation coefficient (r): Measures the strength and direction of a LINEAR relationship between two continuous variables.
- Range: −1.0 to +1.0
- r = +1.0: Perfect positive linear correlation (as X increases, Y increases proportionally)
- r = −1.0: Perfect negative linear correlation (as X increases, Y decreases proportionally)
- r = 0: No linear correlation
- r > 0: Positive correlation (both variables move in the same direction)
- r < 0: Negative correlation (variables move in opposite directions)
Interpretation guidelines (approximate):
- |r| < 0.3: Weak correlation
- 0.3 ≤ |r| < 0.7: Moderate correlation
- |r| ≥ 0.7: Strong correlation
Spearman rank correlation: Used when data are ordinal or when the relationship is monotonic but not necessarily linear. Based on ranks rather than raw values.
Correlation ≠ Causation
This is arguably the most important single concept in biostatistics — and one of the most frequently tested on the NBDHE.
A statistically significant correlation between two variables does NOT mean that one caused the other. Alternative explanations for a correlation include:
- Confounding: A third variable causes both. Toothbrushing frequency may correlate with lower BMI, but the real confounder is likely health-consciousness/socioeconomic status.
- Reverse causation: The direction of causality might be opposite to what is assumed. Does poor oral health cause diabetes, or does diabetes cause poor oral health? (Evidence supports bidirectionality.)
- Coincidence: In large datasets with many variables tested, some correlations will reach statistical significance purely by chance (multiple comparison problem).
To establish causation, you need: temporality (cause precedes effect), biological plausibility, dose-response relationship, consistency across studies, and ideally experimental evidence (RCT).
The p-Value
The p-value is the most widely misunderstood statistical concept. Here is what the p-value IS and IS NOT:
What the p-value IS: The probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is true.
In simpler terms: "If there were truly no effect (null hypothesis), how likely would it be to see data like these (or more extreme) purely due to random chance?"
What the p-value IS NOT (common NBDHE traps):
- ❌ The probability that the null hypothesis is true
- ❌ The probability that the results are due to chance
- ❌ The probability that the alternative hypothesis is false
- ❌ 1 − (probability that the alternative hypothesis is true)
- ❌ A measure of the size or importance of an effect
Conventional thresholds:
- p < 0.05: "Statistically significant" (results unlikely under the null; by convention, reject the null)
- p ≥ 0.05: "Not statistically significant" (fail to reject the null — note: you do NOT "accept" the null)
The p = 0.05 trap: If you run 20 independent tests where the null hypothesis is true, on average, one will produce p < 0.05 purely by chance. This is why exploratory analyses (fishing expeditions) without correction for multiple comparisons are problematic.
Statistical Significance ≠ Clinical Significance
A common board question presents a study with p < 0.001 (highly statistically significant) but an effect size that is trivially small. The correct interpretation: the result is statistically significant but may not be CLINICALLY significant.
- Statistical significance: Addresses whether the observed effect is likely to be real (not due to chance), given the sample size. With large enough sample sizes, even trivially small effects become statistically significant.
- Clinical significance: Addresses whether the effect is large enough to matter in practice. A new toothpaste that reduces plaque index by 0.05 on a 0–3 scale might be statistically significant with a sample of 10,000 subjects (p < 0.001), but is this reduction clinically meaningful? Probably not.
Always consider effect size (Cohen's d, absolute risk reduction, number needed to treat) alongside p-values.
Clinical/Board Application
Board-style question: "A new salivary diagnostic test for periodontitis has 95% sensitivity and 90% specificity. A patient tests positive. What does the positive result indicate, and which mnemonic applies?"
Answer: SpPin — a highly Specific test, when Positive, rules IN the disease. With 90% specificity, false positives are relatively uncommon, so a positive result is reasonably strong evidence for periodontitis (though PPV also depends on the pretest probability/prevalence).
Common Traps
- Trap: "The p-value is 0.03, so there is a 3% chance the null hypothesis is true." WRONG. The p-value is P(data | null), not P(null | data). You cannot flip the conditional probability.
- Trap: "The p-value is 0.08, so the null hypothesis is true / the intervention has no effect." WRONG. You fail to reject the null; you do not accept it. The study may simply have been underpowered.
- Trap: "A study found r = 0.45 between flossing frequency and lower CRP levels, therefore flossing reduces systemic inflammation." WRONG. This is correlation, not causation. A confounder (e.g., overall health-conscious behavior) could explain both.
- Trap: Confusing sensitivity and PPV. Sensitivity = "if diseased, test positive." PPV = "if test positive, is diseased." They answer opposite conditional questions.

Eli explains
The same idea, in plain words
Explain it like I’m 10
Think of a cancer screening test like a metal detector at the airport:
Sensitivity is how good the detector is at finding actual weapons. High sensitivity means it catches almost everything — but it might beep at your belt buckle too. SnNout: If a super-sensitive detector stays silent (Negative), you can be very confident there's no weapon (rules OUT).
Specificity is how good the detector is at ignoring harmless things. High specificity means when it beeps, it's almost certainly a real weapon, not your keys. SpPin: If a super-specific detector beeps (Positive), you should take it seriously (rules IN).
Correlation vs. Causation: Ice cream sales and drowning deaths both go up in summer. They're correlated. Does ice cream cause drowning? No — hot weather causes both.
p-value: Imagine flipping a coin 10 times and getting 9 heads. The p-value asks: "If the coin were fair (null hypothesis), how often would I see 9 or more heads just by luck?" Not: "What's the chance the coin is fair?" Those are completely different questions.
Key takeaways
- Se = TP/(TP+FN); Sp = TN/(TN+FP)
- PPV = TP/(TP+FP); NPV = TN/(TN+FN) — these vary with prevalence
- SnNout: high Sensitivity, Negative → rules OUT
- SpPin: high Specificity, Positive → rules IN
- Correlation r: −1 to +1; measures linear relationship
- Correlation ≠ causation (confounders, reverse causation, coincidence)
- p-value: probability of data (or more extreme) given the null is true; NOT probability the null is true
- Statistical significance ≠ clinical significance
- Q1: A caries detection device has sensitivity of 88% and specificity of 75%. What percentage of TRULY CARIOUS surfaces will the device correctly identify?
- A. 88% ✓ — Sensitivity is the proportion of true positives among those with disease: TP/(TP+FN). A sensitivity of 88% means the device correctly identifies 88% of surfaces that truly have caries.
- B. 75%
- C. Depends on the prevalence of caries in the population
- D. Cannot be determined from the information provided
- Q2: A dental researcher reports, "The correlation between sugar-sweetened beverage consumption and DMFT was r = 0.52, p < 0.001. Therefore, sugar-sweetened beverages cause dental caries." Which of the following is the most accurate critique?
- A. The correlation is too weak to be meaningful
- B. Correlation does not establish causation; confounding variables could explain the association ✓ — Even a strong, statistically significant correlation does not prove causation. Confounders (socioeconomic status, access to care, overall diet quality, oral hygiene habits) could explain the observed association.
- C. The p-value is not statistically significant
- D. The sample size was too small
- Q3: Which of the following correctly states what a p-value of 0.04 means?
- A. There is a 4% probability that the null hypothesis is true
- B. There is a 96% probability that the alternative hypothesis is true
- C. If the null hypothesis were true, the probability of obtaining results at least as extreme as those observed is 4% ✓ — The p-value is P(data at least as extreme | null hypothesis true). It is a statement about the data given the null, not about the null given the data.
- D. The probability that the results are due to chance is 4%
Quick check
3 questions here. Answers stay hidden until you check.
A dental researcher reports, "The correlation between sugar-sweetened beverage consumption and DMFT was r = 0.52, p < 0.001. Therefore, sugar-sweetened beverages cause dental caries." Which of the following is the most accurate critique?
Which of the following correctly states what a p-value of 0.04 means?
Study toolsYou’ll learn to
You’ll learn to
- Construct and interpret a 2×2 contingency table for diagnostic test evaluation
- Calculate and interpret sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV)
- Apply SnNout and SpPin mnemonics to clinical decision-making
- Interpret Pearson correlation coefficients (r) and Spearman rank correlation
- Explain why correlation does NOT imply causation
- Define p-value correctly and identify common misinterpretations
- Distinguish statistical significance from clinical significance
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
