MCAT Foundations · Research Methods, Statistics, and Scientific Reasoning
Inferential Statistics and Hypothesis Testing
On this page 5 sections
In 30 seconds
Inferential statistics lets researchers draw conclusions about populations from sample data. The logic of inference is: assume nothing is happening (null hypothesis), collect data, and ask whether the data are surprising enough to reject that assumption. The p-value quantifies how surprising the data are under the null -- a small p-value (typically < 0.05) means the data are unlikely if the null is true, leading researchers to reject the null in favor of the alternative. But inference carries risk: Type I errors (false positives) and Type II errors (false negatives) are the two ways inference can be wrong. Confidence intervals estimate the plausible range of a population parameter. Specific tests match specific data types: t-tests compare two group means, chi-square tests analyze categorical frequency data. Statistical power is the probability of detecting a real effect when one exists. On the MCAT, passages present experimental results with p-values, confidence intervals, and significance claims; you must interpret whether the statistical evidence supports the authors conclusions and whether the appropriate test was used.
The college version
Null and Alternative Hypotheses
Hypothesis testing begins with two mutually exclusive statements. The null hypothesis (H0) is the default position: there is no effect, no difference, no association. The alternative hypothesis (Ha or H1) is the researcher's prediction: there is an effect, difference, or association. For example, in a drug trial, H0: the drug has no effect on blood pressure; Ha: the drug lowers blood pressure. The null is assumed true until evidence suggests otherwise -- this is analogous to 'innocent until proven guilty.' The alternative can be one-tailed (directional: 'the drug lowers blood pressure') or two-tailed (non-directional: 'the drug changes blood pressure in either direction'). One-tailed tests have more power to detect an effect in the specified direction but cannot detect an effect in the opposite direction. Two-tailed tests are more conservative and are the default in most research because they guard against surprise findings in either direction. The MCAT often asks whether a one-tailed or two-tailed test is appropriate given the research question -- if the passage does not specify a directional prediction, a two-tailed test is typically correct.
p-Values and Statistical Significance
The p-value is the probability of obtaining results at least as extreme as those observed, assuming the null hypothesis is true. It is not the probability that the null is true, nor is it the probability that the results are due to chance -- both are common misinterpretations. A p-value of 0.03 means: if the null were true, we would see results this extreme only 3% of the time. The significance level (alpha, usually 0.05) is the threshold for rejecting the null. If p < alpha, the result is 'statistically significant' and the null is rejected. If p >= alpha, the result is not significant and the null is not rejected -- note that the null is never 'accepted,' only 'not rejected,' because absence of evidence is not evidence of absence. The MCAT commonly tests p-value interpretation: a significant result (p < 0.05) means the observed difference or association is unlikely to be due to sampling variability alone. It does not automatically mean the result is important, large, or clinically meaningful -- statistical significance and practical significance are different concepts. A very large sample can produce a statistically significant result for a trivially small effect.
Type I and Type II Errors
Type I Error (false positive, alpha): rejecting a true null hypothesis. This is concluding there is an effect when none exists. The probability of a Type I error equals the significance level (alpha), typically 0.05. If alpha is set to 0.05, there is a 5% chance of a Type I error each time a test is run. Running multiple tests inflates the familywise Type I error rate; corrections like the Bonferroni correction divide alpha by the number of tests to control this. Type II Error (false negative, beta): failing to reject a false null hypothesis. This is missing a real effect. The probability of a Type II error depends on sample size, effect size, and alpha. Power = 1 - beta, so larger samples and larger effects produce higher power and lower Type II error risk. Type I and Type II errors trade off: decreasing alpha (making it harder to reject the null) reduces Type I errors but increases Type II errors. The MCAT tests this by asking: 'If researchers set alpha to 0.01 instead of 0.05, what happens to the chance of a Type II error?' (Answer: it increases.) Or: 'Which error is more dangerous in this clinical context?' -- missing an effective drug (Type II) may be worse than approving an ineffective one (Type I), or vice versa depending on the drug's safety profile.
Confidence Intervals
A confidence interval (CI) estimates the range within which a population parameter is likely to fall. A 95% CI means: if the study were repeated many times, 95% of the calculated intervals would contain the true population parameter. It does not mean there is a 95% probability that the parameter is within the interval -- the parameter is fixed, not random. CIs provide more information than a binary significance decision: they show precision (narrower = more precise, from larger samples) and effect size. A CI for a difference between two means (e.g., treatment minus control) that includes zero is equivalent to a non-significant result at the same confidence level. If the 95% CI for the difference is (-2, 8), zero is inside the interval, meaning we cannot reject the null of no difference at alpha = 0.05. If the interval is (3, 9), zero is excluded, indicating a significant difference. Key formula for a mean: CI = sample mean +/- (critical value x standard error). The critical value depends on the distribution (z for known population SD, t for estimated SD) and the confidence level. The MCAT tests CI interpretation by asking: 'If the 95% CI for a mean difference is (-0.5, 4.2), what can we conclude?' (We cannot reject the null; the difference is not statistically significant at the 0.05 level.)
t-Tests
t-Tests compare means between two groups or compare a single group mean to a known value. There are three main types. One-sample t-test: compares a sample mean to a known population mean. Independent-samples t-test (unpaired t-test): compares means of two different, independent groups (e.g., treatment vs. control). Paired-samples t-test (dependent t-test): compares means from the same group at two time points or under two conditions (e.g., pre-test vs. post-test scores). The t-statistic formula: t = (difference between means) / (standard error of the difference). A larger t-statistic (in absolute value) makes the p-value smaller. The t-distribution is wider and heavier-tailed than the normal distribution when sample sizes are small; as degrees of freedom increase, it approaches the normal distribution. Degrees of freedom for a two-sample t-test is approximately n1 + n2 - 2. Key assumptions: the data are approximately normally distributed (check this for small samples; large samples are robust via the Central Limit Theorem), and for the independent-samples t-test, variances should be roughly equal (homogeneity of variance; Welch's t-test can be used if this assumption is violated). The MCAT tests t-test selection: 'A researcher measures anxiety scores before and after a meditation intervention in the same 30 participants. Which test?' The answer is a paired-samples t-test because the same participants are measured twice.
Chi-Square Tests
Chi-square (chi-squared) tests analyze categorical (frequency) data rather than continuous measurements. There are two main types. Chi-square goodness-of-fit test: tests whether observed frequencies in a single categorical variable match expected frequencies based on a theoretical distribution. For example, testing whether a die is fair by comparing observed rolls of each face to the expected proportion (1/6 each). Chi-square test of independence (or association): tests whether two categorical variables are associated. Data are arranged in a contingency table (e.g., gender by smoking status). The null hypothesis is that the two variables are independent (no association). The chi-square statistic measures the discrepancy between observed and expected frequencies: chi-squared = sum[(O - E)^2 / E], where O = observed frequency and E = expected frequency under the null. A large chi-square value produces a small p-value, leading to rejection of the null. Key assumption: expected frequencies should be at least 5 in at least 80% of cells, and no cell should have an expected frequency less than 1. If this assumption is violated, use Fisher's exact test instead. The MCAT tests chi-square by asking you to identify the appropriate test for categorical outcome data (e.g., 'Is political party affiliation associated with vaccine status?') and to interpret a contingency table with a reported chi-square statistic.
Statistical Power
Statistical power is the probability of correctly rejecting a false null hypothesis -- that is, the probability of finding an effect that actually exists. Power = 1 - beta, where beta is the Type II error rate. Four factors determine power: (1) Effect size -- larger effects are easier to detect and yield higher power. (2) Sample size -- larger samples reduce standard error and increase power. This is why researchers conduct power analyses before studies to determine the required sample size. (3) Significance level (alpha) -- a more lenient alpha (e.g., 0.05 vs. 0.01) increases power because it is easier to reject the null, but this also increases the Type I error rate. (4) Variability in the data -- less variability (smaller standard deviation) increases power because the signal is clearer against the noise. Power is conventionally set at 0.80, meaning researchers design studies to have an 80% chance of detecting the effect if it exists. An underpowered study may miss a real effect (Type II error), while an overpowered study (extremely large sample) may detect trivially small effects that are statistically significant but practically meaningless. The MCAT tests power by asking: 'What happens to the probability of a Type II error if the sample size is doubled?' (It decreases because power increases.)
How it works
When an MCAT passage reports inferential statistics, follow a systematic interpretation: (1) Identify the null hypothesis -- what 'no effect' looks like. (2) Check the p-value: if p < 0.05, the result is significant and the null is rejected. (3) Examine the confidence interval: does it include zero (for differences) or one (for ratios)? If so, the result is not significant at the matching confidence level. (4) Confirm the test is appropriate -- t-test for means, chi-square for frequencies, ANOVA for 3+ groups. (5) Watch for error-type traps: a non-significant result does not prove the null is true (possible Type II error), and a significant result does not guarantee practical importance. (6) Assess power: small samples increase Type II risk; very large samples can make trivial effects significant. (7) Check for multiple comparisons -- if many tests were run, significance may be inflated unless corrected.
How it works
When an MCAT passage reports inferential statistics, follow a systematic interpretation: (1) Identify the null hypothesis -- what 'no effect' looks like. (2) Check the p-value: if p < 0.05, the result is significant and the null is rejected. (3) Examine the confidence interval: does it include zero (for differences) or one (for ratios)? If so, the result is not significant at the matching confidence level. (4) Confirm the test is appropriate -- t-test for means, chi-square for frequencies, ANOVA for 3+ groups. (5) Watch for error-type traps: a non-significant result does not prove the null is true (possible Type II error), and a significant result does not guarantee practical importance. (6) Assess power: small samples increase Type II risk; very large samples can make trivial effects significant. (7) Check for multiple comparisons -- if many tests were run, significance may be inflated unless corrected.
Comparisons
- B/B (Experimental passages): Passage figures display error bars (which often represent 95% CIs or SEM); overlapping error bars do not necessarily mean non-significance, but non-overlapping bars usually indicate significance.
- C/P (Data analysis): Graphs of physical measurements typically report means with standard deviation or standard error; understanding the difference (SD describes spread of data; SEM describes precision of the mean estimate) is testable.
- P/S (Research methods): The P/S section directly tests hypothesis testing logic, p-value interpretation, and the ability to evaluate whether a study's statistical conclusions are justified.
- RM-008 (Descriptive Statistics): Descriptive stats (mean, SD, distribution shape) are the inputs to inferential tests. You cannot interpret a t-test without understanding what the mean and standard deviation represent.
- RM-010 (Correlation and Regression): Inferential tests for correlation (testing whether r is significantly different from zero) and regression (testing whether slopes differ from zero) extend the hypothesis testing framework to relationships between continuous variables.
Common confusions
- Misinterpreting p-value: "p = 0.04 means there is a 4% chance the null is true" is WRONG. The p-value assumes the null is true and gives the probability of the data under that assumption.
- Non-significance = no effect: A p-value > 0.05 does not prove the null. It may reflect low power (small sample, small effect). The null is not accepted; it is simply not rejected.
- Confusing standard deviation and standard error: SD describes the spread of individual data points. SEM = SD / sqrt(n) and describes the precision of the sample mean. SEM is always smaller than SD for n > 1. CIs use SEM, not SD.
- Overlapping confidence intervals: Slightly overlapping 95% CIs for two independent group means does not guarantee a non-significant difference between them. The test of the difference is more conservative than individual CIs.
- Multiple comparisons problem: Running 20 t-tests on the same data at alpha = 0.05 produces roughly 1 expected false positive by chance alone. Without correction (e.g., Bonferroni), reported significant findings may be Type I errors.
- One-tailed vs. two-tailed: Switching to a one-tailed test after seeing the data to get significance is p-hacking and inflates Type I error. The choice must be justified before data collection.
- Chi-square for small samples: Using chi-square when expected frequencies are less than 5 violates the test's assumptions and yields unreliable p-values. Fisher's exact test is the alternative.
Quick review
- Null hypothesis (H0): default position of no effect, no difference, no association
- Alternative hypothesis (Ha): researcher's claim of an effect or difference
- p-value: probability of observing data at least as extreme, assuming H0 is true
- Alpha (significance level): threshold for rejecting H0, conventionally 0.05
- If p < alpha: reject H0 (statistically significant result)
- If p >= alpha: fail to reject H0 (not significant; H0 is never accepted)
- Type I error (alpha): false positive -- rejecting a true H0
- Type II error (beta): false negative -- failing to reject a false H0
- Power = 1 - beta: probability of correctly rejecting a false H0
- 95% CI: if repeated, 95% of intervals contain the true parameter
- CI includes null value (0 for difference, 1 for ratio): not significant
- One-sample t-test: compare sample mean to known population mean
- Independent t-test: compare means of two different groups
- Paired t-test: compare means of same group at two time points
- Chi-square goodness-of-fit: one categorical variable vs. expected distribution
- Chi-square test of independence: association between two categorical variables
- Power increases with: larger sample, larger effect size, higher alpha, lower variability
- Multiple comparisons inflate Type I error; correct with Bonferroni or similar

Eli explains
The same idea, in plain words
Explain it like I’m 10
Imagine you claim you can tell the difference between Coke and Pepsi by taste alone. I pour 10 cups -- 5 of each, randomly ordered -- and you taste each one and guess. If you are just guessing (the null hypothesis), you would expect to get about 5 right by chance. You get 9 out of 10 right. The p-value is the probability of getting 9 or more right just by guessing -- roughly 1%. That 1% is small enough that I reject the null (you are guessing) and accept the alternative (you really can tell the difference). But here is where it gets tricky: if you actually got 6 right, the p-value might be around 40% -- not surprising under the null. That does not prove you are guessing (Type II error -- you might be good at this but unlucky today, or 10 cups was too few). And if I test 100 people instead of just you, at least one of them will probably get 9 right by pure luck (multiple comparisons inflating Type I error). Confidence intervals give more detail: "You got 90% correct (95% CI: 55% to 100%)." The CI includes 50%, so with only 10 cups, I cannot rule out guessing. A t-test would be inappropriate here because the data are correct/incorrect (categorical), not a continuous measurement. This is a chi-square scenario. Limitation: the taste-test analogy only captures one sample proportion test. Real inferential statistics spans comparisons of means, associations between variables, and estimation of population parameters -- the t-test and chi-square cover the most common MCAT scenarios but are just the beginning.
Study tools & related lessonsRelated
Sources & references
- Introductory Statistics 2e - Chapter 9: Hypothesis Testing with One Sample — OpenStax
- Introductory Statistics 2e - Chapter 10: Hypothesis Testing with Two Samples — OpenStax
- Introductory Statistics 2e - Chapter 11: The Chi-Square Distribution — OpenStax
- Introductory Statistics 2e - Chapter 8: Confidence Intervals — OpenStax
- MCAT Content Outline: Scientific Inquiry and Reasoning Skills — AAMC
This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
