MCAT Foundations · Research Methods, Statistics, and Scientific Reasoning
Correlation and Regression
On this page 5 sections
In 30 seconds
Correlation and regression are statistical tools for quantifying relationships between variables, and they appear throughout MCAT passages -- from biochemistry dose-response curves to psychology survey analyses. The correlation coefficient r measures the strength and direction of a linear relationship between two variables on a scale from -1 to +1. Linear regression extends this by fitting a line that predicts one variable from another. The MCAT tests whether you can interpret scatterplots, distinguish strong from weak correlations, recognize that correlation does not imply causation, and identify spurious correlations driven by confounding variables. These skills are essential for evaluating whether a passage's conclusions are justified by its data.
The college version
Correlation Coefficients
The Pearson correlation coefficient r quantifies the linear relationship between two continuous variables. r ranges from -1 to +1. r = +1 indicates a perfect positive linear relationship (as X increases, Y increases proportionally). r = -1 indicates a perfect negative linear relationship (as X increases, Y decreases proportionally). r = 0 indicates no linear relationship. Key properties: r is unitless and unaffected by changes in scale (multiplying all X or Y values by a constant does not change r). r is symmetric: the correlation between X and Y equals the correlation between Y and X. r is sensitive to outliers -- a single extreme point can dramatically inflate or deflate r. The coefficient of determination r^2 represents the proportion of variance in Y that is explained by X (e.g., r = 0.8 gives r^2 = 0.64, meaning 64% of Y's variance is accounted for by X). The MCAT rarely requires calculating r but frequently asks you to interpret given r values in passage data.
Scatterplots
A scatterplot displays each data point as a dot on a two-dimensional graph, with one variable on the x-axis and the other on the y-axis. The pattern of dots reveals the direction, form, and strength of the relationship. An upward-sloping cloud indicates a positive correlation; a downward-sloping cloud indicates a negative correlation. A tight clustering around an imaginary line indicates a strong correlation (|r| close to 1); a loose, scattered cloud indicates a weak correlation (|r| close to 0). A curved pattern indicates a nonlinear relationship -- r may be near 0 even when a strong non-linear relationship exists, because r only measures linear association. Outliers appear as isolated points far from the main cluster and warrant investigation. The MCAT frequently presents scatterplots in passage figures and asks you to characterize the relationship (positive/negative, strong/weak, linear/nonlinear) or to identify which of several scatterplots corresponds to a given r value.
Linear Regression
Linear regression fits a straight line through a scatterplot to model the relationship between a predictor variable X and a response variable Y. The regression equation takes the form Y-hat = b0 + b1*X, where b0 is the y-intercept (predicted Y when X = 0) and b1 is the slope (the predicted change in Y for a one-unit increase in X). Unlike correlation, regression is directional: regressing Y on X gives a different line than regressing X on Y. The least-squares method finds the line that minimizes the sum of squared vertical distances from points to the line (residuals). The regression line always passes through the point (mean of X, mean of Y). Residual plots (residuals vs. fitted values) help assess whether a linear model is appropriate -- random scatter around zero supports linearity; patterns suggest nonlinearity or unequal variance. The MCAT may present regression output tables with slopes and intercepts and ask you to interpret what the slope means in context (e.g., 'for each additional year of education, predicted income increases by b1 dollars').
Strength and Direction of Correlation
Correlation strength is judged by the absolute value of r, not the sign. r = 0.9 and r = -0.9 are equally strong; they differ only in direction. Conventional benchmarks: |r| < 0.3 is weak, 0.3 <= |r| < 0.7 is moderate, |r| >= 0.7 is strong. These cutoffs are context-dependent -- in some fields, r = 0.3 is clinically meaningful; in others, r = 0.9 is expected. Direction specifies whether variables move together (positive: both increase or both decrease) or in opposite directions (negative: one increases as the other decreases). Example: hours studied and exam score typically show a moderate-to-strong positive correlation. Number of cigarettes smoked per day and lung function (FEV1) show a negative correlation. The MCAT tests whether you can move fluidly between verbal descriptions ('there is a strong negative association'), numerical r values, and scatterplot patterns, and whether you recognize that a weak correlation does not necessarily mean 'no relationship' -- it may be nonlinear.
Spurious Correlations
A spurious correlation is a statistical association between two variables that arises not from a direct causal relationship but from coincidence, confounding, or a shared underlying cause. Classic examples include the correlation between ice cream sales and drowning deaths (both increase in summer -- the confounding variable is hot weather) and the correlation between a country's chocolate consumption and its number of Nobel laureates (likely confounded by national wealth and research funding). Spurious correlations are dangerous because they can mislead researchers and the public into inferring causation where none exists. The MCAT tests this concept by presenting passage data showing a correlation and asking whether the design (often observational) supports a causal conclusion, or by asking what confound could explain an observed association. The antidote to spurious correlation is sound experimental design with randomization, or statistical control for confounds in observational studies.
Prediction versus Causation
Correlation and regression enable prediction: if you know X, you can predict Y using the regression equation. Prediction does not require causation. A model predicting restaurant tips from the day of the week may be accurate even though the day of the week does not 'cause' higher tips -- it is a proxy for other variables like weekend dining patterns. Causation requires three criteria: (1) covariation (X and Y are correlated), (2) temporal precedence (X precedes Y in time), and (3) elimination of alternative explanations (no confounds). Observational studies can establish covariation and sometimes temporal precedence but cannot rule out confounds, so they support prediction but not causation. Only randomized experiments with proper controls can establish causation. The MCAT consistently distinguishes these: a passage may report r = 0.75 between exercise frequency and life satisfaction, then ask 'which of the following conclusions is most justified?' The correct answer will describe an association or prediction, not a causal claim. Regression coefficients from observational data are predictive, not causal.
How it works
The MCAT correlation-and-regression reasoning framework: (1) Look at the scatterplot or r value to characterize direction (positive/negative) and strength (strong/moderate/weak). (2) If a regression line is provided, interpret the slope -- what does a one-unit increase in X predict for Y? (3) Check for nonlinearity: does r capture the full relationship, or is there curvature that r misses? (4) Identify study design: is this an experiment (manipulation + randomization) or an observational study (measured variables only)? (5) Based on design, determine whether causal claims are justified or only predictive/associational claims. (6) Scan for confounds that could produce a spurious correlation. This six-step scan catches the majority of correlation and regression questions across all MCAT sections.
How it works
The MCAT correlation-and-regression reasoning framework: (1) Look at the scatterplot or r value to characterize direction (positive/negative) and strength (strong/moderate/weak). (2) If a regression line is provided, interpret the slope -- what does a one-unit increase in X predict for Y? (3) Check for nonlinearity: does r capture the full relationship, or is there curvature that r misses? (4) Identify study design: is this an experiment (manipulation + randomization) or an observational study (measured variables only)? (5) Based on design, determine whether causal claims are justified or only predictive/associational claims. (6) Scan for confounds that could produce a spurious correlation. This six-step scan catches the majority of correlation and regression questions across all MCAT sections.
Comparisons
- C/P (Chemistry/Physics): Calibration curves use linear regression (absorbance vs. concentration); r^2 values near 1.0 indicate good linearity for Beer's Law. Enzyme kinetics plots (Lineweaver-Burk) use linearized regression to extract Km and Vmax.
- B/B (Biology/Biochemistry): Dose-response curves, growth rate analyses, and gene expression correlations all rely on regression. Passages may report r values between mRNA levels and protein abundance, or between drug dose and physiological response.
- P/S (Psychology/Sociology): Survey studies report correlations between personality traits (r = 0.40 between extraversion and happiness), socioeconomic variables, or health behaviors. The entire P/S section tests your ability to not over-interpret correlational data as causal.
- CARS-like reasoning: Passages in any section may present a claimed relationship and supporting data. The skill of asking 'is this correlation or causation?' and 'what confound could explain this?' is a core scientific reasoning competency tested throughout the exam.
Common confusions
- "r = 0 means no relationship." r = 0 means no linear relationship. Two variables may have a perfect nonlinear relationship (e.g., a parabola or a sine wave) and still give r near zero. Always check the scatterplot.
- "A strong correlation proves causation." r = 0.9 does not imply X causes Y. Ice cream sales and drowning deaths are strongly correlated. Only experimental manipulation with randomization supports causal claims.
- "r is the slope of the regression line." r and the regression slope b1 are related (b1 = r * (SDy / SDx)) but distinct. r is standardized and unitless; b1 is not. A steep slope does not imply a strong correlation.
- "Linear regression works for any data." Regression assumes a roughly linear relationship, independent residuals, and constant variance. Applying linear regression to curved data produces a misleading line -- always inspect the scatterplot first.
- "r^2 = 0.49 means the model is bad." In social sciences, r^2 of 0.49 (explaining 49% of variance in human behavior) is often considered very strong. Context and field-specific norms determine what counts as a meaningful effect size.
- "If two variables are correlated, changing one will change the other." This mistakes prediction for intervention. Even with a perfect correlation (e.g., shoe size and reading ability in children -- both increase with age), manipulating shoe size will not improve reading scores. The confound (age) drives both.
Quick review
- Pearson r: ranges -1 to +1. Measures strength (absolute value) and direction (sign) of a linear relationship. Unitless and symmetric.
- r^2 (coefficient of determination): proportion of variance in Y explained by X. r = 0.7 gives r^2 = 0.49 (49% of variance explained).
- Scatterplot: X on horizontal axis, Y on vertical. Upward slope = positive, downward = negative, tightness = strength, curvature = nonlinear.
- Linear regression: Y-hat = b0 + b1*X. b0 = intercept, b1 = slope (predicted change in Y per one-unit increase in X). Least-squares minimizes squared residuals.
- Regression is directional (Y on X differs from X on Y); correlation is symmetric. The regression line always passes through (x-bar, y-bar).
- Correlation strength benchmarks: |r| < 0.3 weak, 0.3-0.7 moderate, > 0.7 strong. Context-dependent; always check field-specific norms.
- Spurious correlation: association driven by coincidence or a confounding third variable, not by a direct causal link. Hot weather -> ice cream sales and drowning.
- Causation requires: (1) covariation, (2) temporal precedence (X before Y), (3) no confounds. Only experiments with randomization establish causation.
- Prediction does not require causation. A regression model can predict Y from X accurately even when X does not cause Y.
- Outliers can dramatically inflate or deflate r. Always inspect the scatterplot; never trust r alone without visualizing the data.

Eli explains
The same idea, in plain words
Explain it like I’m 10
Imagine you run a lemonade stand and want to know what drives sales. You record the daily high temperature and how many cups you sell. When you plot this on a graph -- temperature on the bottom, cups sold going up the side -- the dots form a pattern climbing upward (a scatterplot). On hot days, you sell more; on cool days, fewer. That upward pattern is a positive correlation. The correlation coefficient r (a number between -1 and +1) measures how neatly the dots line up: r close to +1 means they follow a near-perfect upward line. Now, suppose you also notice that on days when more people wear shorts, you sell more lemonade. Is there a causal link -- does wearing shorts make people thirsty? No. Hot weather causes both shorts-wearing and lemonade thirst. That is a spurious correlation: two things move together because a third factor (temperature) drives both. Statisticians use regression to predict one thing from another (temperature helps predict sales), but prediction is not the same as proving what causes what. The limitation: this analogy treats the world as if one variable predicts another cleanly. Real data is messier -- lemonade sales also depend on weekends, nearby events, and whether you smiled at customers, none of which show up in a simple temperature-versus-sales graph.
Study tools & related lessonsRelated
Sources & references
- MCAT Content Outline: Scientific Reasoning and Research Methods — Association of American Medical Colleges (AAMC)
- Psychology 2e, Chapter 2: Psychological Research, Section 2.3 (Analyzing Findings) — OpenStax
- Biology 2e, Chapter 1: The Science of Biology, Section 1.1 (The Science of Biology) — OpenStax
- Introductory Statistics 2e, Chapter 12: Linear Regression and Correlation — OpenStax
- Simply Psychology: Variables in Research — Simply Psychology
This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
