MCAT Foundations · Research Methods, Statistics, and Scientific Reasoning

Sampling and Generalizability

10 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 5 sections
  1. In 30 seconds
  2. The college version
  3. Eli explains
  4. Study tools
  5. Sources & references

In 30 seconds

Every study studies a sample, but every study claims to say something about a population. The bridge between the two is sampling. A population is the complete set of individuals the researcher wants to draw conclusions about; a sample is the subset actually studied. The quality of that bridge -- whether the sample fairly represents the population -- determines a study's external validity and generalizability. Random sampling is the gold standard because it gives every population member an equal chance of selection, eliminating systematic bias. Selection bias occurs when the sampling method overrepresents or underrepresents certain groups, producing a sample that does not reflect the population. A representative sample mirrors the population's key characteristics (age, sex, SES, etc.). Generalizability asks: to whom, to what settings, and under what conditions do the results apply? The MCAT frequently presents studies and asks you to evaluate whether the sampling method supports the stated conclusions. A study with perfect internal validity can still collapse if the sample is not representative of the population the authors claim to be studying.

The college version

Population versus Sample

The population is the entire group of interest -- all patients with Type 2 diabetes, all registered voters in the US, all neurons in the hippocampus. The sample is the subset of the population that is actually studied. Because studying entire populations is usually impractical, researchers draw inferences from the sample to the population. The key question is whether the sample accurately reflects the population. Key terms: the sampling frame is the list or method from which the sample is drawn (e.g., a voter registration list, a hospital patient database). If the sampling frame systematically excludes part of the population (e.g., unregistered voters, patients without insurance), the sample cannot represent the full population regardless of how participants are selected from within the frame. The MCAT tests this by asking whether a study's sampling frame matches its stated population of interest.

Random Sampling

Random sampling (probability sampling) is a method where every member of the population has a known, nonzero chance of being selected. Simple random sampling -- every individual has an equal probability of selection -- is the most straightforward form. Other probability methods include: stratified random sampling (dividing the population into strata such as age groups or sexes, then randomly sampling within each stratum to ensure representation of subgroups), cluster sampling (randomly selecting entire clusters such as schools or neighborhoods, then studying all members within each cluster), and systematic sampling (selecting every kth individual from a list, e.g., every 10th name). Random sampling differs fundamentally from random assignment: random sampling determines who is in the study (affecting external validity), while random assignment determines which condition participants receive (affecting internal validity). A study can have random assignment without random sampling -- for example, a convenience sample of college students randomly assigned to treatment and control conditions has high internal validity but limited generalizability.

Selection Bias

Selection bias is systematic error introduced when the sample differs from the population in ways relevant to the research question, due to how participants were selected. Common forms include: self-selection bias (volunteers differ from nonvolunteers -- they tend to be more motivated, healthier, and more educated), survivorship bias (studying only those who survived or persisted, missing those who dropped out or died), healthy-user bias (in observational studies, people who engage in one healthy behavior also tend to engage in others, making it hard to isolate effects), and nonresponse bias (people who decline to participate differ systematically from those who agree). Attrition (participants dropping out over time) can also introduce selection bias in longitudinal studies if dropouts differ from completers. The MCAT often asks you to identify whether selection bias provides an alternative explanation for a study's results, or whether a study's sampling method prevents certain conclusions.

Representative Samples

A representative sample mirrors the population's distribution on relevant characteristics -- age, sex, race/ethnicity, socioeconomic status, education level, and any other variable relevant to the research question. Representatives is not all-or-nothing; a sample can be representative on some dimensions but not others. For example, a sample that matches the population on age and sex but overrepresents college graduates may produce valid conclusions about age/sex effects but biased conclusions about attitudes or behaviors correlated with education. Oversampling is a technique where researchers intentionally sample a subgroup at a higher rate than its population proportion to ensure enough statistical power for subgroup analyses; results are then weighted back to population proportions. Convenience samples (studying whoever is easiest to recruit, such as psychology undergraduates) are rarely representative of broader populations. The MCAT tests this by asking whether a study's sample demographics justify its claims about the general population.

External Validity

External validity is the degree to which a study's results generalize beyond the specific sample, setting, procedures, and time point of the original research. It has several dimensions: population validity (do results apply to people beyond the sample?), ecological validity (do results apply to real-world settings beyond the lab?), temporal validity (do results hold across time periods?), and treatment variation validity (do results hold with different operationalizations of the IV and DV?). High internal validity and high external validity often trade off: a tightly controlled laboratory experiment with random assignment maximizes internal validity but may have low ecological validity because the sterile lab setting does not resemble real-world conditions. Field experiments and naturalistic observation have higher ecological validity but lower internal validity. The MCAT tests this trade-off: a passage may describe a highly controlled experiment and then ask whether its conclusions apply to a real-world clinical setting.

Generalizability

Generalizability is the extent to which research findings from a specific study can be applied to other populations, settings, times, and operationalizations. It is the practical endpoint of external validity. Key threats to generalizability include: restricted samples (e.g., studying only one sex, one age group, one species), artificial settings (lab conditions that do not resemble real life), and narrow operational definitions (using a single measure of a broad construct). Replication across diverse samples, settings, and measures is the strongest evidence for generalizability. Meta-analyses that aggregate findings across many studies provide the most robust assessment of how general a finding is. The MCAT often presents a single study and asks you to identify which populations or settings the findings can and cannot be generalized to, based on the sampling method and study design. A common trap: assuming that a statistically significant finding in one sample automatically generalizes to all populations.

How it works

When evaluating a study's sampling on the MCAT, ask four questions: (1) What is the target population the researchers want to generalize to? (2) What is the actual sampling frame and method -- is it random, stratified, convenience, or self-selected? (3) Does the sample match the population on key characteristics, or is there evidence of selection bias? (4) Given the sampling, what populations and settings can the results legitimately generalize to? Remember: random sampling enables generalizability (external validity), while random assignment enables causal inference (internal validity). A study can be internally valid but not generalizable, or vice versa. The MCAT rewards precise language about which claims are supported by a given sampling method.

How it works

When evaluating a study's sampling on the MCAT, ask four questions: (1) What is the target population the researchers want to generalize to? (2) What is the actual sampling frame and method -- is it random, stratified, convenience, or self-selected? (3) Does the sample match the population on key characteristics, or is there evidence of selection bias? (4) Given the sampling, what populations and settings can the results legitimately generalize to? Remember: random sampling enables generalizability (external validity), while random assignment enables causal inference (internal validity). A study can be internally valid but not generalizable, or vice versa. The MCAT rewards precise language about which claims are supported by a given sampling method.

Comparisons

  • P/S (Research methods): The entire P/S section tests sampling concepts -- identifying whether a study used random sampling, representative samples, and whether findings are generalizable.
  • B/B (Experimental passages): Animal studies use specific strains and controlled environments; expect questions about whether findings from a mouse model generalize to humans.
  • C/P (Clinical research): Drug trials often use restricted samples (e.g., excluding pregnant women, children). Questions may ask about the implications for generalizability.
  • RM-002 (Variables and Controls): Random sampling (who is in the study) versus random assignment (which condition they get) is a high-yield distinction spanning both topics.
  • RM-003 (Study Design and Causality): Observational studies can achieve representativeness through large population-based sampling but cannot establish causation. Experiments establish causation but often use convenience samples.
  • RM-005 (Bias and Confounding): Selection bias is the primary bias connecting sampling to validity; it is covered in depth here and revisited in the bias topic.

Common confusions

  • Mistaking random assignment for random sampling: Random assignment assigns participants to conditions (internal validity). Random sampling selects participants from a population (external validity). They are independent design choices.
  • Assuming a large sample guarantees representativeness: A large biased sample is still biased. Sample size improves precision but does not correct for systematic selection bias.
  • Overgeneralizing from WEIRD samples: Most psychology research uses Western, Educated, Industrialized, Rich, and Democratic (WEIRD) samples. Findings may not generalize to other populations.
  • Confusing the sampling frame with the population: If the sampling frame is 'registered voters with listed phone numbers,' the population being sampled is NOT 'all adults' -- it is only those on the list.
  • Ignoring attrition in longitudinal studies: If 40% of participants drop out and dropouts differ from completers (e.g., sicker, less motivated), the remaining sample may no longer represent the original population.
  • Believing stratified sampling eliminates all selection bias: Stratification ensures balance on the stratification variables but does nothing for unmeasured confounds or variables not used in stratification.

Quick review

  • Population: entire group of interest; sample: subset actually studied.
  • Random sampling: every population member has a known, nonzero chance of selection.
  • Random sampling enables external validity; random assignment enables internal validity.
  • Selection bias: systematic error from how participants were chosen -- sample does not represent population.
  • Self-selection bias: volunteers differ from non-volunteers on motivation, health, education.
  • Sampling frame: the list or method from which the sample is drawn; must match the target population.
  • Representative sample: mirrors population on key characteristics (age, sex, SES, etc.).
  • Convenience sample: whoever is easiest to recruit; rarely representative.
  • Stratified random sampling: random sampling within predefined strata to ensure subgroup representation.
  • External validity: degree to which results generalize beyond the study's sample, setting, and procedures.
  • Ecological validity: whether results apply to real-world settings (versus lab conditions).
  • Generalizability: practical extent to which findings apply to other populations, settings, and times.
  • Attrition: participant dropout over time; threatens representativeness if dropouts differ from completers.
  • WEIRD samples: Western, Educated, Industrialized, Rich, Democratic; dominant in psychology research; limits global generalizability.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Imagine you want to know the average height of all students at a university with 30,000 people. You cannot measure everyone, so you take a sample. If you stand outside the gym and measure the first 100 people who walk out, your sample is biased -- it overrepresents athletes, who are likely taller than average. Your conclusion ('the average student is 6 feet tall') would be wrong. If instead you get a complete student roster (your sampling frame), assign every student a number, and use a random number generator to pick 100 students to measure, you have a random sample. On average, your sample will reflect the true student population height. But even a perfect random sample of this university's students cannot tell you the average height of students at another university, or of nonstudents the same age. That is the limit of generalizability: your findings apply to the population you sampled from, and extending beyond that requires new evidence. Limitation: real-world sampling is messier than random number generators. People decline to participate (nonresponse bias), rosters are incomplete, and the population itself changes over time. Even the best sampling plan produces an approximation, not a perfect mirror.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsRelated

Sources & references

  1. Psychology 2e - Chapter 2: Psychological Research — OpenStax
  2. MCAT Content Outline: Scientific Reasoning — AAMC
  3. Biology 2e - Chapter 1: The Study of Life — OpenStax
  4. Simply Psychology: Sampling Methods — Simply Psychology

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.