MCAT Foundations · Research Methods, Statistics, and Scientific Reasoning

Reliability, Validity, and Measurement

12 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 5 sections
  1. In 30 seconds
  2. The college version
  3. Eli explains
  4. Study tools
  5. Sources & references

In 30 seconds

Reliability and validity form the foundation of all scientific measurement. A study can only produce meaningful conclusions if its instruments measure consistently (reliability) and actually capture what they claim to measure (validity). The MCAT devotes the Research Methods section to these concepts because every passage-based question ultimately asks: are these findings trustworthy? Reliability comes in three forms: test-retest (stability over time), internal consistency (items measuring the same construct), and inter-rater (agreement between observers). Validity encompasses construct validity (does the test measure the right thing?), internal validity (did the IV truly cause the DV change?), and external validity (do results generalize?). The relationship is hierarchical: a measure can be reliable without being valid, but a valid measure must be reliable. On the MCAT, you evaluate studies by applying these lenses to determine whether reported conclusions are justified.

The college version

Reliability

Reliability is the consistency and reproducibility of a measurement. A reliable measure yields the same result under identical conditions. Formally, any observed score can be decomposed as: Observed Score = True Score + Measurement Error. Reliability reflects how much of the observed score is true score versus error. High reliability means low measurement error. Reliability is necessary but not sufficient for validity: a bathroom scale that consistently reads 10 pounds too high is reliable (consistent) but not valid (inaccurate). Reliability is quantified through correlation coefficients: test-retest correlations, Cronbach's alpha for internal consistency, and Cohen's kappa for inter-rater agreement. Key MCAT point: when a passage reports that a measure had 'good reliability' (r > 0.70 or alpha > 0.70), you know the instrument produces consistent measurements -- but you still need to ask whether it is measuring the right construct.

Test-Retest Reliability

Test-retest reliability assesses the stability of a measure over time. Researchers administer the same test to the same participants at two time points and correlate the scores. A high correlation (r > 0.70) indicates that the measure produces stable results. This method is appropriate for constructs expected to be stable (e.g., IQ, personality traits), not for constructs that naturally fluctuate (e.g., mood, hunger). Threats to test-retest reliability include: (1) practice effects -- participants improve on the second administration because they remember items; (2) maturation -- natural developmental changes between administrations; (3) carryover effects -- the first test influences responses on the second. The time interval is critical: too short risks practice effects; too long risks genuine change in the construct. Solution: use alternate (parallel) forms of the test for the second administration when practice effects are a concern. The MCAT often asks you to identify why a test-retest correlation is low and whether it indicates a problem with the measure or actual change in the construct.

Internal Consistency

Internal consistency measures whether all items within a single test tap the same construct. If a depression inventory has 20 items, all 20 should correlate with each other -- they are all measuring facets of depression. The most common statistic is Cronbach's alpha, which ranges from 0 to 1; values above 0.70 are considered acceptable, above 0.80 good, and above 0.90 excellent. Very high alpha (above 0.95) may indicate item redundancy rather than desirable consistency. Split-half reliability is an alternative method: split the test into two halves (e.g., odd vs. even items), correlate the halves, and apply the Spearman-Brown prophecy formula to estimate reliability for the full test. Low internal consistency suggests items are measuring different constructs, which threatens construct validity. The MCAT tests this by presenting a scale and asking you to evaluate whether its items form a coherent measure. Key pitfall: a high Cronbach's alpha does not guarantee the test measures what it claims -- it only means the items are internally consistent.

Inter-Rater Reliability

Inter-rater reliability (IRR) assesses the degree of agreement between two or more independent observers rating the same phenomenon. It is critical whenever measurement involves subjective judgment, such as coding behavioral observations, scoring essay responses, or diagnosing from clinical interviews. The appropriate statistic depends on the data type: Cohen's kappa for categorical ratings (corrects for agreement expected by chance), intraclass correlation coefficient (ICC) for continuous ratings, and percent agreement for simple binary judgments (though percent agreement does not correct for chance). A kappa of 0.60-0.80 indicates substantial agreement; above 0.80 is excellent. Threats to IRR include: unclear rating criteria, insufficient rater training, and rater drift over time. The MCAT tests IRR by asking whether disagreements between raters threaten the validity of study conclusions. For example, if two coders disagree 30% of the time when classifying participant responses, the study's dependent variable may be too unreliable to support its claims.

Construct Validity

Construct validity is the degree to which a test or measurement actually captures the theoretical construct it claims to measure. It is the overarching validity question: does this instrument measure intelligence, or does it measure test-taking skill? Construct validity is established through multiple lines of evidence: (1) Convergent validity -- the measure correlates strongly with other established measures of the same construct (e.g., a new IQ test should correlate with the WAIS-IV). (2) Discriminant (divergent) validity -- the measure does NOT correlate with measures of unrelated constructs (e.g., an IQ test should not correlate strongly with a measure of extraversion). (3) Content validity -- the test items comprehensively sample the construct's domain (e.g., a depression scale must cover mood, cognitive, somatic, and behavioral symptoms). (4) Face validity -- the test appears, on its surface, to measure what it claims (this is the weakest form of validity evidence). The MCAT tests construct validity by asking whether a study's operational definition truly captures the intended construct, and whether alternative operationalizations would yield different results.

Internal Validity

Internal validity is the degree to which a study can establish that the independent variable (IV) -- and only the IV -- caused the observed change in the dependent variable (DV). It answers: was the experiment sound? A study with high internal validity rules out alternative explanations. Key threats to internal validity include: (1) History -- events outside the study occurring between pre-test and post-test. (2) Maturation -- natural changes in participants (aging, fatigue, hunger). (3) Testing effects -- the pre-test itself changes post-test performance (practice or sensitization). (4) Instrumentation -- changes in the measurement instrument or procedure. (5) Regression to the mean -- extreme pre-test scores naturally move toward the average on re-test. (6) Selection bias -- pre-existing differences between groups. (7) Attrition (mortality) -- differential dropout across conditions. Defenses: random assignment (makes groups equivalent at baseline), blinding (prevents expectation bias), and standardized procedures (prevents instrumentation drift). Internal validity is the prerequisite for causal claims. The MCAT frequently presents a study with a plausible confound and asks whether the causal conclusion is justified -- this is a direct test of your ability to evaluate internal validity.

External Validity

External validity is the degree to which study findings generalize beyond the specific sample, setting, and time of the research. It answers: do these results apply to other people, places, and circumstances? Types of external validity include: (1) Population validity -- generalizing to other populations (e.g., do results from college sophomores apply to older adults?). (2) Ecological validity -- generalizing to real-world settings (e.g., does behavior in a sterile lab mirror behavior at home?). (3) Temporal validity -- generalizing across time periods (e.g., do findings from 1980 apply today?). The central trade-off in research design is between internal and external validity: tightly controlled lab experiments maximize internal validity (we know X caused Y) but often sacrifice external validity (the setting is artificial). Field studies and naturalistic observation maximize external validity but lose control over confounds. The MCAT tests this trade-off by asking you to identify which type of validity a design choice strengthens or weakens, and whether a study's limitations threaten its generalizability. A study can have perfect internal validity and near-zero external validity -- the MCAT wants you to recognize this distinction.

How it works

When evaluating measurement quality on the MCAT, apply the reliability-validity hierarchy: (1) Check reliability -- does the measure produce consistent results? Look for test-retest correlations, Cronbach's alpha, or inter-rater agreement statistics. If reliability is poor, validity cannot be high. (2) Check construct validity -- does the operational definition capture the intended construct? Look for convergent and discriminant validity evidence. (3) For experimental passages, evaluate internal validity -- are confounds controlled? Was random assignment used? Could alternative explanations account for the results? (4) Assess external validity -- to whom do these findings apply? Consider the sample, setting, and whether the trade-off with internal validity was worthwhile. This four-step framework -- reliability, construct validity, internal validity, external validity -- provides a systematic lens for every research methods question.

How it works

When evaluating measurement quality on the MCAT, apply the reliability-validity hierarchy: (1) Check reliability -- does the measure produce consistent results? Look for test-retest correlations, Cronbach's alpha, or inter-rater agreement statistics. If reliability is poor, validity cannot be high. (2) Check construct validity -- does the operational definition capture the intended construct? Look for convergent and discriminant validity evidence. (3) For experimental passages, evaluate internal validity -- are confounds controlled? Was random assignment used? Could alternative explanations account for the results? (4) Assess external validity -- to whom do these findings apply? Consider the sample, setting, and whether the trade-off with internal validity was worthwhile. This four-step framework -- reliability, construct validity, internal validity, external validity -- provides a systematic lens for every research methods question.

Comparisons

  • P/S (Research design): Every P/S passage that describes a study requires you to evaluate measurement reliability and validity; expect direct questions about Cronbach's alpha, test-retest reliability, and construct validity.
  • P/S (Individual differences): IQ tests, personality inventories (Big Five, MMPI), and clinical assessments all raise reliability and validity questions -- the MCAT tests whether you understand what a test score actually means.
  • B/B (Experimental passages): Internal validity is the core concern; passages ask whether observed biological effects are due to the treatment or a confound.
  • C/P (Measurement): Instrument calibration and precision are forms of reliability; accuracy is a form of validity -- the instrument must be both precise (reliable) and accurate (valid).
  • RM-002 (Variables and Controls): Controls, blinding, and randomization are the tools that establish internal validity.
  • RM-005 (Bias and Confounding): Confounds are the primary threats to internal validity; this topic provides the framework for evaluating whether they were adequately managed.
  • RM-004 (Sampling and Generalizability): Sampling methods determine external validity; random sampling supports population validity while convenience sampling limits it.

Common confusions

  • Reliability without validity: A measure can be perfectly reliable (consistent) but completely invalid (measuring the wrong thing). The MCAT loves presenting a consistent but mis-targeted measure and asking whether it is 'good.' The answer: reliable but not valid.
  • High Cronbach's alpha does not equal construct validity: A scale can have alpha = 0.95 and still measure the wrong construct. Internal consistency only tells you items cohere; it says nothing about whether they measure what they claim.
  • Confusing test-retest with inter-rater reliability: Test-retest is same measure, same people, different times. Inter-rater is same measure, same time, different raters. The MCAT will describe one and offer the other as a distractor.
  • Internal vs. external validity trade-off: The MCAT will describe a highly controlled lab study and ask whether it generalizes (no -- poor external validity), or describe a field study and ask about causation (cannot conclude causation -- poor internal validity). Know which validity a design choice optimizes.
  • Internal validity is about causation, not measurement: Internal validity asks 'did X cause Y?' not 'did we measure Y correctly?' Construct validity asks the measurement question. Students conflate these on the MCAT.
  • Face validity is the weakest evidence: A test 'looking like' it measures something (face validity) is not a strong argument for construct validity. The MCAT expects you to prioritize convergent and discriminant validity evidence.
  • Regression to the mean as a confound: When a study selects participants based on extreme scores and then re-tests them, improvement may be statistical regression, not a treatment effect. The absence of a control group makes this indistinguishable from a treatment effect.

Quick review

  • Reliability = consistency/reproducibility of measurement. Necessary but not sufficient for validity.
  • Test-retest reliability = same test, same people, different times. Watch for practice effects and maturation.
  • Internal consistency = items within a test measure the same construct. Assessed by Cronbach's alpha (target > 0.70).
  • Inter-rater reliability = agreement between independent observers. Assessed by Cohen's kappa or ICC.
  • Construct validity = does the test measure the theoretical construct? Evidenced by convergent and discriminant validity.
  • Internal validity = did the IV (and only the IV) cause the DV change? Requires confound control and random assignment.
  • External validity = do findings generalize to other populations, settings, and times? Trades off with internal validity.
  • Reliable does NOT imply valid (consistent but wrong target). Valid DOES imply reliable (must be consistent to hit target).
  • Face validity = weakest form; content validity = items cover domain; convergent/discriminant = strongest evidence.
  • Cohen's kappa corrects inter-rater agreement for chance; percent agreement does not.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Imagine you are practicing darts. Reliability is how tightly your darts cluster together. Throw three darts: they all land in a tight group in the upper-left corner. Your throws are reliable -- you are consistent. But the bullseye is at the center, so your accuracy (validity) is terrible. Now imagine you adjust your aim and all three darts cluster tightly around the bullseye. Now you are both reliable AND valid. In research, test-retest reliability is like throwing darts on Monday and Wednesday -- do they cluster in the same spot? Internal consistency is like checking whether all three darts in one round land together (coherence within a single test). Inter-rater reliability is like having two different people score your dart throws -- do they agree on where each dart landed? Construct validity asks: is the dart board even the right target for measuring dart skill, or should we be measuring something else entirely? Internal validity asks: did your new throwing technique actually cause the improvement, or was the wind different that day? External validity asks: do your bullseyes on this practice board in your garage predict your performance at a tournament? Limitation: real measurement is messier than darts. Psychological constructs like 'intelligence' or 'depression' do not have a visible bullseye -- the target itself is defined by theory, and different theories draw the bullseye in different places. That is why construct validity is never fully settled; it is always a matter of accumulating evidence.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsRelated

Sources & references

  1. Psychology 2e - Chapter 2: Psychological Research (Sections 2.2-2.3: Approaches to Research, Analyzing Findings) — OpenStax
  2. MCAT Content Outline: Scientific Inquiry and Reasoning Skill 3 (Research Methods and Scientific Reasoning) — AAMC
  3. Reliability In Psychology Research: Definitions and Examples — Simply Psychology
  4. Validity In Psychology Research: Types and Examples — Simply Psychology

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.