Education · Assessment

Reliability and Validity in Assessment

Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Quick check
  8. Study tools
  9. Sources & references

In 30 seconds

Two questions sit underneath every test. Reliability asks whether the scores would come out the same if you ran the measurement again. Validity asks whether the meaning you attach to those scores, and the decision you make with them, is actually supported by evidence. Most courses teach the second one wrongly. Validity is not something a test owns. It is a judgment about one interpretation, for one use, and somebody has to argue for it.

Why this matters

Every claim a test makes about a person is an inference, and these two ideas are how the profession checks inferences. Reliability tells you how much of a score is noise, which is the difference between a defensible cut score and a coin flip near the boundary. Validity tells you whether the number means what the score report says it means. The vocabulary is also your entry point to the rest of measurement: item review, fairness analysis, standard setting, program evaluation, and every research paper that reports a coefficient alpha in its methods section. Learning the current framing now saves you from the most expensive error in the field, which is treating a validation study done for one purpose as a license to use the same scores for a different one.

The college version

Validity is a claim about an interpretation, not a property of a test

The 2014 Standards for Educational and Psychological Testing, the joint AERA, APA and NCME document the field treats as its reference point, define validity as the degree to which evidence and theory support the interpretations of test scores for proposed uses. Two words carry the weight: interpretations and uses. What gets evaluated is the interpretation, not the instrument. The Standards say outright that the unqualified phrase 'the validity of the test' is incorrect, and their opening standard asks developers to articulate each intended interpretation for a specified use and supply evidence for each one separately. The practical consequence is the part students usually miss. One mathematics test can be read three ways: this student has mastered the year's curriculum, this student would benefit from a particular course, this student is likely to succeed at college-level work. Those are three different claims, and each needs its own evidence. Evidence that supports course placement does not license using the same scores to evaluate the teacher. Validation starts with a , the characteristic the test is meant to measure, and then asks what would have to be true for the intended reading of the scores to hold. Two standing threats organize that search. is a test that misses parts of what it claims to cover. is scores being pushed around by something outside the construct: reading load on a science item, writing speed on a test of historical reasoning, familiarity with an unusual response format.

Five sources of evidence, and what the old three types were reaching for

Older textbooks hand you a menu: content validity, criterion validity in predictive and concurrent flavors, and construct validity. That taxonomy comes from the 1954 APA Technical Recommendations, and Cronbach and Meehl's 1955 paper is the canonical statement of the construct variety, invoked whenever a test is read as measuring an attribute that is not operationally defined. Messick spent the following decades arguing that these were never separate certificates but complementary evidence bearing on a single question: does the score mean what we say it means, and does the proposed use follow. The 2014 Standards adopt that unified position explicitly. Validity is a unitary concept, and the document deliberately abandons the historical nomenclature in favor of types of validity evidence. Five sources are named. Test content covers the fit between the themes, wording and format of the items and the specified content domain. Response processes ask whether test takers actually engage the processes the construct assumes, so a claim about mathematical reasoning is weakened if students are matching a memorized algorithm; this source also covers whether judges apply the criteria they are meant to. Internal structure asks whether the relationships among items match the structure the interpretation assumes, one dimension or several, and analyses live here. Relations to other variables covers convergent, discriminant and criterion evidence, which is where the old criterion-related category went. Consequences of testing is the one to get right, because the Standards take a narrower line than the phrase suggests: consequences bear on validity when they can be traced to construct underrepresentation or construct-irrelevant variance. Untraceable consequences may still be excellent reasons not to use a test, but they are value judgments about use rather than validity evidence. None of the five sources is compulsory: which evidence you need depends on which propositions your interpretation rests on.

Reliability: consistency across replications, and where the error comes from

The Standards pair reliability with the word precision and define it as consistency of scores across replications of the testing procedure. That makes 'what counts as a replication?' the first question rather than a technicality. For an attribute not expected to change, two administrations a day apart are replications. For a state such as mood they are not, because the difference between them is real change rather than error. For a knowledge test, two forms built to the same specification count; for an inventory whose wording is fixed, a reworded form is a different test. Classical test theory formalises this as observed score equals plus error, where the true score is the hypothetical average over infinitely many independent replications and the reliability coefficient is the proportion of observed score variance that is true score variance. Generalizability theory splits error into components for items, occasions and raters, so you can see which source is costing you. Item response theory handles the same problem through information functions, which give precision at each point on the scale. One consequence mirrors the point about validity. Reliability is not a fixed attribute of an instrument; it is a property of scores obtained from a particular group under particular conditions. Nimon, Zientek and Henson put it directly: reliability inheres in scores, not in tests, so quoting a manual's coefficient as though it described your own data is an error. The Standards agree from the other direction, calling blanket statements that a test 'is reliable' rarely, if ever, acceptable documentation.

Four families of reliability evidence, and why they are not interchangeable

Test-retest uses the same form on two occasions and captures instability across occasions, including practice and memory effects. Parallel or alternate forms uses two forms in independent sessions and captures the error introduced by sampling different items. works from the relationships among items within a single administration: split-half with a Spearman-Brown adjustment back up to full length, KR-20 for dichotomously scored items, coefficient alpha in general, with the Standards' glossary treating KR-20 as simply the dichotomous special case of alpha. Inter-rater agreement is required whenever human judgment enters scoring. These indices are not substitutes for one another, and the Standards say so directly: each defines measurement error differently, so a high alpha tells you nothing about stability across days and a high test-retest correlation tells you nothing about whether the items hang together. Reporting one and letting readers assume the others is a routine failure. Inter-rater agreement needs particular care. Raw percent agreement flatters the scoring, because two raters who never spoke will agree a great deal by accident, especially when one category is common. discounts that: kappa equals observed agreement minus expected chance agreement, divided by one minus expected chance agreement, where the expected value comes from the raters' own marginal rates. That correction is why kappa is preferred. It is not a cure-all, though. Because expected agreement depends on the marginal distribution, the same percent agreement can produce very different kappas depending on how lopsided the categories are, which is precisely the information percent agreement was concealing.

The standard error of measurement, and the band around a score

A reliability coefficient describes a group. The brings the same information down to an individual score and expresses it in the score's own units. Under classical test theory it is the standard deviation of scores multiplied by the square root of one minus the reliability coefficient. Conceptually it estimates how much one person's observed score would bounce around across repeated administrations, which lets you report a band rather than a point. That formula also explains something that surprises people. Reliability depends on the spread of the group being measured, so a selective population with a narrow ability range yields a low reliability coefficient even when the measurement is exactly as precise as before. Tighe and colleagues demonstrated this on the MRCP(UK) postgraduate medical examinations: restricting the sample to candidates who had already passed dropped reliability from about .90 to about .70 while the standard error of measurement barely moved. Their conclusion is worth carrying around: a reliability coefficient describes an assessment together with the particular people who sat it, whereas the standard error behaves much more like a property of the measuring instrument itself. The Standards add one refinement. A single average standard error hides the fact that precision varies along the scale, and conditional standard errors at each score level are more informative, especially in the region where a cut score concentrates decisions. And a band is a statement about how scores from this procedure vary, not a probability statement about where one person's unobservable true score sits.

What coefficient alpha does and does not tell you

Alpha is the most reported and most misreported statistic in educational measurement, and two claims routinely made for it are false. First, alpha is not a measure of unidimensionality. Sijtsma showed that alpha is a function of the number of items and the average inter-item covariance and nothing else, so unidimensional item sets can produce almost any alpha and multidimensional sets can produce identical ones. 'Alpha was .88, confirming the scale is unidimensional' is not a weak inference; it is a non sequitur. Dimensionality is an internal-structure question needing a factor analysis. Second, alpha is not generally equal to reliability. It equals reliability only under , meaning every item relates to the true score with the same weight, differing at most by an additive constant. Real item sets rarely satisfy that. Under the usual classical assumptions, including errors that are uncorrelated across items, alpha is a lower bound and so tends to understate reliability; the simulation work of Trizano-Hermosilla and Alvarado puts the underestimate in the region of one to eleven per cent as tau-equivalence violations worsen. If item errors are correlated, as they are with shared reading passages or chained items, even the lower-bound guarantee lapses. McDonald's omega, computed from a factor model, does not require tau-equivalence and is generally the better default; where items are strongly skewed, the greatest lower bound performs better than either. The practical reading for a student working through a methods section: an alpha is a floor estimate resting on assumptions the authors almost certainly did not test.

How reliability constrains validity without ever guaranteeing it

The textbook line, that a test can be reliable without being valid, is true and so compressed that it teaches the wrong lesson. Unpack it in both directions. Reliability constrains. Random error attenuates relationships, and under classical assumptions the correlation between two measures cannot exceed the square root of the product of their reliabilities. Scores with reliability .64 therefore cannot correlate above .80 with any criterion, however perfectly that criterion is measured, and cannot exceed .72 with a criterion whose own reliability is .81. Both ceilings were confirmed by simulation while this lesson was written. Noise also widens every score band, which is where it bites on classification. The Standards make the point without algebra: inconsistent scores limit accurate prediction, diagnosis and decision making. The bound does assume errors are uncorrelated, and Nimon and colleagues show that when errors correlate with true scores or with each other an observed correlation can exceed its true-score counterpart, so treat it as a property of the classical model rather than a law of nature. Reliability does not guarantee. Reliability indices are blind to systematic error. A test that consistently rewards fast reading while claiming to measure science reasoning will look excellent on every reliability coefficient you compute, because construct-irrelevant variance is stable rather than random, and stability is what those coefficients reward. Consistency tells you the measurement repeats; it says nothing about what is repeating. And the constraint is a ceiling, not a threshold. No reliability figure exists above which an interpretation becomes valid. The Standards ask for precision appropriate to the stakes, allowing less where a decision is reversible or corroborated elsewhere, and leave the judgment to the user.

Fairness as part of the validity argument

The 2014 Standards treat fairness as a foundation alongside validity and reliability and define it in validity terms: design every step of the process to minimize construct-irrelevant variance so that score interpretations hold for all examinees in the intended population. Three ideas carry most of the load. Differential item functioning occurs when test takers who are at the same level on the measured attribute, but belong to different groups, answer a particular item correctly at systematically different rates. DIF is a statistical flag rather than a verdict; the Standards insist that calling an item biased requires a substantive explanation of why it behaves that way, and note that DIF sometimes reveals genuine multidimensionality the test framework anticipated. Accessibility and universal design push the work earlier: build items from the outset so that irrelevant barriers, such as unnecessary reading load, formats that defeat magnification, or language demands beyond the construct, never enter. Accessibility is a legal requirement in some testing contexts. Accommodations and modifications are distinguished in the Standards by what happens to score meaning. An accommodation preserves comparability, so the accommodated scores support the same inferences; a modification changes the construct measured, so the scores differ in meaning from the standard form. Reading items aloud is an accommodation on a mathematics test and a modification on a test of decoding. The line between them is a validity line, and it decides whether the resulting scores belong on the same scale as everyone else's.

Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Picture two different complaints about a bathroom scale. First complaint: you step on it three times in a minute and get 152, 149, 154. That is a consistency problem. The scale is noisy, and no single reading can be trusted on its own. Second complaint: you step on it once, get a steady 151 every single time, and then somebody uses that number to decide whether you are healthy. That is a validity problem, and notice it has nothing to do with noise. The number is rock solid. What is wrong is the meaning someone attached to it and the decision they made on the strength of it. Almost every muddle in this topic comes from mixing up those two complaints. Consistency is about the measuring. Validity is about the claim.

Picture it like this

So treat validity as a court case rather than a certificate. Nobody hands a test a validity badge that it carries around forever. Somebody has to stand up and argue: here is what I say this score means, here is the use I want to put it to, and here is my evidence, that the content matches the domain, that students reason rather than pattern-match, that the items hang together the way I claimed, that the scores predict what I promised. New use, new case, new evidence.

Where the picture stops working

The courtroom picture oversells the finality. A verdict closes a case, whereas validation never quite closes, because new evidence can reopen it and a test can drift out of validity as its population or its curriculum changes. It also suggests a single visible opponent, when the usual failures are quiet ones: a construct drawn too narrowly at the start, or an irrelevant demand nobody thought to look for. And a court returns guilt or innocence, while validity is a matter of degree. The honest output is a level of support, weighed against how much the decision costs when it goes wrong.

Worked example

Two calculations, both executed rather than quoted. First, precision. A district reading test reports scale scores with a standard deviation of 100 in the tested grade and a reliability coefficient of .91. The standard error of measurement is 100 times the square root of (1 minus .91), which is 100 times the square root of .09, which is 30 scale points. A student scores 512 and the proficiency cut sits at 500. A one-standard-error band runs from 482 to 542; a 1.96-standard-error band runs from about 453 to 571. Both straddle the cut. The reported score is above 500 but the measurement is not precise enough to establish that the student's standing is, and the honest report of that fact is the band, not the point. Two cautions travel with it: the band describes how scores from this procedure vary, not the probability that a true score sits inside it, and a single average standard error understates the error at some score levels, so the conditional standard error near the cut is the number that actually matters. This is a statement about measurement, not advice about what to do with any individual student. Second, agreement. Two raters independently mark 100 essays pass or fail. Rater A passes 80, rater B passes 85, and the two agree on 87 of the 100. Eighty-seven per cent sounds strong. But chance agreement is (.80 times .85) plus (.20 times .15), which is .71, because both raters pass most essays and would collide constantly even while guessing. Cohen's kappa is (.87 minus .71) divided by (1 minus .71), which is .16 divided by .29, which is .55, weak agreement on McHugh's scale. The scoring is far less consistent than the headline number suggested, and the reason is sitting in the margins: when one category dominates, percent agreement is close to worthless.

Key takeaway

Reliability is how much of a score survives a repeat of the measurement; validity is whether the meaning you attach to that score, for the use you have in mind, is supported by evidence. Unreliability caps how much validity evidence you can ever gather, and no amount of reliability makes an interpretation valid.

Quick check

3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.

Question 1 of 3foundational

Which statement matches how the 2014 Standards for Educational and Psychological Testing define validity?

Choose an answer, then check it.
Question 2 of 3intermediate

A journal article states: 'Coefficient alpha was .89, confirming that the scale is unidimensional.' What is wrong with the inference?

Choose an answer, then check it.
Question 3 of 3intermediate

Scores on a test have a standard deviation of 60 and a reliability coefficient of .75. What is the standard error of measurement, and what does a one-standard-error band around a reported score of 300 span?

Choose an answer, then check it.
Practice all 5

Keep learning

Ready to build on this? Continue to the next lesson.

Practice this lesson
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related

You’ll learn to

  • Define reliability as consistency of scores across replications of a testing procedure, and validity as the degree to which evidence and theory support a specific score interpretation for a proposed use.
  • Distinguish the five sources of validity evidence in the current Standards from the superseded content/criterion/construct taxonomy, and explain what the older terms were reaching for.
  • Apply the standard error of measurement to build and interpret a band around a reported score, and compute Cohen's kappa from a rater agreement table.
  • Explain why unreliability places a ceiling on validity evidence without any level of reliability guaranteeing a valid interpretation.
  • Evaluate which reliability index fits a given source of measurement error, including why kappa is preferred to raw percent agreement and why coefficient alpha is weaker evidence than it is usually taken to be.
  • Analyze fairness evidence, including differential item functioning, accessibility, and the accommodation/modification distinction, as part of a validity argument.

Common mistakes

  • Saying a test 'is valid' or asking whether an instrument 'has validity'.

    Validity attaches to a specific interpretation of scores for a specific use, supported by evidence. The 2014 Standards call the unqualified phrase incorrect, and require each intended interpretation to be validated separately, so a test validated for course placement is not thereby validated for teacher evaluation.

  • Reporting coefficient alpha as evidence that a scale is unidimensional.

    Alpha depends only on the number of items and the average inter-item covariance, so it cannot distinguish a unidimensional set from a multidimensional one. Dimensionality is internal-structure evidence and needs a factor analysis; alpha at best estimates a lower bound on reliability, and only under assumptions almost nobody checks.

  • Treating high percent agreement between raters as proof that scoring is consistent.

    Raters agree by accident, and the more lopsided the categories the more often they do. Cohen's kappa subtracts the agreement expected from the raters' own marginal rates; 87 per cent agreement can correspond to a kappa near .55, and the gap between the two numbers is the chance agreement percent agreement was hiding.

  • Concluding from 'a test can be reliable but not valid' that reliability is irrelevant to validity.

    Unreliability caps the evidence you can gather: under classical assumptions a correlation cannot exceed the square root of the product of the two reliabilities, so scores with reliability .64 cannot correlate above .80 with any criterion. What the slogan really says is that reliability indices are blind to systematic error, so a test can consistently measure the wrong thing.

  • Quoting the reliability coefficient printed in a test manual as the reliability of your own data.

    Reliability is a property of scores from a particular group under particular conditions, not a fixed attribute of an instrument. A narrower ability range lowers the coefficient even when precision is unchanged, which is why the standard error of measurement often travels better across populations than the coefficient does.

Easily confused

Reliability vs. Validity

Reliability asks whether scores repeat across replications of the testing procedure; validity asks whether a particular interpretation of those scores, for a particular use, is supported by evidence. Reliability is estimated with coefficients and standard errors; validity is argued from several sources of evidence and never certified once and for all.

Percent agreement between raters vs. Cohen's kappa

Percent agreement counts the times two raters matched. Kappa subtracts the matches expected from their marginal rates before dividing by the room left for improvement, so it answers how much agreement exceeded chance. The two diverge most when one category dominates, which is exactly when percent agreement is most flattering.

Coefficient alpha vs. McDonald's omega

Alpha equals reliability only under essential tau-equivalence and is otherwise a lower bound, typically understating it. Omega is computed from a factor model that allows items to load unequally, so it does not need that assumption and is generally the better default; with strongly skewed items the greatest lower bound outperforms both.

Random error vs. Construct-irrelevant variance

Random error is unpredictable fluctuation and is what reliability coefficients detect. Construct-irrelevant variance is a systematic pull from something outside the construct, so it is stable across replications, invisible to reliability indices, and can even raise them while corrupting the interpretation.

The 1954-era types of validity vs. The 2014 sources of validity evidence

The older scheme named content, criterion and construct validity as if a test could hold one and lack another. The Standards now treat validity as unitary and list five sources of evidence, of which criterion evidence is one strand of relations to other variables, so a test cannot possess a type of validity, only accumulate evidence for a reading of its scores.

Key vocabulary

Construct
The characteristic or concept a test is designed to measure, such as reading comprehension or mathematical reasoning; naming it precisely is the first step of validation.
Construct underrepresentation
A test's failure to capture important parts of the characteristic it claims to assess, which narrows what its scores can legitimately be taken to mean.
Construct-irrelevant variance
Systematic influence on scores from something outside the intended characteristic, such as reading load on a science item or writing speed on a reasoning task.
True score
In classical test theory, the hypothetical average of a person's scores over infinitely many independent replications of the same testing procedure.
Standard error of measurement
An estimate, expressed in score units, of how far an individual's observed score would typically fall from the average of repeated administrations.
Internal consistency
The degree to which items within one administration agree with each other, used as an indirect estimate of how much scores would vary across equivalent forms.
Essential tau-equivalence
The assumption that every item relates to the underlying true score with the same weight, differing at most by an additive constant; coefficient alpha equals reliability only when it holds.
Cohen's kappa
An agreement index for categorical judgments that subtracts the agreement two raters would reach by chance from the agreement they actually reached.
Differential item functioning
A pattern in which examinees of equal standing on the measured attribute but different group membership answer one item correctly at systematically different rates.

Sources & references

  1. Standards for Educational and Psychological Testing (2014 edition) — Joint Committee on the Standards for Educational and Psychological Testing of the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education; published by AERA
  2. Validity of Psychological Assessment: Validation of Inferences From Persons' Responses and Performances as Scientific Inquiry Into Score Meaning (ETS Research Report RR-94-45) — Samuel Messick, Educational Testing Service; ERIC full text ED380496 (also published in American Psychologist, 50(9), 741-749, 1995)
  3. Construct Validity in Psychological Tests (Psychological Bulletin, 52, 281-302) — Lee J. Cronbach and Paul E. Meehl; Classics in the History of Psychology, an internet resource developed by Christopher D. Green, York University
  4. On the Use, the Misuse, and the Very Limited Usefulness of Cronbach's Alpha (Psychometrika, 74(1), 107-120) — Klaas Sijtsma; PubMed Central (PMC2792363)
  5. Interrater Reliability: The Kappa Statistic (Biochemia Medica, 22(3), 276-282) — Mary L. McHugh; PubMed Central (PMC3900052)
  6. The Standard Error of Measurement Is a More Appropriate Measure of Quality for Postgraduate Medical Assessments Than Is Reliability: An Analysis of MRCP(UK) Examinations (BMC Medical Education, 10:40) — Jennifer Tighe, I. C. McManus, Nicholas G. Dewhurst, Liliana Chis and John Mucklow; PubMed Central (PMC2893515)
  7. The Assumption of a Reliable Instrument and Other Pitfalls to Avoid When Considering the Reliability of Data (Frontiers in Psychology, 2012) — Kim Nimon, Linda Reichwein Zientek and Robin K. Henson; PubMed Central (PMC3324779)
  8. Best Alternatives to Cronbach's Alpha Reliability in Realistic Conditions: Congeneric and Asymmetrical Measurements (Frontiers in Psychology, 7:769) — Italo Trizano-Hermosilla and Jesus M. Alvarado; Frontiers Media

EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.

Researched 2026-08-18

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.