Education · Assessment

Rubrics and Performance Assessment

Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Quick check
  8. Study tools
  9. Sources & references

In 30 seconds

A asks the student to actually do the thing: write the argument, run the investigation, play the piece, design the circuit. Someone then judges what was produced. A rubric is the judging instrument, and it has two parts: criteria naming the qualities that matter, and descriptions of what each criterion looks like at several levels of quality. Rubrics make expectations visible and scoring steadier, but they cannot rescue a badly chosen task or a criterion that counts features instead of describing quality.

Why this matters

Most of what schools claim to teach cannot be shown by choosing option C. Writing, argument, laboratory work, clinical reasoning, and performance are constructs you can only observe by watching someone do them, which means the score depends on human judgment and therefore on how that judgment is structured. Learning to build a rubric well is the difference between feedback a student can act on and a number they can only accept. The same skill transfers past the classroom: hiring rubrics, clinical assessment forms, peer review criteria, and grant scoring sheets all live or die on whether their descriptors describe quality or merely count things.

The college version

Why performance assessment exists

Some things a student knows can be inferred from a well-built multiple-choice item. Others cannot. If the claim you want to make is that a student can construct a historical argument from conflicting documents, design an experiment that isolates a variable, or hold a coherent conversation in a second language, no selected-response item elicits that behavior; at best it samples a fragment of the knowledge the behavior draws on. Grant Wiggins put the contrast sharply in 1990: conventional testing relies on indirect proxy items from which we hope to infer performance, while a performance assessment examines the performance directly. If the outcome you value is problem-posing, experimental investigation, document-based inquiry, or the sustained revision of a piece of writing until it works for a reader, then, he argued, the assessment should be built out of those challenges. The practical consequence is that the score no longer comes from an answer key. It comes from a person looking at a product or a performance and making a judgment. Everything difficult about performance assessment follows from that one fact.

Authentic assessment, and an honest caveat

Authenticity is a claim about the task, not about the score. An authentic task resembles the way the skill is actually used: the challenge is ill-structured, the student has room to plan and revise, the audience and purpose are real enough to matter, and the performance obligations are known in advance so the student can rehearse. Wiggins made that last point deliberately, contrasting it with tests whose validity depends on secrecy. It is a genuinely good design goal. It is not, by itself, evidence that your interpretation of the score is warranted. A task can look exactly like professional practice and still produce scores that mostly reflect the particular prompt, the particular rater, or how much background knowledge the scenario assumed. Authenticity is an argument for building the task a certain way; validity is a separate argument you have to make with evidence about how the scores behave. Treat the two as different conversations, and be suspicious of any claim that a task is valid because it is realistic.

Rubric anatomy, and what a rubric is not

Susan Brookhart's working definition is the one to hold onto: a rubric has two parts, criteria that say what to look for in the work, and performance level descriptions that say what each criterion looks like in work of varying quality. Criteria are the qualities; levels are the points along the continuum; descriptors are the sentences that populate the cells. Take away the descriptions and you no longer have a rubric. A checklist asks a yes-or-no question about each criterion: the citation is present or it is not. A rating scale keeps the criteria but replaces the descriptions with a bare scale, whether numeric (1 to 5), evaluative (excellent, good, fair, poor), or frequency-based (always, usually, sometimes, never). All three tools can be used defensibly for some purpose, and all three have criteria. The difference is that only the rubric tells a student what better work would actually look like. In Brookhart's review of 51 rubrics drawn from 46 higher-education studies, several of the tools authors called rubrics turned out, on inspection, to be rating scales or point schemes.

Analytic, holistic, and single-point

An judges each criterion separately, producing a profile: strong on evidence, weak on organization. That separation is what makes analytic rubrics good for feedback, and it is why they dominate the research literature, accounting for 29 of the 51 rubrics in Brookhart's review. A requires one judgment across all criteria at once. It is faster and less cognitively demanding, which suits high-volume grading and situations where nobody will act on the detail, but it cannot tell a student which quality to work on. A lists only the proficient level for each criterion, with an open column on either side: one for evidence that the work falls short, one for evidence that it goes beyond. Jarene Fluckiger proposed it as a way to put students into goal setting and self-assessment rather than level-matching. Its cost is real: because the shortfall and excellence columns are written fresh for each student, it takes more of the rater's time and gives less scaffolding for agreement between raters. Choose the form by the decision you need to support, not by habit.

Writing descriptors that describe quality

The classic failure is a rubric whose levels differ only in quantity: many errors, some errors, few errors, no errors. It looks rigorous and teaches nothing. A student told they have some errors learns neither what an error is nor what the next level requires, and two raters can only agree if they already share an unstated definition of the counting unit. The same failure wears other costumes. Levels that differ by effort words (minimal, adequate, thorough) rate the student rather than the work. Criteria that restate the assignment directions describe compliance rather than learning; Brookhart's example contrasts a criterion reading has three sources with one reading uses a variety of relevant, credible sources. A criterion of the first kind can be satisfied by a student who understood nothing. Her review found seven of the 51 rubrics used rating-scale language or counted occurrences instead of describing quality, and she judged those tools more useful for grading than for learning. The working test for a descriptor is simple: could a student read it and know what to do differently tomorrow?

Rating is a measurement problem

Once a person assigns the score, the person becomes part of the instrument. Measurement research names the recurring patterns. Severity and leniency are the tendencies to score consistently below or above what the performance warrants. Centrality is the retreat to the middle categories, which flattens real differences. The halo effect, first noticed in rating research over a century ago, is letting a general impression or a judgment on one trait carry into the others; it quietly collapses an analytic rubric back into a holistic one, and because it makes ratings repetitious it inflates reliability estimates rather than lowering them, which is why agreement statistics alone will not catch it. Misfit is simply inconsistency: scoring that does not follow a stable pattern. These are not rare. In Stefanie Wind's simulation study, having roughly ten percent of raters display an effect produced substantial changes in how students were classified, and severity did more damage than centrality or inconsistency. The remedies are procedural rather than clever. Train raters on the rubric using anchor performances that have been pre-scored. Moderate: have raters score a common set independently, then discuss discrepancies until the interpretation converges. Monitor continuously, because trained raters drift. In NAEP, supervisors backread a target of at least five percent of each scorer's daily work to check that the guide is still being applied as written, and trainers use calibration sets to pull teams back to the standard.

The threat that is easy to miss: task sampling

It is tempting to think the main risk is disagreement between raters. The evidence points elsewhere. When Richard Shavelson, Xiaohong Gao, and Gail Baxter analyzed elementary mathematics and science performance assessments with generalizability theory, the variance components for raters and for rater interactions were zero or negligible once raters were trained; one well-trained rater was essentially enough. The dominant source of error was the task. Students who did well on one hands-on investigation did not reliably do well on the next, so a score built on a single task generalizes poorly to the domain the task was supposed to represent. With one rater and one task, the relative generalizability coefficients in that work were 0.15, 0.21, and 0.32 across three data sets, and reaching roughly .80 was estimated to need about 23 tasks in one study, 15 in another, and 8 in a third. Those numbers belong to those specific assessments, not to every performance task ever written, but the direction of the finding has held up. The practical lesson is not to abandon performance assessment; it is to stop making confident claims about a student's general ability from one performance, and to accumulate evidence across several tasks before treating a score as a stable estimate.

Specificity versus construct validity

There is a real trade-off hiding inside rubric design. Jonsson and Svingby's review of 75 studies found that scoring reliability improves when rubrics are analytic, topic-specific, and paired with exemplars or rater training. Push that logic far enough and you get a rubric that specifies exactly which facts, moves, and sentences a response should contain, at which point raters agree almost perfectly and the assessment has quietly changed. A task-specific rubric that lists the required content cannot be shared with students in advance without giving away the answer, and it converts a complex performance into a compliance checklist. The construct you are now measuring is the ability to follow a specification, not the ability to reason. General rubrics apply to a family of tasks, which is what lets them be shared, taught, and reused, but they leave more room for raters to disagree. The design question is not how specific can I be, but what is the least specification that gets raters to agree without changing what counts as good work.

Rubrics belong in the room before the performance

A rubric handed out with the grade is a receipt. A rubric handed out before the work is an instructional document, and that is where its value comes from. Panadero and Jonsson's review of 21 studies of formative rubric use proposed several routes by which rubrics can help: they make expectations transparent, reduce anxiety, structure feedback, support self-efficacy, and give students something concrete to self-regulate against. Effects, though, depend on implementation, not on the document. In studies where rubrics were simply distributed without teaching students to use them, there was typically no significant effect on performance. Used well, the rubric becomes the shared language for self-assessment (score your own draft, then justify the level) and peer assessment (name the criterion, cite the evidence). Two cautions are worth carrying. Brookhart found that only about 56 percent of studies reported using the rubric with students at all, so much of what is published as rubric research is really grading-tool research. And every study in her review reported positive outcomes, with no relationship between rubric quality and results, which she attributed partly to publication bias. Rubrics are worth building carefully; the literature supporting them is more enthusiastic than it is precise.

Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Imagine you want to know whether someone can bake. You could give them a written quiz about flour and ovens, or you could ask them to bake a cake. The quiz is quick and easy to mark, but it never shows you a cake. Performance assessment is the second choice: make the person do the real thing. The problem with the real thing is that there is no answer key. Somebody has to taste the cake and decide how good it is, and people disagree. A rubric is how you organize that disagreement. First you decide what actually matters about a cake, maybe texture, flavor, and how it looks. Those are your criteria. Then, for each one, you write down what a great version looks like, what an okay version looks like, and what a poor version looks like. Those descriptions are the important part. Writing three errors, two errors, one error is not a rubric; it is counting. And here is the trick most people miss: you show the rubric to the baker before they bake, not after.

Picture it like this

A rubric is like the score sheet a diving judge uses. The judges are not making it up as they go. They agreed beforehand what counts, they practiced on recorded dives until their scores lined up, and the divers know the standards too.

Where the picture stops working

The analogy breaks in three places. Diving judges score the same narrow skill on a defined list of dives, so one dive tells you a lot; a school performance task samples a huge domain, so one essay tells you much less about writing in general than one dive tells you about diving. Diving scores exist to rank competitors, while a classroom rubric mostly exists to tell a learner what to do next. And a diving judge never has to worry that spelling out the standard too precisely would change the sport, whereas an over-specified rubric really can turn thinking into box-ticking.

Worked example

A biology department is unhappy with its lab report rubric. The Data Analysis criterion has four levels: Many errors, Some errors, Few errors, No errors. Diagnose it before rewriting. First, it counts rather than describes, so a student at Few errors learns nothing about what to do next. Second, agreement between raters rests on an undefined unit; one instructor counts a wrong test as one error, another counts every affected sentence. Third, the ceiling is the absence of mistakes, so a cautious student who runs the simplest possible analysis outscores one who attempted something harder. Now rewrite each level as observable performance. Level 4: chooses a test that matches the design and data type, states the assumptions it requires, and interprets the result in terms of the biological question, including the direction and size of the effect. Level 3: chooses an appropriate test and interprets it biologically, but leaves assumptions unstated or does not address effect size. Level 2: reports a test result correctly but interprets it only as significant or not significant. Level 1: reports numbers with no test, or applies a test that does not match the design. Then check the repair against the specificity trade-off. Had the department instead written uses a Welch two-sample t-test, raters would agree almost perfectly and the criterion would now measure whether students followed an instruction, not whether they can choose and defend an analysis.

Key takeaway

Performance assessment buys you access to constructs that selected-response items cannot reach, and pays for it in judgment; a rubric is worth building only if its descriptors describe quality well enough for a student to act on and for two raters to converge, and even a perfect rubric cannot make one task speak for a whole domain.

Quick check

3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.

Question 1 of 3foundational

What two components must be present for a scoring tool to count as a rubric rather than a checklist or a rating scale?

Choose an answer, then check it.
Question 2 of 3intermediate

An instructor wants students to know which specific aspect of their oral presentation to improve before the next one. Which rubric form best fits that purpose, and why?

Choose an answer, then check it.
Question 3 of 3intermediate

A history rubric's Use of Evidence criterion has four levels reading many errors, some errors, few errors, and no errors. What is the primary defect, and what repair addresses it?

Choose an answer, then check it.
Practice all 5

Keep learning

Ready to build on this? Continue to the next lesson.

Practice this lesson
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related

You’ll learn to

  • Define performance assessment and explain which constructs require it rather than selected-response items.
  • Distinguish a rubric from a checklist and from a rating scale, and identify the criteria, levels, and descriptors in a given rubric.
  • Evaluate performance level descriptors and rewrite ones that count features, rate effort, or restate assignment directions.
  • Explain how rater training, calibration, and moderation control rater severity, centrality, and drift.
  • Analyze task sampling variability as the dominant threat to generalizing from a single performance task.
  • Apply rubrics with students before performance, including in self-assessment and peer assessment, and state honestly what the evidence supports.

Common mistakes

  • Writing performance levels that differ only in quantity words, such as many errors, some errors, and few errors.

    Describe what the work looks like at each level. A descriptor should let a student name the specific change that would move their work up one level, which a count of unspecified errors never does.

  • Building criteria out of the assignment directions, such as has three sources or is five paragraphs long.

    Write criteria about the learning the task is meant to reveal, for example uses a variety of relevant, credible sources. Directions belong in the instructions; a student can satisfy them perfectly while understanding nothing.

  • Calling a checklist or a 1-to-5 rating scale a rubric.

    A rubric needs both criteria and descriptions of quality at each level. Checklists ask has or has not; rating scales attach numbers or labels without saying what they look like. Both are legitimate tools, but neither shows a student what better work is.

  • Treating a score on one performance task as a stable estimate of the underlying ability.

    Task sampling variability is the largest source of error in performance assessment. Generalizability studies of elementary mathematics and science tasks found that reaching acceptable dependability took roughly eight to twenty-three tasks. Collect several performances before making a confident claim.

  • Sharing the rubric only when returning grades.

    Give the rubric before the work and teach students to use it, in self-assessment and peer assessment. Reviews of formative rubric use found no reliable performance effect in studies where rubrics were merely distributed without instruction in using them.

Easily confused

Rubric vs. Checklist

Both list criteria, but a rubric describes what each criterion looks like across levels of quality, while a checklist records only presence or absence. Checklists are efficient for procedural steps that are genuinely binary; they cannot communicate what better work would be.

Analytic rubric vs. Holistic rubric

Analytic rubrics require one judgment per criterion and yield a diagnostic profile, so they support feedback. Holistic rubrics require a single overall judgment, so they are faster and better suited to grading at volume where nobody will act on the breakdown.

Rater error vs. Task sampling error

Rater error comes from the judge (severity, centrality, halo, drift) and is largely controllable through training, moderation, and monitoring. Task sampling error comes from which exercise the student happened to get, and is reduced only by using more tasks.

General rubric vs. Task-specific rubric

A general rubric applies to a family of similar tasks and can be shared with students in advance, supporting learning. A task-specific rubric names the content a particular response should contain, which raises rater agreement but cannot be shared without giving away the answer.

Key vocabulary

Performance assessment
An assessment in which the student produces the behavior or product of interest, such as an essay, an investigation, or a recital, and a judge evaluates what was produced.
Authentic assessment
A design goal for tasks: the challenge resembles how the skill is actually used, is ill-structured, allows planning and revision, and is known to the student in advance.
criterion (in a rubric)
One named quality that a judge looks for in the work, such as use of evidence or control of technique, chosen because it reflects the intended learning.
Performance level descriptor
The sentence in a rubric cell that says what one named quality looks like in work at a particular point along the continuum from weak to strong.
Analytic rubric
A scoring instrument that requires a separate judgment for each named quality, producing a profile of strengths and weaknesses rather than one overall number.
Holistic rubric
A scoring instrument that asks for one judgment taking all named qualities into account at once, trading diagnostic detail for speed.
Single-point rubric
A format that states only the proficient standard for each named quality, leaving open space on either side to record evidence of shortfall and of work that exceeds the target.
Moderation
A procedure in which several judges score the same sample of work independently, then compare and discuss discrepancies until their interpretations of the scoring guide converge.
Rater drift
The gradual movement of a trained judge away from the agreed interpretation of a scoring guide over a long scoring session, corrected by calibration sets and periodic review.
Task sampling variability
The extent to which a student's score changes depending on which particular exercise they were given, treated as measurement error when the score is meant to represent a whole domain.

Sources & references

  1. The Case for Authentic Assessment. ERIC Digest (ED328611) — Grant Wiggins; ERIC Clearinghouse on Tests, Measurement and Evaluation / American Institutes for Research, prepared under U.S. Department of Education OERI contract R88062003
  2. Appropriate Criteria: Key to Effective Rubrics — Susan M. Brookhart, Duquesne University; Frontiers in Education 3:22
  3. Sampling Variability of Performance Assessments (CSE Technical Report 361) — Richard J. Shavelson, Xiaohong Gao and Gail P. Baxter, University of California, Santa Barbara / National Center for Research on Evaluation, Standards, and Student Testing (CRESST), UCLA
  4. The Use of Scoring Rubrics: Reliability, Validity and Educational Consequences (EJ796733) — Anders Jonsson and Gunilla Svingby; Educational Research Review 2(2), 130-144; ERIC catalog record
  5. The Use of Scoring Rubrics for Formative Assessment Purposes Revisited: A Review (EJ999454) — Ernesto Panadero and Anders Jonsson; Educational Research Review 9, 129-144; ERIC catalog record
  6. Examining the Impacts of Rater Effects in Performance Assessments — Stefanie A. Wind, University of Alabama; Applied Psychological Measurement 43(2), 159-171; open-access copy via PubMed Central (PMC6376535)
  7. How Good Are Our Raters? Rater Errors in Clinical Skills Assessment (ED494160) — Cherdsak Iramaneerat and Rachel Yudkowsky, University of Illinois at Chicago; conference paper archived in ERIC, Institute of Education Sciences, U.S. Department of Education
  8. NAEP Technical Documentation: Backreading — National Center for Education Statistics, Institute of Education Sciences, U.S. Department of Education
  9. NAEP Technical Documentation: Scoring Monitoring — National Center for Education Statistics, Institute of Education Sciences, U.S. Department of Education
  10. Single Point Rubric: A Tool for Responsible Student Self-Assessment (repository record) — Jarene Fluckiger, University of Nebraska at Omaha; The Delta Kappa Gamma Bulletin 76(4), 18-25; DigitalCommons@UNO record
  11. Rubric Best Practices, Examples, and Templates — DELTA (Digital Education and Learning Technology Applications), NC State University

EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.

Researched 2026-08-18

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.