Education · Assessment
Formative Versus Summative Assessment
On this page 9 sections
In 30 seconds
The same ten-question quiz can be formative or summative. What decides is not the instrument but the use made of the evidence: an assessment is formative when something actually changes because of what it showed, and summative when it records where a learner stood at an end point. Michael Scriven coined the pair in 1967 for judging curricula; Benjamin Bloom moved it onto student learning two years later. Every other term in this unit hangs off that distinction.
Why this matters
Assessment is where students first meet the gap between what a course says it values and what it actually rewards. Teacher preparation programs, licensure exams and school improvement plans all use this vocabulary, and they use it loosely: districts buy products labeled formative that are nothing of the kind. Getting the distinction right protects you from that. It also transfers. Anyone who trains, coaches, reviews code or supervises a clinical rotation faces the same question about whether an evaluation exists to improve the work or to certify it, and the same temptation to attach a score to advice that would have worked better without one.
The college version
Purpose, not instrument
Michael Scriven introduced the pair of terms in 1967, and not about students at all. He was writing about how to evaluate curricula, and he separated two roles evaluation might play: one in the ongoing improvement of a curriculum while it was still being developed, and one that let administrators decide whether the finished curriculum represented a sufficient advance on the available alternatives to justify adopting it. He proposed calling those roles formative and summative.
Two years later Benjamin Bloom moved the distinction onto student learning. He acknowledged the traditional role of tests in judging and classifying students, and set against it a use of formative evaluation that provides feedback and correctives at each stage of the teaching-learning process - brief tests used by teachers and students as aids to learning. Bloom then added the sentence everything downstream depends on: such tests may be graded and used in the judging function, but formative evaluation is much more effective when it is separated from grading and used primarily as an aid to teaching.
Read those origins carefully and the most common student error dissolves. Neither Scriven nor Bloom made formative a property of an instrument. A test is not formative; a use is. Dylan Wiliam states the criterion sharply: assessments are formative if and only if something is contingent on their outcome and the information is actually used to alter what would have happened without it. An exit ticket nobody reads is summative in effect - a record and nothing more. A state accountability test whose results genuinely reshape next summer's teacher workshops has been used formatively, however high its stakes were for the students who sat it.
So the distinction is not about timing, format, formality, or low stakes. Wiliam describes a continuum of cycle lengths instead: short-cycle within a single lesson, five seconds to an hour; medium-cycle between lessons, a day to two weeks; long-cycle between instructional units, four weeks to a year or more. What makes any of them formative is not the length of the loop, where it happens, who runs it, or even who responds. It is that evidence is evoked, interpreted in terms of learning needs, and used to adjust.
For learning, as learning, of learning
A second vocabulary sits alongside the first and is often mistaken for a synonym. Writing for Manitoba Education and the Western and Northern Canadian Protocol, Lorna Earl and Steven Katz set out three purposes of classroom assessment.
Assessment for learning Evidence gathered so that a teacher can modify and differentiate instruction and give students feedback that advances their work. Full entry → gives teachers information they use to modify and differentiate teaching. It asks not only what students know but how, when and whether they apply it, and it feeds both instructional adjustment and feedback to students.
Assessment as learning The use of assessment to build metacognition, with students monitoring their own understanding against criteria and adjusting their own strategies. Full entry → puts the student in the driver's seat. Its object is metacognition: students monitoring their own understanding, judging their work against criteria, and adjusting their strategies. Earl and Katz note that systematic assessment as learning has historically been rare, and it is the category most schools underuse. The 2018 revision of the Council of Chief State School Officers' definition makes the same move, extending the goal of the process to supporting students to become self-directed learners.
Assessment of learning The summative purpose: confirming what students know and can do, and producing statements of proficiency that other people will act on. Full entry → is summative: it confirms what students know and can do and produces statements of proficiency that other people - the next teacher, an admissions office, a licensing board - will rely on.
Two cautions. Some writers use only two categories and fold assessment as learning into assessment for learning, so the three-way split is a convention, not a law. More importantly, Earl and Katz argue that it is very difficult and sometimes impossible to serve three assessment purposes at once. An assessment designed to let students expose their confusion without penalty is not the same artifact as one designed to survive scrutiny as a proficiency claim, and trying to do both usually means doing the first one badly.
The neighboring labels
The FAST SCASS collaborative of state assessment directors published a paper for the express purpose of untangling the labels, and each one names something real.
Summative assessment An assessment whose results are used to record or certify the level a student, class or program has reached at an end point, typically after instruction has concluded. Full entry → provides information about the level of student, school or program success at an end point in time, after instruction concludes. Its results serve evaluative judgment, inferences about mastery, grades, and accountability requirements.
Interim assessment A test given periodically through the school year, also called benchmark or quarterly, used to predict later test performance, evaluate programs, or report individual student data. Full entry → - also called benchmark, common, or quarterly assessment - covers tests administered periodically through the year to serve some mix of three functions: predicting performance on a later high-stakes test, evaluating programs, and supplying teachers with individual student data. Many such products were marketed as being formative assessments in themselves. FAST SCASS quotes Margaret Heritage's diagnosis that the core problem is the false but widespread assumption that Formative assessment The process of eliciting evidence of learning and using it to change what a teacher or student does next; the label attaches to the use made of the evidence, not to the instrument. Full entry → is a particular kind of measurement instrument rather than a process. Their own verdict is blunter: interim tests give periodic snapshots, but because they are not proximate to learning as it develops they do not inform ongoing teaching, and at the time of writing there was no meaningful evidence that they enhance student learning.
Diagnostic assessment A detailed evidence-gathering procedure that identifies which specific subskills or pieces of enabling knowledge a student is missing, normally reserved for students who are not progressing. Full entry → is a procedure detailed enough to indicate which specific subskills or pieces of enabling knowledge a student does and does not possess. It answers what is not being learned, rather than whether learning is happening. Because it is time-intensive it is normally reserved for students who are not progressing, and it feeds the formative process when results arrive fast enough to act on.
Three further labels round out the set. Curriculum-embedded assessment is built into the materials or activities, so students are never told to stop and take a test. Universal screening is a brief measure given to everyone two or three times a year to flag students who may be at risk. Progress monitoring is a short measure repeated weekly or biweekly to track growth, usually with alternate forms.
Notice that every one of these is defined by an instrument, a population, or a schedule - not by a purpose. That is why each can be used formatively or summatively, and why the purpose test still has to be applied on top of the label. FAST SCASS makes the point about curriculum-embedded work explicitly: it is more likely to be formative when it is not graded, and once it enters the grading process the only thing distinguishing it from any other summative assessment is how smoothly it was blended into the lesson.
Feedback is not automatically good
The engine of formative assessment is supposed to be feedback, so it matters that feedback is not a reliably positive intervention. Avraham Kluger and Angelo DeNisi's 1996 meta-analysis pooled 607 effect sizes from 23,663 observations. The average effect was positive and moderate, d = .41. But over one third of the feedback interventions decreased performance - a substantial minority of cases in which telling people how they were doing made them worse.
Their explanation, Feedback Intervention Theory, is that feedback works by directing attention, and that its effectiveness falls as attention moves up a hierarchy from the task toward the self. Feedback that says this paragraph has no evidence in it points at the work. Feedback that says you are a careless writer, or simply 71 percent, points at the person. The second kind recruits ego defense, and ego defense does not revise paragraphs.
This is the mechanism behind the oldest practical rule in the field. Bloom said in 1969 that formative evaluation is much more effective when separated from grading. Ruth Butler tested it in 1988, giving fifth- and sixth-grade classes numerical grades, individual comments, or both, and framing the contrast as task-involving versus ego-involving evaluation. The reading that has stuck in teaching-center guidance is that adding a score to a comment removes the benefit of the comment. Black and Wiliam reached the same conclusion from the wider literature: pupils given only marks or grades do not benefit from the feedback, which improves learning when it gives specific guidance on strengths and weaknesses, preferably without any overall marks.
The operational consequence is simple. If you want a comment to do work, protect it from the grade - withhold the score until the student has responded, require a revision before releasing it, or do not score the piece at all. Writing both on the same page and hoping the student reads the left column is, on this evidence, wishful.
What good formative practice actually contains
The 2018 FAST SCASS definition calls formative assessment a planned, ongoing process used by all students and teachers during learning and teaching to elicit and use evidence of student learning, improve understanding of intended disciplinary learning outcomes, and support students in becoming self-directed learners. It names five practices that must be integrated: clarifying learning goals and Success criteria The stated features of acceptable work, shared with students in advance so that both teacher and learner can judge a performance against the same standard. Full entry → within a broader progression of learning; eliciting and analyzing evidence of student thinking; engaging in self-assessment and peer feedback; providing actionable feedback; and using that evidence and feedback to move learning forward.
Strip that to the three parts that fail most often.
First, success criteria stated in advance. Students cannot judge their own work against a standard they have not been shown, and a teacher who has not written the criterion down will grade by impression. It is also what separates feedback from opinion.
Second, evidence from every student. A question answered by the three students who raise their hands tells you about three students. Wiliam's illustration is a multiple-choice question about thesis statements where every student holds up a lettered card, so the teacher reads the whole room at once and knows who chose what. He calls the resulting instant a Moment of contingency The instant in a lesson at which evidence from students has been read and instruction can genuinely change direction as a result. Full entry →: the point at which instruction can actually change direction. Contrast it with the class discussion that feels productive and yields no readable evidence about anyone.
Third, a decision that follows. This is what the definition is really about, and the part routinely skipped. Evidence that produces no change in what happens next was not formative, however carefully it was gathered.
The word planned in the 2018 definition is doing deliberate work here. FAST SCASS argues that a teacher who appears to be adjusting on the fly is often enacting one of several plans prepared beforehand - anticipating the likely misconceptions, writing the question that will surface them, and deciding in advance what each pattern of answers will trigger. Improvisation of that quality is rehearsed.
How strong is the evidence?
You will meet a number attached to this topic, and you should handle it carefully. In their 1998 Phi Delta Kappan article Inside the Black Box, Paul Black and Dylan Wiliam wrote that typical effect sizes of the formative assessment experiments were between 0.4 and 0.7, larger than most found for educational interventions, and that improved formative assessment helps low achievers most. That sentence became the most-quoted claim in classroom assessment, repeated in textbooks and district slide decks as though it were a measured constant.
It is not one. Neal Kingston and Brooke Nash reviewed more than 300 K-12 studies of formative assessment and found only 13 reporting enough information to compute an effect size. From the 42 independent effect sizes those studies yielded, a random-effects model gave a weighted mean of 0.20, with estimates of 0.32 in English language arts, 0.17 in mathematics and 0.09 in science. Their conclusion was that the commonly claimed figure is not supported by the existing research base, and that the base itself is too thin.
Randy Bennett's critical review attacks the object rather than the estimate. He concluded that formative assessment does not yet name a well-defined set of artifacts or practices, that the breadth of the definitions in use permits implementations with wildly different outcomes, and that the magnitude of the common quantitative claims is suspect because it traces to sources that are untraceable, flawed, dated or unpublished. If the treatment is not well defined, an average effect size for it means little.
The original authors have qualified their own case. Writing in 2006, Wiliam said his own reading of the research suggested medium-cycle formative assessment had shown only modest impact, and that the studies which did show impact tended to be those that changed teachers' day-to-day and minute-to-minute practice. Federal guidance is similarly restrained: the 2009 IES practice guide on using student achievement data assigns a Low level of evidence to all five of its recommendations. Low, the panel is careful to say, means it did not identify a body of research demonstrating effects on achievement - not that the practice is unimportant.
So hold this position. The mechanism is well motivated: eliciting evidence, interpreting it, and acting on it is how teaching becomes responsive rather than scheduled. The size of the effect is unsettled, depends heavily on what was implemented, and should not be quoted as a fixed number.

Eli explains
The same idea, in plain words
Explain it like I’m 10
Imagine two ways of weighing yourself. If you step on the scale, look at the number, and change what you eat this week, the weighing changed something. If you step on the scale on the last day of a study so a researcher can write the number down, the weighing recorded something. It is the same scale and the same body. What differs is what happens next. School assessments work exactly this way. A quiz is formative when the teacher or the student does something different because of the answers. The same quiz is summative when its only job is to go in the grade book. So the useful question is never whether something is a formative assessment. The useful question is who is going to change what they do because of this, and how soon.
Picture it like this
Formative assessment is tasting the soup while you cook. Summative assessment is the diner tasting the soup after it is served. Same spoon, same soup, different job. The cook can still add salt; the diner can only deliver a verdict, and the verdict arrives too late to rescue that bowl.
Where the picture stops working
The analogy hides three things. The cook tastes for herself, but the strongest versions of formative assessment put the spoon in the student's hand and teach them to judge their own work. Soup is not discouraged by criticism and students are; a third of feedback interventions in one large review made performance worse. And the diner's verdict can be formative for the next batch, which is the whole point: it is the use, not the moment, that decides the label.
Worked example
Two teachers give the same ten-item quiz on Tuesday. Ms Okafor collects the papers, scores them out of ten, records the scores, and hands them back Wednesday with the marks circled. Her class moves on to the next section as planned. Mr Bhatt writes no score. He tallies which item each student missed and finds that nineteen of his twenty-six students missed item 7 in the same way, subtracting before distributing. He cancels Wednesday's planned lesson, reteaches distribution for fifteen minutes, and ends the period with a two-item re-check. He returns the quizzes with a written comment and no number, because a number would have crowded the comment out. Same instrument, same students, same day. Ms Okafor ran a summative assessment: the evidence became a record. Mr Bhatt ran a formative one: the evidence changed the next lesson. Note what Mr Bhatt needed and Ms Okafor did not - a response from every student, an interpretation of those responses in terms of learning needs, and a decision that actually followed.
Key takeaway
Formative and summative name what you do with the evidence, not kinds of tests. If nothing changed because of what an assessment showed, it was not formative - whatever it was called, however low the stakes, and however carefully the data was collected.
Quick check
3 questions here, of 5 in this lesson’s practice set. Answers stay hidden until you check.
What does it mean to say that formative assessment is a process rather than an instrument?
A district buys a quarterly benchmark test marketed as a formative assessment. Score reports reach teachers three weeks after each administration, once the unit has ended, and no instructional change follows. How is this best classified?
Study tools & related lessonsYou’ll learn to · Common mistakes · Easily confused · Key vocabulary · Related
You’ll learn to
- Define formative and summative assessment in terms of the purpose and use of the evidence rather than the type of instrument.
- Explain how Scriven's 1967 distinction for curriculum evaluation became Bloom's distinction for evaluating student learning.
- Distinguish assessment for learning, assessment as learning and assessment of learning, and place diagnostic, interim and benchmark assessment among them.
- Apply the purpose test to classify a real assessment scenario, including cases where the same instrument is used both ways.
- Analyze why feedback effects vary in direction, and why attaching a grade to a comment tends to blunt the comment.
- Evaluate the evidence behind popular claims about the size of formative assessment's effect on achievement.
Common mistakes
Treating formative and summative as categories of test, so that quizzes and exit tickets are filed as formative and exams as summative.
They are categories of use. The same quiz is formative if its results change the next lesson and summative if they only get recorded. Classify the decision that follows, not the paper.
Calling any low-stakes or ungraded activity formative, even when nobody looks at the results.
Being ungraded is not sufficient. An assessment is formative only if something is contingent on its outcome and the information actually alters what would otherwise have happened.
Accepting a vendor's or district's label - treating a quarterly benchmark product as formative because it is sold that way.
Interim and benchmark tests are a separate category defined by their schedule and function. They inform ongoing teaching only when results arrive in time to change instruction and someone changes it.
Quoting an effect of 0.4 to 0.7 standard deviations as the established benefit of formative assessment.
That range comes from a 1998 summary article. A later meta-analysis found a weighted mean of about 0.20 from the few studies with usable data, and critics have shown the original figures trace to sources that cannot be checked.
Assuming more feedback is always better, so long as it is delivered kindly.
In the largest meta-analysis of feedback interventions, over a third lowered performance. Feedback aimed at the person rather than the task is the failure mode, and a grade written beside a comment tends to pull attention toward the person.
Easily confused
Formative use of evidence vs. Summative use of evidence
Formative use changes what happens next for these learners; summative use records where they stood. The same instrument, and even the same administration, can be put to either use.
Assessment for learning vs. Assessment as learning
Both are formative in purpose, but for learning positions the teacher as the one who acts on the evidence, while as learning trains the student to monitor and adjust their own work.
Diagnostic assessment vs. Interim or benchmark assessment
Diagnostic work is fine-grained and targeted at students who are stuck, identifying which subskill is missing. Interim tests are broad periodic snapshots of the whole cohort, better at prediction and reporting than at telling a teacher what to do on Thursday.
A grade vs. Descriptive feedback
A grade summarizes the person's standing; descriptive feedback describes the work and what to do to it. Delivered together, the grade tends to win the student's attention and the feedback's benefit is lost.
Key vocabulary
- Formative assessment
- The process of eliciting evidence of learning and using it to change what a teacher or student does next; the label attaches to the use made of the evidence, not to the instrument.
- Summative assessment
- An assessment whose results are used to record or certify the level a student, class or program has reached at an end point, typically after instruction has concluded.
- Assessment for learning
- Evidence gathered so that a teacher can modify and differentiate instruction and give students feedback that advances their work.
- Assessment as learning
- The use of assessment to build metacognition, with students monitoring their own understanding against criteria and adjusting their own strategies.
- Assessment of learning
- The summative purpose: confirming what students know and can do, and producing statements of proficiency that other people will act on.
- Diagnostic assessment
- A detailed evidence-gathering procedure that identifies which specific subskills or pieces of enabling knowledge a student is missing, normally reserved for students who are not progressing.
- Interim assessment
- A test given periodically through the school year, also called benchmark or quarterly, used to predict later test performance, evaluate programs, or report individual student data.
- Success criteria
- The stated features of acceptable work, shared with students in advance so that both teacher and learner can judge a performance against the same standard.
- Moment of contingency
- The instant in a lesson at which evidence from students has been read and instruction can genuinely change direction as a result.
- Descriptive feedback
- Comment that tells a learner what is strong, what is weak and what to do next, directed at the work rather than at the person.
Sources & references
- Formative Assessment: Getting the Focus Right (Educational Assessment, 11(3-4), 283-289) — Dylan Wiliam, Educational Testing Service; author manuscript deposited in UCL Discovery, University College London
- Revising the Definition of Formative Assessment — Formative Assessment for Students and Teachers (FAST) State Collaborative on Assessment and Student Standards, Council of Chief State School Officers (2018); copy hosted by the Michigan Department of Education
- Distinguishing Formative Assessment from Other Educational Assessment Labels — Formative Assessment for Students and Teachers (FAST) SCASS, Council of Chief State School Officers (2012); copy hosted by the Michigan Department of Education
- Rethinking Classroom Assessment with Purpose in Mind: Assessment for Learning, Assessment as Learning, Assessment of Learning — Lorna Earl and Steven Katz (Aporia Consulting) with the Western and Northern Canadian Protocol for Collaboration in Education assessment team; Manitoba Education, Citizenship and Youth (2006)
- The Effects of Feedback Interventions on Performance: A Historical Review, a Meta-Analysis, and a Preliminary Feedback Intervention Theory (Psychological Bulletin, 119(2), 254-284) — Avraham N. Kluger and Angelo DeNisi; record in the Hebrew University of Jerusalem research information system
- Enhancing and Undermining Intrinsic Motivation: The Effects of Task-Involving and Ego-Involving Evaluation on Interest and Performance (British Journal of Educational Psychology, 58(1), 1-14) — Ruth Butler; ERIC record EJ380489, Institute of Education Sciences, U.S. Department of Education
- Feedback to Support Students' Motivation and Learning — Distance Education and Learning Design, College of Education and Human Ecology, The Ohio State University
- Inside the Black Box: Raising Standards Through Classroom Assessment (Phi Delta Kappan, 80(2), 1998) — Paul Black and Dylan Wiliam; republished by Kappan Online, Phi Delta Kappa International
- Formative Assessment: A Meta-Analysis and a Call for Research (Educational Measurement: Issues and Practice, 30(4), 28-37) — Neal Kingston and Brooke Nash; ERIC record EJ951173, Institute of Education Sciences, U.S. Department of Education
- Formative Assessment: A Critical Review (Assessment in Education: Principles, Policy & Practice, 18(1), 5-25) — Randy Elliot Bennett; ERIC record EJ912798, Institute of Education Sciences, U.S. Department of Education
- Using Student Achievement Data to Support Instructional Decision Making (IES Practice Guide, NCEE 2009-4067, September 2009) — Institute of Education Sciences, What Works Clearinghouse, U.S. Department of Education
EliExplains lessons are original prose written from the open, credible references above. See Copyright & Licensing.
Researched 2026-08-18
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.

