Instructional Design Foundations
Lesson 9 of 12

Lesson 9 Unit D: Judging the work

Assessment Fundamentals

Assessment is where a design proves or fails to prove that it worked, and it is the part of the field most often done badly by people who are otherwise good. This lesson gives you the vocabulary (formative and summative, validity and reliability), the craft of writing an item that measures what it claims, the three rubric types and when each fits, and Kirkpatrick's four levels for judging a whole program. Unit D is about judging the work; this lesson is the one that makes your judgments defensible.

By the end you should be able to distinguish formative from summative assessment and state the purpose of each (9.1); define validity and reliability and identify which is threatened in a described case (9.2); write an aligned selected-response item that avoids the common flaws (9.3); choose an analytic, holistic or single-point rubric for a described purpose and justify it (9.4); and name Kirkpatrick's four levels and the evidence each requires (9.5). It carries Unit D's first constructed response.

Read alongside this lesson: Mastering Performance-Based Objectives and Assessments (10 minutes, reread from Lesson 5 with assessment in mind) and Evaluating Course Effectiveness (9 minutes). Reference only: Understanding the Quality Matters Rubric; Quality Matters is a licensed standard, so this course teaches the idea of a course-quality standard and names QM and SUNY's open OSCQR as two examples, rather than teaching you to apply QM's rubric itself.

9.1 Formative and summative

Formative assessment happens during learning and exists to change what happens next: to tell the learner what to work on and the instructor what to reteach. Its value is in the feedback, not the score; a formative check with no feedback is a waste of a question. The checkpoints in this course, the "check yourself" prompts, a mentor's comment on a draft: all formative.

Summative assessment happens at the end of a unit of learning and exists to certify: did the learner meet the objective, yes or no, to what degree. Its value is in the judgment, so it has to be fair and defensible in a way a formative check does not. The unit assessments in this course are summative; so is a certification exam, a capstone review, a practicum sign-off.

The distinction is about purpose, not format. The same ten multiple-choice items are formative if the learner gets feedback and tries again, and summative if the score goes on a record. Designs go wrong when they confuse the two: a summative test given with no prior formative practice is an ambush, and a formative check treated as a gate punishes the learning it was supposed to support. Lesson 4's evidence on retrieval practice is the argument for frequent formative assessment; this lesson's evidence on validity is the argument for careful summative assessment.

9.2 Validity and reliability

Validity asks whether the assessment measures what it claims to measure. An assessment of "can conduct a difficult conversation" that consists of multiple-choice items about conversation models measures knowledge of the models, not the ability to conduct one. It may be a fine assessment; it is not valid for the claim. The everyday form of validity for a designer is alignment: does the item ask for the same behavior, at the same kind and depth of thinking, as the objective promised? Lesson 5's three-column table is a validity check.

Reliability asks whether the assessment measures consistently: would the same learner get the same result on a different day, on a parallel form, from a different scorer? A multiple-choice test with a fixed key is reliable by construction, as long as the items are unambiguous. A performance judged by a human is only as reliable as the rubric and the scorer's calibration, which is why the Work Rubric asks two assessors to score the same round and compare, and why this course's constructed responses get the same treatment in the first cohort.

The two are independent and both are needed. A bathroom scale that is always two kilograms heavy is reliable and not valid. A scale that gives a different reading each time you step on it is neither. An outcome claim that rests on an assessment that is not both valid and reliable is a claim the evidence does not support, whatever the pass rate says, and the program's honesty rule (no outcome claim beyond the evidence) is enforced here.

Check yourself: a company reports that 96 percent of staff "passed" data-privacy training. The assessment was five true-or-false questions shown immediately after the content, with the answers visible in the previous screen. Validity or reliability, and why?

Validity, first: the assessment measures whether a person can find an answer on the previous screen, not whether they can handle personal data correctly at work. Reliability is also weak (five true-or-false items give a high score to guessing), but the deeper problem is that a 96 percent pass rate on this instrument says nothing about the claim. Report the number as what it is: completion, not competence.

9.3 Writing an item that measures

You will write hundreds of selected-response items and review thousands. Most first drafts share a small set of flaws, and knowing them turns item writing from a chore into a craft.

Start from the objective. An item exists to measure one objective at its stated Bloom's process and DOK. Write the objective's number beside the item before you write the item. If you cannot, the item is measuring something else.

The stem poses one clear problem, in positive language, with enough context to be answerable and no more. A stem that can be answered without reading the options is testing recall of a phrase, not the objective. A stem with a negative ("which of the following is NOT") is answered wrong by careful readers who miss the NOT; if you must, capitalize and bold it, and prefer not to.

The key is the one option a subject expert would agree is best. "Best" matters: at DOK 3 there may be several defensible answers and one strongest, and the stem should say "best" or "most" so the learner knows that is the game.

The distractors are the whole art. Each one should be the answer a plausible misconception would produce, so that a learner who holds that misconception picks it, and you learn which misconception they hold. Distractors that no one would choose are dead weight. Common flaws that give the key away: the key is longer or more detailed than the distractors; the key repeats a word from the stem; the distractors do not agree grammatically with the stem; "all of the above" or "none of the above" (the first is chosen by anyone who recognizes two true options, the second measures nothing about the objective).

The feedback, for a formative item, names the misconception the chosen distractor represents and points to where in the material it is corrected. "Incorrect, see section 3" is not feedback. "That is the definition of reliability; the case describes an instrument measuring the wrong thing, which is a validity problem. Reread 9.2." is.

The DOK check. Cover the options. Can the stem be answered from a definition? Then it is DOK 1 whatever it looks like. A DOK 2 item gives a short case and asks the learner to classify it. A DOK 3 item gives a case with competing considerations and asks for the best course of action with a defensible reason. A checkpoint can live at DOK 1 and 2. A unit assessment that is all DOK 1 has been written to be easy to write.

9.4 Three kinds of rubric

A rubric is the instrument for scoring a performance or a product that a key cannot score. Three forms, each for a different purpose.

Analytic rubrics break the performance into criteria and describe several levels of each. They give detailed, criterion-by-criterion feedback and are the right tool when the learner needs to know which part to fix, and when several scorers need to agree. The Work Rubric is analytic: four criteria, four levels, descriptors in every cell. Their cost is length, and the risk is that a long rubric gets scored by impression anyway.

Holistic rubrics describe the whole performance at each level in one paragraph. They are fast, and right for a summative judgment where an overall verdict is what matters and the parts are hard to separate. Their cost is feedback: a holistic score tells the learner where they landed, not why.

Single-point rubrics describe only the standard, in one column, and leave space on either side for the scorer to note what fell short and what exceeded. They are the right tool when learners will self-assess before submitting, when the scorer is one person reading many short pieces, and when the goal is a clear "meets or not yet" with a note. The constructed responses in this course use single-point rubrics for exactly those reasons, and the program's own rubric guidance says a criterion a learner cannot self-assess against is a vague criterion.

Check yourself: a mentor will read forty two-paragraph analyses and needs to tell each learner what to fix. Which rubric type?

Single-point. One reader, many short pieces, feedback the learner can act on, and a standard the learner can hold up against their own draft before submitting. An analytic rubric would be more precise and slower than the volume allows; a holistic one would give a level with no note.

9.5 Kirkpatrick's four levels

Donald Kirkpatrick's framework evaluates a training program rather than a learner, and it is the one every corporate stakeholder knows, so you must be fluent in it.

LevelQuestionEvidenceHonest note
1. ReactionDid they find it valuable and engaging?Post-course surveys, ratingsThe cheapest to collect and the least informative; satisfaction and learning are weakly related
2. LearningDid they acquire the knowledge, skill, attitude?Assessments, before and after where possibleOnly as good as the assessment's validity (9.2)
3. BehaviorDo they do it differently on the job?Observation, manager reports, performance data, weeks or months laterRarely measured because it is hard; it is the level that answers the sponsor's real question
4. ResultsDid the organization's outcome change?The business metric the needs analysis namedAttribution is genuinely difficult; be honest about what else changed

Two uses. As a planning tool, run it backwards: decide at Plan which result (Level 4) and which behavior (Level 3) the project exists for, and how you will see them, before you decide what to teach. That is question three and question five of minimum viable analysis in Kirkpatrick's vocabulary. As a review tool, ask what level a program's evidence actually reaches; most reach Level 1 and claim Level 4, and naming that gap politely is a mark of a designer worth hiring.

Course-quality standards

Several organizations publish standards for judging the quality of a whole course, usually online: alignment of objectives, activities and assessments; learner support; accessibility; navigation; clarity of expectations. Quality Matters is the most widely used in higher education and is a licensed rubric; SUNY's OSCQR is an open alternative. You should know that such standards exist, what they cover, and that alignment is at the heart of every one of them, because a client may ask you to design to one. This course does not teach you to apply QM's rubric, both because it is licensed and because the underlying discipline is the alignment table you already know.

Constructed response CR-D1: write two items

Unit D's first written response, read by your mentor. As a document, uploaded here.

The objective your items must measure: "Given a short description of a training program and the data it collected, the learner will classify each data point to its Kirkpatrick level." (Apply, DOK 2.)

The task. Write two multiple-choice items that measure this objective. For each: a stem, four options, the key marked, and a one-line rationale for every distractor naming the misconception it catches.

Meets when both items test the given objective and not an adjacent one (an item asking for a definition of Level 3 tests recall, not classification); each has one keyed answer a subject expert would agree is best; every distractor has a rationale naming a misconception; and none of the common flaws appear: grammatical cluing, implausible distractors, all-of-the-above, unmarked negatives, a stem answerable without the options.

Before the checkpoint

Seven questions. Formative versus summative, validity versus reliability with a case, the item-writing flaws, the three rubric types, the four levels: all should be within reach.

Sources named in this lesson: Donald Kirkpatrick (the four levels); Quality Matters and SUNY OSCQR (course-quality standards); the linked 24/7 Teach articles. Verify years against the primary sources before citing.