Instructional Design Foundations
Lesson 11 of 12

Lesson 11 Unit D: Judging the work

AI in the Instructional Design Workflow

The learning model's shortest sentence about AI is the one to hold onto: AI lowered the cost of a first draft and raised the cost of trusting one. This lesson is about that second cost. It gives you the program's stance on what AI does well and poorly in instructional design work, a method for judging an AI-generated draft through the four Evaluation lenses, the rules about what never goes into a model, and the discipline of verifying before using. Then it asks you to do it, on a draft with real errors planted in it.

By the end you should be able to state what current AI does well and poorly in ID work, in the model's terms (11.1); judge an AI-generated artifact through the four lenses and name specific failures (11.2); state the data hygiene rules (11.3); and apply verify-before-using to identify the claims in a draft that must be checked (11.4). It carries Unit D's second constructed response.

Read first: The Instructional Designer's Guide to Becoming an AI Translator and Evaluator (12 minutes). Read for the stance, not the tool: Why All 24/7 Teach Instructional Designers Are Trained on How to Use ChatGPT; it names one tool and the tool landscape has moved, but the argument that fluency is a job requirement has not. The program does not organize around any tool; the bootcamp uses several and expects you to pick up the next one yourself.

11.1 What AI does well, what it does poorly, and why the job moved

Current generative models produce a competent first draft of almost any instructional artifact in seconds: an objective set, a module outline, a storyboard, a quiz, a case, a job aid, a facilitator script. The drafts are fluent, well organized, and confident. That is what they do well, and it is not nothing: the entry layer of the field, the work a new designer used to spend years on, is now nearly free.

What they do poorly is the layer above. A model does not know your learners; it knows a plausible average of learners. It does not know what your analysis found, so its draft answers a generic question rather than the one you were asked. It cannot tell whether its own draft matches how people learn, because it has no way to check its output against the evidence; it can only produce something that sounds like the evidence. It states things that are not true with the same fluency as things that are. And it will happily produce a beautiful artifact aligned to nothing, which is the Alignment criterion's exact description of the failure that speed makes likelier: makes a lot and checks little.

The Senior Mindset article in the mindset track summarizes research showing that people who trusted the AI more thought less critically about its output, and that the people least equipped to catch a flawed draft were the least likely to look for the flaw, because fluency gives a novice no reason for suspicion. That is the trap, and this course's whole Unit B was built to arm you against it: you cannot catch a violation of the spacing effect if you have never heard of it.

So the job moved. The designer's value is now the judgment applied to the draft: whether it serves what the analysis found, whether it will work, what to change. The program calls this being an AI translator and evaluator: translating a need into something a model can help with, and evaluating what comes back against standards the model does not have. The linked article is the full argument.

11.2 Judging an AI draft through the four lenses

The Evaluation criterion's four lenses were written for judging any finished work. They are especially sharp on AI drafts because each lens catches a characteristic failure of generated content.

LensAsk of the draftThe characteristic AI failure it catches
Learning science validationDoes this match how people actually learn?Content presented once and tested at the end; long undivided passages; decorative "engagement"; objectives at Remember with a quiz at Remember for a skill that needed practice. Run Lesson 4's findings and Lesson 6's frameworks against it.
Contextual relevanceDoes it fit the real application need?Generic examples that fit no one; assumptions about the audience the analysis contradicts; the right answer to a question nobody asked. Run Lesson 8's five questions against it.
Instructional integrityWill it produce the intended outcome?Objectives that cannot be measured; activities that do not practice the objective; assessments misaligned to both. Run Lesson 5's alignment table against it.
Implementation feasibilityCan it be executed given real constraints?A ninety-minute design for a fifteen-minute slot; interactions the authoring tool cannot build; a video budget nobody has; an accessibility failure that will fail procurement. Run Lesson 10 against it.

Method: read the draft once for what it is. Then read it four more times, once per lens, writing down every specific failure with the lens that caught it and the lesson that explains why it is a failure. Do not fix as you go; a list of named failures is worth more than a half-repaired draft, because it is what you would bring to a Demo. Then decide what to keep, what to change and what to discard. A draft that fails all four lenses is still a faster start than a blank page, and knowing exactly how it fails is the skill.

11.3 Data hygiene: what never goes into a model

The rules are short and they are not negotiable, because the consequences are legal and reputational and they land on your employer and your learners, not on the model.

  • Personal information about learners never goes in: names attached to performance, assessment results, health or disability information, contact details. In K-12 and higher-education work this is the Family Educational Rights and Privacy Act (FERPA) in the United States and equivalents elsewhere, and the penalties are real. In corporate work it is employee data governed by privacy law and by the employer's own policy.
  • Client confidential content never goes in without written permission: proprietary procedures, unreleased products, internal data, anything under a non-disclosure agreement. Assume a public model may retain and learn from what you paste unless your organization has a contract that says otherwise.
  • Use the tools your organization has approved, in the configuration it has approved. An enterprise account with a data-processing agreement is a different thing from a personal account with the same name on it.
  • Anonymize and abstract before you prompt. "A new charge nurse missed items on handoff" is a fine prompt. The nurse's name, unit and incident report are not.

When in doubt, ask before pasting, and write down what you were told. A designer who can say "I checked our policy and used the approved tool with anonymized data" is safe. One who cannot is a liability, however good the draft was.

11.4 Verify before using

A generated draft will contain claims: statistics, citations, quotations, tool capabilities, regulatory requirements, historical facts. Some will be right. Some will be invented with the same confidence, including citations to papers that do not exist and quotations nobody said. The discipline is: a claim in a draft is a claim to be checked, not a fact to be used.

Practically: highlight every number, every named source, every "studies show," every "the regulation requires," every quotation. For each, find the primary source or remove the claim. "Research shows spaced practice improves retention" survives if you can point at Cepeda; "a 2019 study found a 43 percent improvement" does not survive unless you have the study in front of you. This course models the rule on itself: publication years are omitted throughout and every lesson tells you to verify against the primary source, because the writer did not have every source open and would rather leave a year out than guess one.

The same rule applies to what a draft says about tools and standards. A model's account of what an authoring tool can do, or which WCAG version a contract requires, is a hypothesis. Check the documentation.

Check yourself: a generated storyboard includes "According to the Association for Talent Development, 68 percent of learners prefer microlearning." What do you do?

Treat it as unverified. Find the ATD source. If it exists and says that, cite it properly with the year and the report. If you cannot find it in ten minutes, delete the sentence; the storyboard does not need it, and a stakeholder who asks for the source and finds there is none has learned something about you.

Constructed response CR-D2: judge the draft

Unit D's second written response, read by your mentor. About 200 words, as a document, uploaded here.

The context. The hospital handoff brief from Lesson 7: new charge nurses, experienced clinicians, fifteen minutes at most, shared workstations with no headphones, handoff checklist errors causing medication delays. A designer prompted a model and received the draft below.

Module: Excellence in Shift Handoffs (45 minutes)

Objectives. By the end of this module, learners will: (1) understand the importance of effective handoff communication; (2) be familiar with the SBAR framework; (3) appreciate how handoff errors impact patient safety; (4) complete the unit handoff checklist accurately.

Screen 1. Welcome animation with upbeat music (autoplay). Narrator: "Welcome to Excellence in Shift Handoffs!" Stock photo of smiling nurses.

Screens 2 to 9. Narrated slides explaining the history of handoff communication, the SBAR framework, and the eight most common handoff errors. Narrator reads each slide's text aloud. Key risks highlighted in red.

Screen 10. "Did you know? Studies show that 80 percent of serious medical errors involve miscommunication during handoffs (Joint Commission, 2017)."

Screen 11. Knowledge check: five true-or-false questions on the SBAR framework.

Screen 12. Congratulations screen. Certificate of completion.

The task. Name at least two real misalignments or failures in this draft, specifically (which objective, which screen), and tie each to the Evaluation lens that catches it. Flag at least one claim in the draft that must be verified before use. Say in one sentence what you would keep.

Meets when at least two real failures are named specifically and each is tied to the lens that catches it; at least one claim is flagged as needing verification before use; and the response does not accept a fluent element that is wrong. Praising the draft's polish without finding the planted problems is not yet. There are more than two problems in this draft; a strong response finds several and ranks them.

Before the checkpoint

Five questions. The stance, the four lenses as applied to a draft, the data rules, and the verification discipline. Lesson 12 is the Unit D assessment.

Sources named in this lesson: the two linked 24/7 Teach articles; the learning model's four lenses; the Senior Mindset article's summary of research on AI and critical thinking (which itself flags its findings as contested). The statistic and citation inside the sample draft are part of the exercise and are not asserted by this course.