Teachers are told to review AI output before using it. But what does review actually mean?
Most teachers check three things. Is the material accurate? Is it aligned to the standard? Is the tone appropriate? If all three pass, the material goes into the lesson.
That is a read-through, not an audit. A read-through catches errors. It does not catch architecture. A lesson material can pass all three of those checks and still fail the student cognitively. The output looks professional. The content is correct. The cognitive design is wrong.
In The Offloading Trap, I argued that teachers need to audit AI-generated materials, not just read them (Rhoads, 2026). That post established the four-check audit: cognitive science alignment, curriculum alignment, pacing alignment, and slop detection. This post goes deeper into the first check. It names two specific cognitive-science tests that are invisible unless you know to look for them, the two highest-leverage checks a teacher can run in ten minutes on any AI-generated material.
Test 1: The Working-Memory Check
Working memory has finite capacity. This is not a theory. It is one of the most replicated findings in cognitive science (Sweller, 1988). Every piece of processing a student does that is not directly serving the learning goal is extraneous load. It eats the cognitive capacity that should go to the actual learning.
AI-generated materials fail this check constantly. The output looks engaging. It is cognitively inefficient. A teacher prompts AI for an explainer on the water cycle. It comes back with a whimsical multi-sentence narrative about a raindrop’s “journey,” complete with a subplot, before it ever names evaporation, condensation, or precipitation. A student spends working memory following the story instead of building the actual process model. The explainer looks engaging. The student is processing the narrative, not the science.
This is not just a text problem. AI-generated images, infographics, and slide decks fail the same check. A cluttered diagram with decorative icons, extra arrows, and unrelated illustration elements is extraneous load in image form. A teacher pulling AI-generated visuals for a lesson faces the same question: does any element of this require processing that is not directly serving the learning goal? If the answer is yes, strip that element or move it. Every decorative icon, every narrative aside, every unrelated visual is stealing cognitive capacity from the actual learning.
Stadler, Bannert, and Sailer (2024) ran an experiment comparing ChatGPT-aided research to standard web search. ChatGPT users showed significantly lower cognitive load. The work felt easier. But their arguments were lower quality with shallower reasoning. The same load and reasoning-depth tradeoff shows up when students use AI directly. It is worth assuming it applies to teacher-facing AI output too, though no one has tested that specific claim yet. The felt ease is not the learning. The cognitive ease is the problem.
Test 2: The Worked-Example Check
A worked example is not a correct answer with steps. A worked example makes expert reasoning visible. It shows why each step, what would go wrong if you did it differently, and what the common misconception looks like. Non-examples are as important as examples. Showing the wrong path teaches students to recognize and avoid it. The worked-example effect is one of the most heavily replicated findings in instructional design research (Atkinson, Derry, Renkl, & Wortham, 2000).
AI does not build worked examples. It builds answer keys wearing a worked-example costume.
A teacher asks AI for a worked example of building a thesis statement from a prompt. It hands back a finished thesis: “While both novels explore identity, Novel A uses setting to fragment the protagonist’s self-concept while Novel B uses dialogue to reconstruct it.” Correct and clean. But it never shows the reasoning, how it picked the two comparison points, why a broader framing did not work. And it gives no non-example of an unsupportable, too-broad thesis to mark the boundary. A student sees a correct product without understanding the reasoning that produced it. The example demonstrates a result. It does not build a schema. If the worked example shows steps without reasoning, or if it lacks a non-example showing what goes wrong on the wrong path, it is an answer key. Add the reasoning yourself, or write the non-example by hand.
No study has yet tested whether AI-generated worked examples surface expert reasoning the way well-designed human ones do. A 2025 paper on instructional designers and ChatGPT notes that current AI-assisted content generation tends to function as an “isolated content generator” rather than something aligned to instructional frameworks. That is suggestive, not causal. I am being honest about the gap because the gap is the point: nobody has verified this, so someone has to check it by hand. This is the same move I made in The Measurement Problem (Rhoads, 2026). The research has not caught up. The teacher has to check by hand.
The Ten-Minute Audit
These two checks are not the whole science of learning. They are the two highest-leverage checks a teacher can run in ten minutes on any AI-generated material. The audit has five checks across three passes:
- Is the material accurate? (read-through check)
- Is it aligned to the learning standard? (read-through check)
- Does any element require processing that is not directly serving the learning goal? (working-memory check: if yes, strip it)
- Does the worked example show the reasoning behind each step, or just the steps themselves? (worked-example check)
- Is there a non-example showing what goes wrong if you take the wrong path? (worked-example check: if no, add one or write it yourself)
The read-through catches accuracy. The two tests catch architecture. No platform needed. No rubric needed. A teacher can do this at the point of use, standing at their desk, in the time it takes to drink coffee.
This is the skill an AI-readiness engagement teaches. Not which tool to use. How to check what the tool produces.
How This Connects
In Cognitive Offloading vs. Cognitive Outsourcing, I drew the category distinction: cognitive offloading means the tool handles mechanical work so the student can focus on cognitive work (Rhoads, 2026). Cognitive outsourcing means the tool does the thinking. These two tests are the diagnostic layer underneath that distinction. The working-memory check tells you if the material is adding non-essential processing. The worked-example check tells you if it is showing reasoning or just answers. One tells you what category your AI-generated material falls into. The other tells you why.
What This Means
The evidence base for cognitive load theory is decades deep (Sweller, 1988; Sweller, van Merrienboer, & Paas, 1998). The evidence base for AI’s impact on K-12 learning is near zero. The Stanford SCALE review screened 818 papers and found zero high-quality causal studies in U.S. schools (Stanford SCALE, 2026). The teacher’s job right now is to audit AI output against settled science, because the AI-specific research will not arrive in time for this school year. The science is already there. The AI is already there. The audit is the bridge.
If your district is building its AI instructional strategy for the coming year, I work with leadership teams and instructional coaches to design and implement AI auditing frameworks that connect cognitive science to classroom practice. This includes the two checks in this post, the full four-check audit from The Offloading Trap, and the district-level measurement framework from The Measurement Problem. Schedule a consultation to build your district’s AI auditing framework.
Science of learning, pedagogy, technology. In that order. Always.
Subscribe to Navigating Education
Weekly insights on AI, cognitive science, and evidence-based instruction, delivered every week. No spam.