Claude for Teachers launched this month. Teachers can pull in assessment data, lesson plans, and standards from all 50 states, then ask an AI to build personalized instruction overnight.
That’s powerful. It’s also the exact moment to be honest about what happens between the prompt and the classroom.
I’ve written about this distinction from the student’s perspective, cognitive offloading vs. cognitive outsourcing, where the question is whether the tool does the mechanical work or the thinking. This is the teacher’s side of the same equation. When a teacher uses AI to generate instructional materials, the same line applies. The tool can offload the mechanical work of formatting, generating, and assembling. Or it can outsource the instructional decisions that should belong to the teacher.
The audit is not all-or-nothing. Every check you run converts a little outsourcing into offloading. Every check you skip leaves the tool thinking for you.
Every AI-generated instructional material should pass four checks before it reaches students. Not one. Not two. Four. Anything less is outsourcing dressed up as efficiency.

Science of learning → pedagogy → technology. In that order. Always.
Check 1: Cognitive Science Alignment

Does the output align with how learning actually works? Five principles matter here.
Cognitive load. AI loves to add colorful borders, decorative graphics, and extra steps to a worksheet. That looks engaging. It is increasing cognitive load. The brain is processing the decoration instead of the content. Sweller’s framework calls this extraneous load, and it’s the enemy of learning (Sweller, 1988; Sweller, van Merriënboer, & Paas, 1998). Minimize noise. Channel effort toward building understanding. If the output adds clutter, it fails this check.
Formative assessment. An AI-generated quiz that auto-scores is producing data. A teacher who uses that data to see misconception patterns is producing feedback. Formative assessment works when it produces information that changes the next instructional move (Black & Wiliam, 1998; Hattie & Timperley, 2007). Does the output help you see what students are thinking, or does it just score what they wrote?
Dual coding. Verbal and visual channels together strengthen retention (Paivio, 1971; Clark & Paivio, 1991). But irrelevant visuals increase cognitive load and impair learning (Mayer, 2009). A labeled diagram alongside text is dual coding. A colorful infographic that doesn’t map to the learning goal is decoration. The output should pair visuals with meaning, not fill space.
Spaced practice. Distributed practice across days outperforms massed practice, even when total study time is held constant (Cepeda, Pashler, Vul, Wixted, & Rohrer, 2006; Dunlosky, Rawson, Marsh, Nathan, & Willingham, 2013). I’ve covered this principle in depth elsewhere. The audit question is narrower: does the AI output embed distributed review, or does it default to a 50-question cram session?
Interleaving. Mixing problem types produces better transfer than blocked practice (Rohrer & Taylor, 2007; Bjork & Bjork, 2011). Many AI tools default to blocked practice because it feels easier. The output should mix problem types, not group them in blocks.
If the output fails any of these, it is working against how learning works. No amount of personalization fixes that.
Check 2: Curriculum Alignment
Does the output match what your district actually teaches?
AI tools pull from broad training data. Standards from all 50 states. Textbooks from multiple publishers. Lesson frameworks from dozens of sources. The output will look correct. It may not be correct for your context.
Ask: Does this align with your state standards, not just the topic but the grade-level expectation? Does it match the scope and sequence your district has adopted? Does it use the vocabulary your teachers have been trained on? Does it match the texts your students are actually reading?
An AI-generated lesson on “main idea” that uses passages from a training-data corpus instead of the texts in your curriculum is not curriculum-aligned. It is a generic lesson that happens to share a topic with your curriculum. The difference matters.
Check for accessibility too. Does the output work for your students with IEPs, your English learners, your students reading two grade levels below? AI-generated materials fail here most often. The language is generic, the scaffolds are absent, and the differentiation is surface-level at best. If the output doesn’t serve the students who need the most support, it’s not ready for your classroom.
Check 3: Pacing Alignment
Does the output fit where your students are in the learning progression?
A perfectly designed lesson that arrives a week early or late is not useful. AI tools do not know what you taught yesterday. They do not know what is coming tomorrow. They do not know which skills your students have mastered and which ones need another day.
Pacing is an instructional decision, one of the most consequential ones a teacher makes. When a teacher accepts AI-generated pacing without checking it against the progression, that is outsourcing. And the first students to feel it are the ones who needed the scaffolded sequence most.
Check 4: Slop Detection
Is the output actually good, or does it just look good?
AI-generated materials can be fluent, well-formatted, and confidently hollow. Slop is output that reads like instruction but does not hold up under scrutiny. It comes in specific forms:
- Standard-naming without alignment. The lesson cites RI.4.3 but the task doesn’t require students to analyze connections between events.
- Surface recall masquerading as rigor. The assessment looks demanding but every answer is stated verbatim in the text.
- Activity that fills time without producing evidence. The lesson has students “discuss in groups” but produces no work product that reveals understanding.
- Volume without structure. Fifty practice problems when ten well-sequenced ones would do more.
Slop is not the same as error. An AI output can be factually correct and still be slop. Hollow instruction that occupies time without producing learning. Correctness alone is not a sufficient bar.
The only defense is a teacher who reads every output critically and asks: would I have taught this before AI existed? If the answer is no, the output needs revision or rejection.
The Audit Is Where Professional Judgment Lives
The audit is not extra work on top of the AI output. The audit is where the teacher’s professional judgment lives in an AI-assisted workflow. The alignment. The pacing. The fit between the material and the students in the room.
When a teacher audits an AI-generated lesson against cognitive science, curriculum, pacing, and quality, that teacher is offloading the mechanical work and owning the cognitive work. That is productive offloading.
When a teacher skips the audit because the output looks good, that is outsourcing. The tool did the thinking. The teacher delivered someone else’s decision.
Will every teacher run four checks on every output every time? No. That is an aspiration, not a Tuesday in October. The real goal is building a habit so the audit becomes instinctive, fast, and partial beats skipping entirely. A teacher who runs two checks is better off than a teacher who runs none.
The question is not whether AI belongs in instruction. It does. The question is whether teachers have the framework and the time to audit outputs before they reach students. Four checks. Build the habit. Start with one.
Science of learning → pedagogy → technology. In that order. Always.
If you are a district or school leader: bring this framework to your next instructional leadership team meeting. Ask one question. What would it look like for every teacher in your system to run AI outputs through these four checks before they reach students? Then build the PD plan to get there.
I am currently writing a book that expands on this framework. A deeper treatment of the four checks, the slop taxonomy, and what it looks like when a district builds an AI audit culture from the ground up. More details soon.
References
Bjork, R. A., & Bjork, E. L. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher, R. W. Pew, L. M. Hough, & J. R. Pomerantz (Eds.), Psychology and the real world (pp. 56–64). Worth Publishers.
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education, 5(1), 7–74.
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks. Psychological Bulletin, 132(3), 354–380.
Clark, J. M., & Paivio, A. (1991). Dual coding theory and education. Educational Psychology Review, 3(3), 149–210.
Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4–58.
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112.
Mayer, R. E. (2009). Multimedia learning (2nd ed.). Cambridge University Press.
Paivio, A. (1971). Imagery and verbal processes. Holt, Rinehart, and Winston.
Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science, 35(6), 481–498.
Sweller, J. (1988). Cognitive load during problem solving. Cognitive Science, 12(2), 257–285.
Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1), 59–89.
Sweller, J., van Merriënboer, J. J. G., & Paas, F. (1998). Cognitive architecture and instructional design. Educational Psychology Review, 10(3), 251–296.