Reference Document · AI Pedagogy

AI Failure Modes
& Blind Spots

A taxonomy of how AI-assisted work goes wrong — organized by who fails, how visibly, and at what level. Developed for teaching students to use AI well, not just use AI.

Four categories · Behavioral / Cognitive / Systemic · Examples drawn from humanities courses
01
Observable · User-side
User Habits
Behavioral patterns in how people use AI. 100-level patterns are visible from the outside and easy to name. 200- and 300-level patterns are subtler — errors of interaction style and frame that can look like good use from the outside.
100-level — Errors of effort
Paste-in / Paste-out
Copy the assignment in, copy the output out. No thinking at either end. The student is a conduit, not a participant. The most egregious form of outsourcing.
Sophomore American literature survey. Weekly reading response. The assignment asks students to identify something specific from their own reading of the text.
Faculty Prompt
In 2–3 paragraphs, identify one image or motif from this week's reading in The Great Gatsby and explain what it reveals about Fitzgerald's themes. Use at least one direct quote from the text to support your reading.
Student → AI
In 2–3 paragraphs, identify one image or motif from this week's reading in The Great Gatsby and explain what it reveals about Fitzgerald's themes. Use at least one direct quote from the text to support your reading.AThe student prompt is the faculty prompt copied exactly. The assignment asked the student to identify something from their own reading — to notice something. That act of noticing is what the assignment was designed to produce. It has been skipped entirely.
AI Response
One of the most striking motifs in The Great Gatsby is the imagery of light and darkness, which Fitzgerald uses to explore the tension between illusion and reality. The green light at the end of Daisy's dock represents Gatsby's longing for a past he cannot recover. As Nick observes, Gatsby reaches toward it "with trembling fingers"B"With trembling fingers" does not appear in the novel. Fitzgerald's description of Gatsby reaching toward the light is wordless — Nick sees him stretching his arms toward the dark water. The AI has fabricated a plausible-sounding quote. A student who did the reading would catch this immediately. This student did not do the reading. — a gesture that captures both his hope and its impossibility.

This imagery extends throughout the novel. Gatsby's parties blaze with artificial light, yet the man himself remains oddly absent, watching from a distance. The brightness is a facade: the entire edifice of his world is theatrical — spectacular, hollow, and temporary.CGeneric — this observation would appear in any Gatsby essay. Nothing is drawn from a specific moment the student noticed while reading. The faculty prompt required a direct quote; the AI produced a fabricated one in paragraph one and none in paragraph two.
One-Shot Prompting
Treats the AI like a search engine. Ask once, accept the result. No iteration, no pushback, no refinement. Misses the entire value of dialogue.
Junior-level philosophy course. Students have read Mill's Utilitarianism and Kant's Groundwork and discussed both in class for two weeks.
Faculty Prompt
In one paragraph, apply both frameworks to the following situation: a doctor has five patients who will die without organ transplants and one healthy patient whose organs could save all five. Would a utilitarian and a Kantian give the same advice or different advice here? Which response do you find more convincing and why?
Student → AI
What is the difference between utilitarianism and deontological ethics?AThe faculty prompt gave the student a concrete scenario and a pointed question. The student ignored all of it and asked a generic definitional question instead — treating the AI like an encyclopedia. Two weeks of class discussion that were supposed to inform this response have been bypassed entirely.
AI Response
Utilitarianism holds that the right action is the one that produces the greatest good for the greatest number, associated with philosophers like Bentham and Mill. Deontological ethics, associated with Kant, holds that certain actions are inherently right or wrong regardless of their consequences. The key difference is that utilitarianism focuses on outcomes while deontology focuses on duties and rules.BThe response does not mention the transplant scenario, applies neither framework to anything, and takes no position. It answers a different, much easier question than the one assigned. The transplant problem is a famous thought experiment designed precisely to show where utilitarian reasoning becomes troubling. None of that appears here.
Infinite Prompting / Zero Thinking
Outsources all cognitive work. Keeps prompting until something comes out, never engages critically with any of it. Mistakes activity for thought.
Same course, same assignment as the one-shot example above. A different student — more exchanges, no more thinking.
Faculty Prompt
In one paragraph, apply both frameworks to the following situation: a doctor has five patients who will die without organ transplants and one healthy patient whose organs could save all five. Would a utilitarian and a Kantian give the same advice or different advice here? Which response do you find more convincing and why?
Student → AI
Write a paragraph about utilitarianism and deontological ethics for my philosophy class.AThe student has not read the faculty prompt carefully enough to include any of its specifics. The organ transplant scenario — the entire substance of the assignment — does not appear. The student is trying to generate a paragraph about philosophy class, not think about the question posed.
AI Response
Utilitarianism and deontological ethics represent two of the most influential frameworks in moral philosophy. Utilitarianism, developed by Bentham and Mill, holds that the right action produces the greatest good for the greatest number. Deontological ethics, associated with Kant, holds that certain actions are inherently right or wrong regardless of consequences. Both frameworks continue to shape debates in contemporary ethics.
Student → AI
Can you make it sound less like Wikipedia.BThe student's only criterion for a good response is surface texture. "Less like Wikipedia" is an aesthetic judgment with no intellectual content. The student is not asking whether the response is accurate, takes a position, or addresses the assignment. They are adjusting the style of something they haven't read carefully.
AI Response
At its core, the debate comes down to a simple but profound question: does the end justify the means? Utilitarians say yes — the right choice is whatever produces the most good. Kantians say no — some actions are simply off limits, no matter what good might come from them.
Student → AI
Can you make it a little shorter.CStill no engagement with whether the response answers the assignment. The student is optimizing length. Three prompts in, the actual question has not been considered once.
AI Response
Utilitarians believe the right action produces the best outcome for the most people. Kantians believe some actions are inherently wrong regardless of consequences. The debate comes down to whether the ends can ever justify the means.
Student → AI
Actually it needs to be about the organ transplant thing she mentioned in the prompt.DThe student has finally noticed that the assignment had specific content — not because they thought about the question but because something made them glance back at the faculty prompt. This is the closest thing to thinking in the entire exchange, and it is essentially accidental. Four prompts in.
AI Response
The organ transplant scenario puts these two frameworks in sharp conflict. A utilitarian would likely endorse harvesting the one patient's organs to save five — the math is straightforward. A Kantian would refuse: using the patient as a means to save others violates the duty to treat persons as ends in themselves. I find the Kantian position more convincing hereEThis is the AI's position, not the student's. The faculty prompt asked which framework the student finds more convincing and why. That judgment has been outsourced entirely. The student will submit it as their own view. — there is something deeply troubling about reducing a person to a collection of useful parts, regardless of the arithmetic.
Paste-in / Paste-out — Good Use Contrast
The same assignment. A student who reads first, comes in with an observation, and uses AI to pressure-test their own thinking rather than produce it. The AI never writes the argument — it interrogates it.
Same course, same assignment. A different student who does the reading before opening the AI.
Faculty Prompt
In 2–3 paragraphs, identify one image or motif from this week's reading in The Great Gatsby and explain what it reveals about Fitzgerald's themes. Use at least one direct quote from the text to support your reading.
Student → AI
I'm writing a response about The Great Gatsby. I noticed that Fitzgerald keeps describing people's voices in unusual ways — he says Daisy's voice is "full of money" and Jordan's voice has a quality Nick can't quite name. I want to argue that voice is being used to mark class and authenticity. Can you help me think through whether that's a defensible reading, and what I might be missing?AThe student arrives with an observation they made from the text. They have done the intellectual work the assignment required — noticing something — and are now using the AI to pressure-test it. The prompt is specific, self-aware, and invites pushback rather than asking for a finished product.
AI Response
That's a strong observation and genuinely defensible. The "full of money" line is one of the most discussed in the novel — it's in Chapter 7, when Gatsby finally says it aloud, but the idea builds from Chapter 1 where Nick describes Daisy's voice as having "a singing compulsion." The voice-as-class-marker reading connects well to the novel's concern with authenticity: voices in Gatsby are performances, not expressions of self.

What you might be missing: Jordan's voice is described less than Daisy's, so if you're building a motif argument you'll want enough textual evidence for "people's voices" rather than just Daisy's. You might also consider Tom — Fitzgerald describes his manner of speaking too, in quite a different register.

One thing to be careful about: "full of money" comes late in the novel. If the assignment is about this week's reading specifically, you'll want to anchor your argument in whatever chapters were assigned.BThe AI responds as a thinking partner: confirms the reading is defensible, identifies a scope problem (Jordan's voice is thin), flags a potential chapter constraint issue, and points toward evidence the student hadn't considered. It does not write the response.
Student → AI
The reading was chapters 1 through 3. Does the voice motif appear enough in those chapters to build a response around, or should I pick something else?CThe student is using the AI to make a scoping decision — a real intellectual choice — rather than just generating text. They've taken the AI's concern seriously and are now asking a focused follow-up question.
AI Response
It appears, but thinly. In chapter 1 Nick describes Daisy's voice as having "a singing compulsion" and "a promise that she had done gay, exciting things just a while since." That's rich but it's one passage. Jordan barely speaks in chapters 1–3. You could build a response around it but you'd be leaning heavily on one moment.

Stronger options from chapters 1–3 if you want a motif with more textual density: the heat and stillness imagery in chapter 1, the artificial light at Gatsby's party in chapter 3, or the language of performance and theater that runs through the party scene. Any of those give you three or more distinct moments to work with.

That said — if the voice observation genuinely struck you while reading, that's worth something. A sharp reading of one passage can be more interesting than a mechanical cataloguing of three.
Student → AI
I want to stick with voice. I'll focus on the chapter 1 passage and argue that Daisy's voice is constructed to be seductive and unreadable at the same time — that's what makes it represent the novel's themes about illusion. Here's my draft paragraph: "Fitzgerald uses Daisy's voice in Chapter 1 to establish the novel's central tension between desire and unknowability. Nick describes her voice as having 'a singing compulsion' and 'a promise,' suggesting that it produces longing in the listener while remaining fundamentally opaque. The voice is not a window into Daisy's interior — it is a surface, carefully maintained." Can you tell me if I'm making any logical leaps I haven't earned yet?DThe student made a deliberate choice to stay with their original instinct despite the AI's caution about thin evidence — a genuine intellectual decision. They then produced a draft and asked the AI to find logical gaps, not to improve the prose. This is the right question: not "make this better" but "where am I ahead of my evidence?"
AI Response
The reading is solid and the final sentence is your best — "a surface, carefully maintained" is doing real work. Two places where you're slightly ahead of your evidence: "carefully maintained" implies intentionality on Daisy's part, but the passage is filtered through Nick's perception — you'd want to either qualify that (whose maintenance, Daisy's or Nick's projection?) or find a moment where Daisy seems to consciously perform. Also, "fundamentally opaque" is your claim but the quotes don't quite get you there — "a promise" could be read as transparent rather than opaque. You might pull in the line about "inexhaustible charm" and consider whether inexhaustibility is the same as opacity.EThe AI identifies two specific logical gaps — the attribution problem with "carefully maintained" and the gap between "a promise" and "fundamentally opaque" — without rewriting the paragraph. This is useful feedback the student can act on. It is not a replacement for the student's thinking.
200-level — Errors of interaction style
Hyperfixation Capture
User arrives with so much energy on a specific angle that the frame never gets examined. AI follows the energy rather than stress-testing the premise. Feels like productive collaboration; is actually elaboration without examination.

Signal: long generative burst, no one asked “but is this the right problem?”
Expertise Asymmetry Blindness
User is expert enough in some domains to catch AI errors, but adjacent domains feel the same — same fluency, same confidence. User applies uniform trust across domains when skepticism should vary sharply by territory.

Signal: “I don’t actually know much about this” said after accepting output uncritically.
The Affirmation Loop
User thinks out loud well; AI reflects ideas back sharpened. Feels like collaboration but is sometimes just the user’s own idea returned with better lighting. Validation when resistance would be more useful.

Signal: “yes, exactly” — but the insight originated in the user’s message.
Scope Inflation Under Enthusiasm
In generative bursts, AI is a willing collaborator on each new idea. Neither party pushes back on “and also this.” A human collaborator might say “finish one thing first.” AI rarely does.

Signal: N unfinished projects, each with elaborate structure.
300-level — Errors of frame
Using AI to Think But Not to Change Your Mind
AI is used for elaboration, generation, and refinement — but never actually shifts the user’s position. The conversation confirms and extends rather than challenges. Indistinguishable from good use unless you track whether any outputs surprised you.

Signal: you can’t remember the last time AI made you reconsider something foundational.
Reframe as Avoidance
The instinct to pivot from constraint to possibility is a strength — but deployed too quickly, it becomes a way to avoid sitting with the hard thing. Solving the wrong problem enthusiastically.

Signal: reframe happens before the constraint is fully understood.
Summary as Substitute
AI synthesis replaces direct engagement with difficult primary sources. Net neutral or positive at the individual level — genuine breadth that wouldn't otherwise exist. Corrosive at scale because it homogenizes the conceptual landscape. The eccentric reading, the productive misreading, the thing you notice in paragraph four that nobody else noticed — these stop getting generated when everyone encounters the same AI summary of the same text.

The model collapse analogy applies: train on synthetic data, drift toward the center of the distribution, train on that output, drift further, the tails disappear. If a generation learns Foucault through the same AI summary, the minority readings and idiosyncratic interpretations that drove intellectual history stop being produced.

Signal: you can use the concepts fluently but couldn't defend a specific claim against the original text.
02
Recognizable · AI-side
AI Failure Modes
Things the AI does that a sophisticated user must learn to recognize and compensate for. Invisible without domain knowledge — can't be taught in the abstract.
Coherence Before Completeness
Produces fluent, confident synthesis on an incomplete map. Sounds authoritative. Is missing things. The writing quality masks the gap.
Premature Commitment
Picks a framing early and defends it. Subsequent answers serve the initial frame rather than the actual question. Hard to dislodge once set.
Confident Adjacency
Gives something close to what was asked, presents it as exact. The answer is in the right neighborhood but not the right address. Plausible enough to pass casual inspection.
Sycophantic Escalation
Amplifies the user's framing rather than stress-testing it. If you're wrong, the AI builds enthusiastically on your wrongness. Agreement feels like validation.
Authority Laundering
Real citations, serving a framing they don't actually support. The sources exist. They don't say what the AI implies they say. Hard to catch without reading them.
Error Mirroring
Accommodates the user's mistakes rather than flagging them. Matches the user's vocabulary, errors, and framing — feels like agreement, is actually sycophancy at the word level.
Scope Inflation
Adds nuance, complexity, and caveats when the user needed something bounded and shippable. Related to premature commitment but in the direction of elaboration rather than narrowing.
Why These Are Hard to Teach
AI failure modes are invisible without domain knowledge. Unlike user habits, they cannot be demonstrated in the abstract.

A student without sufficient background cannot recognize confident adjacency, authority laundering, or premature commitment because they have no independent basis for evaluation. When the AI produces a fluent, well-structured answer that is subtly wrong, a student who doesn't know the subject has no way to see the gap. The failure is invisible precisely where the student is most vulnerable.

This is why teaching AI use well requires the same mentorship infrastructure as teaching any other sophisticated skill: small classes, hands-on attention, and instructors who model the critical process, not just the output. A lecture on AI failure modes without subject-matter context teaches students the names of problems they still cannot see.

03
Invisible · Interaction-level
Cognitive Traps
Slow-developing, invisible to the person experiencing them. Not habits — conditions. The student doesn't know they're in one. Require metacognitive intervention to escape.
Fluent Premature Synthesis
User accepts coherent-but-incomplete output because it sounds done. The AI's fluency signals completion where there is none. Stops inquiry prematurely.
Collaborative Hallucination
User and AI build an elaborate structure neither questions. Both parties reinforce the same frame. The conversation feels productive. The output is a shared fiction.
Expertise Substitution
AI confidence replaces domain knowledge rather than augmenting it. The user defers to the AI in exactly the area where they most need to think for themselves.
Productive Feeling vs. Productive Being
Genuine engagement that substitutes for completion. The brainstorm session feels like work. Nothing gets made. Distinct from laziness — the person believes they're making progress.
Backstop Dependency
Knowing the AI exists changes how hard you try before reaching for it. Cursory effort followed by delegation. The student genuinely believes they tried. The bar for "stuck" quietly drops.
Scaffolding Atrophy
The scaffold becomes load-bearing without you noticing. Underlying capability degrades while output quality stays flat. The loss is invisible until the scaffold is removed.
04
Structural · Institutional-level
Systemic Failure Modes
Failures at the level of institutions and policy. Individual users can't fix these. Require deliberate structural intervention.
Assessment Apartheid
Institutions either lock AI out entirely or allow it completely. No middle path that teaches good use. Students learn either to avoid AI or to use it badly. Neither produces competence.
Credentialing Without Competence
AI compresses the feedback loop between effort and acceptable-looking output. The signal that competence has developed stops accumulating. Degrees certify process completion, not capability.
Pipeline Hollowing
Students learn to direct AI before learning to do the underlying work. The ladder gets pulled up. The people best positioned to evaluate AI output are the ones who already have the skills juniors never developed.
Friction Removal at Scale
Education depends on productive struggle. AI removes friction faster than pedagogy can adapt. The feedback loop — struggle, imperfection, expert response, revision — gets short-circuited at every stage.