Evaluate AI Sources with SIFT: A Worksheet for Grades 6-8
Picture this scene from late in a research-writing unit. A teacher is working through a stack of research papers and notices something odd: three students cited the exact same source. “Johnson, A. (2024). The Future of Machine Learning in Classrooms. Journal of Educational Technology, 18(3), pp. 44-61.” Author, title, journal, volume, issue, page range. The citation is formatted so cleanly it could be a textbook example. So the teacher opens a browser tab and searches. And searches again. The journal exists — but volume 18, issue 3 doesn’t. The paper doesn’t. The author profile doesn’t. ChatGPT had assembled a citation that looked exactly like every real citation it had been trained on, and three students copy-pasted it without a second look.
This is not a story about students cheating. It is a story about a skill gap in source evaluation that traditional frameworks were never built to address. SIFT — Stop, Investigate the source, Find better coverage, Trace claims — gives students the four-move sequence that covers this gap. This post walks through how to apply SIFT specifically to AI-generated text (not just general web sources), explains the three-tier triage log that makes the process concrete for grades 6-8, and shows what differentiation looks like by grade level and subject area.
TL;DR: Apply SIFT directly to AI chatbot output: Stop before copying any AI claim into a paper. Investigate whether the source the AI named is real — open a new tab and check. Find better coverage — locate the claim in a source you can actually verify. Trace claims — click through to the original document; if it doesn’t exist, the citation is hallucinated. For grades 6-8, give students a triage log: Verified / Plausible / Fabricated. Three columns, five minutes, every research session.
Why AI Hallucinations Break Traditional Source Evaluation

Standard source-evaluation frameworks assume the source exists. When you teach students to check whether a website is credible, you’re asking them to evaluate the quality of something real. SIFT was designed with that assumption built in: you stop, you investigate, you find better coverage. But you start from the premise that there is a source to evaluate.
AI hallucinations flip the premise entirely. When a language model fabricates a citation — a plausible author name, a real-sounding journal title, a believable volume and page number — there is no underlying document to evaluate. The credibility question is not “is this source reliable?” It is “does this source exist at all?”
This distinction matters enormously for grades 6-8 instruction. Students who have been trained on traditional source evaluation will look at “Johnson, A. (2024). Journal of Educational Technology, 18(3)” and begin evaluating it. They will try to find information about the author. They will look up the journal’s reputation.
They will never think to question whether the citation itself was assembled from whole cloth — because no prior lesson taught them that was possible.
Research published in peer-reviewed literature has documented AI chatbots fabricating references with convincing author names, plausible journal titles, and non-existent DOIs at high rates on certain query types — sometimes affecting the majority of generated references. The problem is structural, not occasional. A 2025 peer-reviewed study on ResearchGate tested AI-generated citations and found a substantial share were fully fabricated — real-looking but non-existent. (For a full set of classroom-ready hallucination examples to project, see AI hallucination examples for grade 7 ELA — five documented real-world cases with a built-in verification routine.)
The starting-point resource for introducing this concept without a full lesson is the FREE: AI Hallucination Fact-Check Worksheet (Grades 6-7-8) — a free companion worksheet for introducing the concept before the full SIFT unit.
How to Apply SIFT to AI-Generated Text, Step by Step

The four moves of SIFT map cleanly onto AI chatbot output, but each move needs a specific adaptation. Here is what each step looks like when the text being evaluated came from a language model, not from a web page.
Stop. Before a student copies any claim from AI output into a paper, they ask one question: “Did I find this fact in the AI response, or in a source the AI named?” These are two entirely different things. A claim the AI stated directly is one category. A citation the AI provided as supporting evidence is another — and it requires separate verification.
The stop moment is also the moment to notice what kind of claim is being made: a statistic (“78% of students improved”), a named study, a direct quote, or a factual assertion. Each type requires a different verification move.
Investigate the source. This is the two-tab move. A student opens a new browser tab, enters the author name + publication title + journal name + year, and runs a search in Google Scholar or a public database.
- If the source exists, it will appear.
- If the journal is real but the specific article is not, that is a partial hallucination — the model borrowed a real venue to dress up a fabricated entry.
- If nothing appears, the citation is fabricated.
This is the step that catches the Johnson (2024) problem from the opening scenario.
For this step, the technique is lateral reading of AI sources — moving away from the AI’s output and checking independent sources before reading further. Fact-checkers and professional researchers do not read a document and then decide if it’s credible. They open multiple tabs simultaneously and triangulate. Lateral reading applied to AI output means treating the AI response as a starting point, not a destination.
Find better coverage. If the AI’s cited source doesn’t exist — or if the student can’t verify the specific claim even in a source that does exist — the next move is to locate the same claim in a source the student can actually reach and read. This is not optional. “The AI said it” is not a citation. “The AI said it and I couldn’t find a source, but the claim seemed plausible” is worse. Find better coverage means identifying what kind of source could verify this claim (a news report, a published study, a government database) and going to get it.
Trace claims. Go to the original document and read what it actually says. Does the source make the claim the AI attributed to it? AI-generated text sometimes cites real sources but misrepresents them — stating that a study found something it didn’t find, or quoting a statistic from a different context. Tracing the claim to its origin catches selective misrepresentation, not just outright fabrication.
The standards this four-step process hits directly, by anchor code:
- CCSS.ELA-LITERACY.W.7.8 — “Gather relevant information from multiple print and digital sources, assess the credibility and accuracy of each source.”
- ISTE 1.3.b — “evaluate the accuracy, perspective, credibility and relevance of information, media, data or other resources.”
The lateral reading technique is the practical application of both.
What Does the Three-Tier Triage Log Actually Look Like?

The triage log is a three-column student record. Every AI-generated source gets one row. Every row gets one of three verdicts.
Verified: The source exists. The student found it in Google Scholar, a library database, or the publisher’s website. The source says what the AI claimed it said. The student can paste a working URL or DOI next to the verdict.
Plausible: The source exists — the journal is real, the author has a publication record — but the student cannot verify the specific claim the AI attributed to it. The article might not be accessible without a library login. The statistic might appear in the abstract but not in the detail the AI provided. Plausible does not mean correct. It means unresolved, and it requires a note about what the student tried and could not confirm.
Fabricated: The source does not exist. The journal is real but the issue or volume number doesn’t match any published edition. The author name produces no results connected to the cited topic. The DOI leads to a 404 or an unrelated paper. Fabricated means do not cite — not “probably fine.”
Here is what a completed row looks like in practice. AI output states: “Johnson, A. (2024) found that 78% of students improved reading fluency after four weeks of AI-assisted practice. Journal of Educational Technology, 18(3), p. 44.” Student verification: search “Johnson AI reading fluency Journal of Educational Technology 2024” in Google Scholar. Result: no matching entry. Search the journal directly at the publisher’s site: volume 18 exists, issue 3 exists, but the article is not there. Verdict: Fabricated. The student does not cite it. The student goes to Find Better Coverage.
The triage log works because it gives students a structured place to record what they actually did during verification — not just a checkbox. A student who writes “I searched and didn’t find it” has produced something a teacher can read and comment on. A student who skips the log produces nothing that can be coached.
The print-ready version with a structured triage log, a grade-level differentiation rubric, and an answer key is the AI Source Evaluation Lesson | SIFT Hallucination Triage | Grades 6 7 8 ($8) — the ready-to-photocopy version with grade-level differentiation and an answer key, designed so a teacher can run it on a Monday with no additional prep.
How to Differentiate SIFT Verification by Grade Level
SIFT is not one lesson — it is a progression. The cognitive demand of full source verification exceeds what most grade 6 students are ready to sustain independently, and the solution is not to skip the process but to scope it.
| Grade | Verification Task | Time Required |
|---|---|---|
| 6 | Verify 1 citation per AI output — focus entirely on “does this source exist?“ | 5-7 minutes |
| 7 | Verify 2 citations — add “does the source say what the AI claims?“ | 8-10 minutes |
| 8 | Full source audit — verify all citations + check for selective quoting | 12-15 minutes |
Grade 6 enters SIFT through the single tightest question: does the source exist? A grade 6 student who runs the two-tab Google Scholar check and records a Verified / Fabricated verdict has accomplished something meaningful. The Investigate step is the whole assignment at this level. Do not ask them to also evaluate credibility, cross-reference claims, and trace quotes in one sitting. That is four skills, not one.
Grade 7 adds the second question: does the source actually say what the AI claims? This is the distinction between a fabricated citation (the source doesn’t exist) and a misrepresented one (the source exists but the specific claim is wrong or decontextualized). A grade 7 student verifying two citations will sometimes discover that both exist — but that one of them doesn’t contain the statistic attributed to it. That discovery is the lesson.
Grade 8 runs the full audit: all citations verified, plus a check for selective quoting. Grade 8 students can handle the additional complexity of finding a source, reading it, and identifying whether the AI’s use of it was accurate or cherry-picked.
The inquiry-based structure here aligns with ISTE 1.3.d — “build knowledge by actively exploring real-world issues and problems” — because a grade 8 student doing a full source audit is doing genuine knowledge-building through investigation, not just completing a worksheet.
A teacher running this with a mixed-grade class could assign different triage-log row counts by grade level and review the same lesson together, debriefing in whole-group at the end.
SIFT Across Subject Areas: ELA, Social Studies, and Science
The triage log does not belong only in ELA research units. AI-generated source hallucinations appear wherever students use chatbots to gather supporting evidence — which now means nearly every subject area.
ELA scenario. A student drafting an analytical essay asks a chatbot to “find three quotes from the novel that show the protagonist changing.” The AI produces three quotes with chapter and page number citations. Two of the quotes do not appear in the novel. One is a real quote but from a different chapter than the one cited. SIFT’s Trace Claims step is the fix: the student goes to the actual text and checks the page number.
In an ELA context, this is a skill students already have — close reading of primary texts — applied to a new problem. The primary text is the novel, not the AI output.
Social studies scenario. An AI response to a research prompt about immigration policy cites “a 2023 Congressional Budget Office report” showing a specific economic impact figure. The CBO is a real institution, and students have been taught it’s a credible source. But when a student runs the Investigate step and searches the CBO’s published reports page, no 2023 report matching that description exists.
The credible-sounding institutional name made the fabricated source harder to catch, not easier. That is the social studies lesson: a real institution attached to a fake document is still a fabricated citation.
Science scenario. An AI response to a question about climate modeling states that “three peer-reviewed studies found…” but provides no author names, no journal names, and no dates. There is nothing specific to investigate. The Find Better Coverage step is the only move available: locate a verifiable study on the same claim through a science database or news source.
A student who has practiced the triage log knows that “the AI mentioned three studies” is not a citation — it is an uncited claim, and uncited claims go in the Plausible column until a verifiable source replaces them.
The through-line across all three scenarios: SIFT’s four steps distribute differently by subject, but the triage log structure stays constant. Verified / Plausible / Fabricated works in every class, every research context, every grade level.
From One-Time Lesson to Daily Research Habit
A single SIFT lesson does not produce a verification habit. It produces one completed triage log. The goal is for the three-column check to become as automatic as re-reading a paragraph before submitting — built into the research process rather than added as an afterthought.
The practical mechanism is a standing research-assignment instruction. After running the SIFT lesson once, every subsequent research assignment for the rest of the year includes one line in the directions: “At least one AI-generated source must go through the triage log before you cite it.” The triage log template — three columns, five rows — lives on the classroom wall and on Google Classroom as a permanent resource students return to rather than a one-time handout they lose.
When the habit is established, something else happens: students begin applying the three-verdict vocabulary without being prompted. A student who identifies a fabricated citation in a social studies assignment and says “this one’s fabricated — the journal doesn’t have a volume 12” has internalized the framework, not just completed a worksheet.
That is what habit formation looks like in a grades 6-8 classroom.
The stakes matter here beyond the assignment. AI4K12 Big Idea #5 addresses Societal Impact — the principle that AI can affect individuals, communities, and society in both positive and negative ways. When a hallucinated citation passes unchecked from a student paper into the work record of a classroom, a school, or eventually a professional context, misinformation circulates as attributed knowledge.
The verification habit students build in seventh grade is the same habit that determines whether a hallucinated source propagates or stops. Teaching it matters beyond the research-writing unit.
For additional free AI literacy resources, the AI literacy free resources page at /free includes downloadable starter materials available with no purchase required. If your students need vocabulary grounding before the triage log makes full sense — particularly the term “hallucination” itself — the AI vocabulary worksheet for middle school covering 60 terms builds the definitional foundation for grades 6-8, including the distinction between hallucination and simple error.
The triage log does not require technology. It does not require a subscription. It requires a photocopy, a browser with Google Scholar, and fifteen minutes of structured practice. Most research units already have that time. The three-column habit just needs a place to land.
For a broader view of where source verification fits in the AI literacy curriculum — alongside vocabulary, prompting, and ethics — the teaching guide maps the full scope for grades 6-12.
This post was drafted with AI assistance and human-finalized.
Quick questions
When you apply SIFT to a website, you're evaluating a source that exists — you check who published it and whether other sites corroborate it. When you apply SIFT to a ChatGPT answer, the source the AI cites may not exist at all. The 'Investigate the source' step becomes a two-part check: does this source exist? And does it say what the AI claims? That second layer is what hallucination triage adds to standard SIFT instruction.
Treat the AI output as a first draft of a source list, not a finished list. Give students a three-column triage log: Claim / AI-cited source / Verified source. For every claim they want to use, they must fill in column three. If they can't find a real source that backs the claim, the claim is cut. This workflow keeps AI as a brainstorming tool while building the verification habit — and it maps directly onto SIFT's 'Trace claims' step.
Get the free AI-Proof Assignment Toolkit
10 ways to redesign any assignment so an AI chatbot structurally can't do it — plus a redesign worksheet, a 45-minute lesson, the “Spot AI Work” card, and parent templates. One email, all 5 pieces.
Prefer the full breakdown? See everything inside the toolkit →