Subscribe to the 10 Minute Teacher Podcast and Cool Cat Teacher Talk anywhere you listen to podcasts.
This is #10 in the Top 10 of the First 1,000 — a countdown of the ten most-downloaded episodes in the history of the 10 Minute Teacher, running weekdays through September 18. Ranked on total downloads including re-airs, this conversation with Michelle Blanchet has been heard 23,379 times.
It originally aired on September 27, 2021 as episode 764. I am re-airing it unchanged because the problem it solves did not go away.
Michelle Blanchet, co-author of The Startup Teacher Playbook, teaches a method for figuring out the specific thing that is wrong — not the vague heavy feeling, but the actual named problem — so that you can make an improvement and feel forward progress. Social emotional learning is not just for students. This one is about you.
I think this episode did well because it hits a topic that happens to all teachers — we get upset! Not only do we get upset but when we do it negatively impacts our ability to be able to connect with students and just do our job. Michelle’s advice resonated with so many teachers! I hope you enjoy this show! ~ Vicki
Listen to the Show
Resources from This Episode
The Startup Teacher Playbook by Michelle Blanchet and Darcy Bakkegard — the book behind this conversation, published by Times 10.
The Educators’ Lab — the organization Michelle founded to support teacher-driven solutions to problems in education.
About Michelle Blanchet
Michelle Blanchet
Michelle is an educator striving to improve how we treat, train, and value our teachers. After ten years of experience working with young people, she founded the Educators’ Lab, which supports teacher-driven solutions to educational challenges. Michelle earned a master’s in international relations from Instituto de Empresa in Madrid. She has taught social studies in Switzerland and the U.S. and has presented at numerous events, including SXSWedu and TEDxLausanne. Michelle is a part of the Global Shaper Community of the World Economic Forum. She has worked with organizations like PBS Education, the Center for Transformative Teaching and Learning, Ashoka, and the Center for Curriculum Redesign. The Startup Teacher Playbook through Times 10 Publications is her first book. She occasionally blogs for Edutopia.
As the 10 Minute Teacher approaches its 1,000th episode, I went back through the download numbers for every show and pulled the ten that teachers listened to most. One airs each weekday through September 18, counting down to number one.
If this one gave you something to try this week, share it with the teacher across the hall.
Disclosure of Material Connection: This post includes affiliate links to books, which means that if you choose to buy I will be paid a commission at no additional cost to you. I am disclosing this in accordance with the Federal Trade Commission’s 16 CFR, Part 255: “Guides Concerning the Use of Endorsements and Testimonials in Advertising.”
Subscribe to the 10 Minute Teacher Podcast and Cool Cat Teacher Talk anywhere you listen to podcasts.
AI with intention. AI with integrity. Is it possible? Tony Frontier, author of AI with Intention, is talking about how we value learning in the age of AI. He shares the language we can use to evaluate what we’re doing, how we’re talking about AI, and how we’re measuring learning, and ultimately, why education needs to change more than ever.
I also open this show with some highlights from the Report of MIT’s Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training and what it means for learning, and they use a term I think we should all be discussing: cognitive surrender. We cannot let our kids cognitively surrender and let AI do the work for them. They also talk about helping students understand that their learning, integrity, and judgment are the product of college, not the grades. So many challenges we have now, but this week has been about sharing the conversations and thought leaders to help you improve the AI conversation in your school. I hope this Thursday helps you be a thought leader in your school!
SPONSORED:Experience AI, a free program co-developed by the Raspberry Pi Foundation and Google DeepMind, sponsored this episode. All opinions are my own and those of the guest.
Experience AI is a free AI literacy program with ready to teach lessons for ages 8 to 16. The lessons are written for teachers of all subjects including: science, math, social studies, English, and art and some don’t even need a computer to use. Support your students to become critical thinkers and better understand AI technologies today. Explore the free resources at Experience AI.
If the next voice is, I must not be very good at this, I must not be very smart — that’s a kid who is going to stop trying. And that’s where we lose kids… That’s a kid who has confused learning with compliance. And we did that to that kid.
Dr. Tony Frontier, Director of the AI Center for Effective Teaching and Learning
The AI Policy Dr. Tony Frontier Says Actually Works
You can listen to the whole conversation in the player above or read the full transcript below. Here is the framework Tony walked through, piece by piece. He has been running focus groups in high schools around the country — urban, suburban, rural, public, private, parochial — and his first finding is the one that should worry us: you cannot tell one school’s transcript from another.
Kids are using AI tools. Typically, kids say they use it twenty-five to forty percent of the time for their schoolwork… And there’s a group of kids say they use it all the time. I’ve had kids say they haven’t read anything since they were a freshman. They just have AI summarize everything.
— Dr. Tony Frontier
The second finding is the one we can actually do something about. Students told him they are not getting instruction. “There’s some vague permission, there’s a lot of outright bans, but there’s very little focused, specific lessons. Here’s how, here’s why.”
The 3 Choices Schools Have Right Now
Tony frames three approaches schools are taking. He says, “if you have access to a search bar, you have access to an AI tool that can do pretty much all of your homework, kindergarten through PhD.” So there are three doors:
Let kids figure it out on their own. “A lot of videos out there on TikTok about how to use AI to cheat.”
Let the tech companies decide. Just let the kids do everything.
Take the lead. “The kids are ahead of us right now and we have to get ahead of the kids.”
✏️✏️✏️
Part 1: Guidelines and Boundaries
Boundaries say what is off-limits. Guidelines describe what good looks like. Tony’s warning is that most schools write only the first kind and think they are done. Wow, are we giving kids guidelines? Are we modeling good uses?
Do you have any good uses of AI that you can model for your students tomorrow? Can you show them the chat and what you did? Can you show them how you figured out AI was wrong on something or why you made other choices than the AI did for the work. This is part of disclosure and transparency, but also modeling and giving guidelines.
Part 2: Define Effective, Ineffective, and Inappropriate Use
Sit down and define what these look like in your school. Sit down with paper and write one-sentence definitions for what this looks like. Tony’s point is that the definitions have to be specific enough to provide direction but general enough not to micromanage a teacher’s judgment or expire the moment a tool changes.
Effective — aligned and intentional.
Ineffective — “I used AI to complete my work, but I can’t explain the work that I completed.”
Inappropriate — privacy violations, cheating, “things that are completely inappropriate.”
This is the part that Tony says many schools skip entirely, but that is essential. Remember that whatever we ask of students, we model. I’ve included quotes from Tony for each of these, but listen to the show for the full context. (And remember his book AI with Intention: Principles and Action Steps for Teachers and School Leaders has many of these items.)
3 Commitments for Students:
Integrity — “Can a student say I used AI tools in ways that are aligned to my teacher’s communicated expectations?”
Transparency — “I document and report any tool or resource I use.”
Explainability — the student “is always accountable to explain, justify, or further explain the work that they did.”
3 Commitments for Teachers
Fidelity — is this aligned to stated priorities for teaching and learning? “I don’t mean everyone’s on page 263 of the teacher’s manual.”
Transparency — “We need to model the transparency we expect from our kids.”
Explainability — “If I used AI to create a slideshow or an assessment, I should be able to say, here’s why this specific academic language is here.”
Tony really implies these aren’t for “gotchability” but a framework for intentionality and admin support within the commitments.
Not because we’re looking to catch people doing the wrong thing… If I’m a teacher and I’m awake at two AM worried about something, and I know I used AI for something the day before, I want to know that I have administrative support.
— Dr. Tony Frontier
Make “Who Helped You With This?” a Normal Question
One point here is important to hear: “I’ve yet to hear a compelling reason why kindergarten through fifth-grade kids should be using LLMs.” He also talks about making transparency with who did the work part of what is done, to level the playing field as much as anything.
We should just make it normal starting in kindergarten. Who helped you with this? My mom helped me, or my older sibling helped me… Because if I don’t know what you can and can’t do independently, I can’t teach you.
— Dr. Tony Frontier
In my AP Computer Science Principles classroom, my students can use AI to write code. However, they must be prepared to explain every line of code, what it does, and how the code works. If they cannot, I reserve the right to have them remove the code to see if it still works. I care more about learning detection than AI detection. In fact, students who resisted AI in that course, in my opinion, did not do as well as those who used it wisely for learning support.
Cognitive surrender — Falling back on AI at the first hint of struggle, and getting the illusion of learning instead of the learning. This is not Tony’s phrase and it is not mine — it comes from the Report of MIT’s Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training (August 25, 2026), which I read from at the top of this episode. MIT’s point is that giving in to that option cheats students out of the cognitive friction that actual learning requires.
Explainability — A student’s accountability to explain, justify, or further explain the work they turned in. ⚠️ Be careful with this one: in the AI industry, “explainability” means something else entirely — whether a human can understand why the model produced a given output. Tony is using it about the student, not the machine. Both meanings are legitimate; they are just not the same word doing the same job. Source: Dr. Tony Frontier, AI Center for Effective Teaching and Learning.
Agency / agentic learner — A learner who asks questions, seeks clarification, wants personal relevance, and can say “help me out here, I need that differently.” ⚠️ Same caution: “agentic AI” in the tech press means software acting on its own. Tony’s “agentic learner” is about the human doing the acting. As he puts it, our most agentic learners are our kindergartners. Source: AI Center for Effective Teaching and Learning.
Humanizer — A tool that rewrites AI-generated text so it will not trip an AI detector. There is no official definition for this one; it is student vocabulary. Tony’s finding is the part worth knowing: some students start using humanizers after being falsely accused — “I didn’t use AI, but once I got accused, now I run everything through a humanizer.”
These are new to my AI Vocabulary List — I’m adding them. Want to find out how many of these words you already know? Play Spy the AI, the free vocabulary game I vibe coded from that same list.
Stop Doing AI Detection. Start Doing Learning Detection.
As I said in the show, AI detection whack-a-mole, to me as a teacher, is pretty pointless. I have moved to at least two oral assessments every grading period. So I as told him: we shouldn’t be about AI detection; we should be about learning detection. He says the same thing in a similar way.
He said “Exactly” — and then told me the session he had just given is called Catch Them Learning. That is great!
Here is why the switch matters, in his words. It changes who carries the burden of proof. I think this table helps explain it.
Chasing cheating
Building integrity
“Did the student actually do the work?”
“Does this work actually represent the student’s true knowledge and skills?”
“Now it’s on me to prove that you cheated.”
“Integrity is a shared responsibility between the student and the teacher.”
A policy about what we’ll do if we catch you
A policy about what we’ll both do to make the work real
What a lot of people realize is their academic integrity policy — it’s really a what we’ll do if we catch you cheating policy. It’s really about what not to do.
— Dr. Tony Frontier
As he said earlier, students need to be modeled what to do with AI and not be given a list of things that you can’t do.
What To Watch Out For: Tony had some lines we need to read.
The arms race you cannot win.(13:48) — “The technology arms race is always going to be one that the students can win. Simply because there are more of them than there are of us.”
False accusations create the behavior.(13:48) — “I’ve had a few kids in a few different schools say, I didn’t use AI, but once I got accused, now I run everything through a humanizer.” As Tony puts it: “it’s like the teacher forced them into using AI.”
Detectors catch the wrong students.(13:33) — my own line, and Tony confirmed it from his focus groups: the only kids who get caught using AI are the ones who don’t know how to use AI.
Policies that try to put the genie back.(17:16) — see the Blockbuster section below.
Where We Lose Kids
I asked Tony a question that I, as a mom of a couple of dysgraphics, wonder: does it make it okay if you can pay to have a human writing coach, but not okay to turn AI into a writing coach? His answer started with IEPs and case-by-case judgment and he says that is important. But then, he makes a great point that speaks to equity.
When my kids came home from school, stuck on their schoolwork, they had a dad with a PhD and a mom with a master’s degree to help them. No one ever asked, right, did your parents help you with this? But we’re hyper-concerned now about, did I help you with this? We should have been asking those questions all along.
— Dr. Tony Frontier
Then he described the moment a student’s inner voice turns. What happens when a kid says “I’m stuck, I don’t get it?” But it can happen often by second or third grade, sometimes as early as the end of first, the next voice becomes “I must not be very smart.” That kid stops trying. And the only tool they can think to reach for is one that hands them the answer. This is not good. (It is the cognitive surrender that MIT talked about in their report and one we need to be watching out to prevent!)
Our Most Agentic Learners Are Kindergartners
Our students with the greatest level of agency in our schools, they’re our kindergartners. They come in, they ask questions, they want to know, they want to tell you. That’s how we want kids to interact with AI tools.
— Dr. Tony Frontier
And the skills turn out to be the same ones we already teach: “Asking questions, seeking clarification. If something isn’t working for your learning, can you stand up and say, help me out here? I need that differently.”
An Image to Help: The Blockbuster Boardroom
When Tony hears schools talk about ramping up policies to catch cheating, this is what it sounds like to him:
Some of the conversations that were probably happening at the Blockbuster video boardroom right around 2010. Where it was like, you know, if we just get this late fee thing right, everything’s gonna be fine. No — that was a transactional solution to a transformational challenge.
— Dr. Tony Frontier
The transformational challenge, he says, is that how people receive and consume entertainment had completely changed. It did not matter how hard Blockbuster chased the late fee policy. It was the wrong conversation.
And perhaps instead of the AI “whack-a-mole” detection impossibility (wrong conversation) we should instead be having a conversation about what truly makes education (the right convo to have, perhaps?)
Tony says: “We have an education model that’s built on 1890s technology. We’re gonna batch process kids, roll them down the conveyor belt” — same page, same pace, same time, because that is how you manufacture things. His argument is that these tools, used well, finally make it possible to stop.
Try This Tomorrow: The Quarter Sheet of Paper
When I told Tony this kind of conversation takes time, he agreed. Then he gave the single most usable thing in the episode. Students upload the assignment the night before. The next day, on a quarter sheet of paper:
Here is an important word from last night’s homework. Write your definition of it.
What was the main idea?
What was your biggest takeaway from last night’s assignment?
No, you cannot grade a hundred and forty of those every day. That is not the point.
It’s not just turning it in. I want to know what you know. I want to know what you don’t know. I want to know what questions you have. It starts to shift the culture.
— Dr. Tony Frontier
I remember in 1991 I started teaching people how to use the internet. I remember people being scared. Others were angry. And I received criticism that I was somehow a bad person because I believed people needed to learn to use the Internet properly and civilize it. Internet used to be a proper noun. In 2016, we started using internet in lowercase. That is how much the internet became part of everything we do. Enough to no longer be a proper noun! It is no longer a question of whether we use the internet; it is how, when, and why. We need intentionality there, too.
The criticism that came with Web 2 stung a little less. But by now, I just move forward. AI is the most powerful technology we’ve ever encountered as a general populace, and it comes with far more responsibility.
This technology is literally different for everyone who uses it. Previously, if I logged in to an app, I might have a similar experience to yours. No more. A thousand tiny nuances from how we speak and type to what we ask can cause our experiences to be far different. The challenges with bringing student-facing and teacher-facing AI to schools are many, but it is here, and we have to move forward and have the conversations that matter. AI detection is just so not the convo to be having.
I hope that you’ll pick up Tony’s book and use what I’ve shared with you here to add to the conversations you’re having at your school.
The AI Center Implementation Toolkit — the toolkit Tony references on air, with the worked examples of effective, ineffective, and inappropriate use. It includes a four-phase planning process plus the full set of guidelines, boundaries, and shared commitments.
The AI Center for Effective Teaching and Learning — Tony directs it. Also home to the AI2TPS teacher-perception survey and the student focus-group protocol, if you want to run this in your own school.
Experience AI — this episode’s sponsor. Free AI literacy lessons from the Raspberry Pi Foundation and Google DeepMind.
About Dr. Tony Frontier
Dr. Tony Frontier was a teacher in Milwaukee Public Schools and a high school principal, K-12 curriculum director, and University Professor. He is currently the Director of the AI Center for Effective Teaching and Learning where he works with schools and districts to ensure AI tools are used in ways that support, rather than undermine, effective teaching and learning. He is author of 6 books and numerous articles. His most recent book, AI with Intention, Principles and Action Steps for Teachers and School Leaders with a foreword by Jay McTighe, and has been praised by John Hattie as “the best book I have read to introduce all to optimize the power of AI.”
Other Shows for Teachers and Administrators Writing AI Policy
Creativity and Art in an AI World — Susan M. Riley on why creativity is the defining skill of the AI era. This is the Monday show where we used the phrase “hard fun,” which MIT uses in this report too.
Disclosure of Material Connection: This is a sponsored episode and blog post. Experience AI has compensated me to share information about the Experience AI program. However, all opinions expressed are my own. I have personally reviewed these resources and only recommend tools I believe offer genuine value to classroom teachers. My endorsement is limited to the educational products and services discussed in this episode. I am disclosing this in accordance with the Federal Trade Commission’s 16 CFR, Part 255: “Guides Concerning the Use of Endorsements and Testimonials in Advertising.” The sponsor has no impact on the editorial content of this show.
Disclosure of Material Connection: This episode includes some affiliate links. This means that if you choose to buy I will be paid a commission on the affiliate program. However, this is at no additional cost to you. Regardless, I only recommend products or services I believe will be good for my readers and are from companies I can recommend. I am disclosing this in accordance with the Federal Trade Commission’s 16 CFR, Part 255: “Guides Concerning the Use of Endorsements and Testimonials in Advertising.” This company has no impact on the editorial content of the show.
Subscribe to the 10 Minute Teacher Podcast and Cool Cat Teacher Talk anywhere you listen to podcasts.
Tim Plaehn is a teacher who writes about honor codes. In his writing class last April, he believed he had two students he knew used AI. So, he gave the whole class an opportunity to “come clean,” and he said his email began dinging while 22 of the 40 students admitted they had used AI on their papers. Tim has authored the book, The Honor Code: Students, Integrity, and our Path Forward , where he talks about integrity and how we need to be discussing integrity with our students. This is not an easy topic but an important one!
SPONSORED:Experience AI, a free program co-founded by the Raspberry Pi Foundation and Google DeepMind, sponsored this episode. All opinions are my own and that of the guest.
Experience AI is a free AI literacy program with ready to teach lessons for ages 8 to 16. The lessons are written for teachers of all subjects including: science, math, social studies, English, and art and some don’t even need a computer to use. Support your students to become critical thinkers and better understand AI technologies today. Explore the free resources at Experience AI.
Tim spent ten years as faculty chair of the honor council at Asheville School. No, their school didn’t publicly shame students, but the situations were hard nonetheless. He was the one who sat across the table from the students. He also knows the consequences if students aren’t held accountable (and sometimes even when they are.)
He also talks about some restorative practices. For example, a letter the student writes to their parents about the situation (that never gets sent.) We have so many issues with students these days and Tim shares some practical ideas and thoughts that can help shape our conversations about honor and integrity in school today.
I want to give you a chance to restore your integrity because you’re headed off to college and it’s going to be all you.
Tim Plaehn, American Studies teacher at Asheville School
Middle and high school teachers who have caught it and don’t know what to do next — and the deans, principals, and department chairs who write the policy they’re working under. If you have ever wondered whether your school’s response to cheating actually changes anybody, this is your twelve minutes.
Key Takeaways for Teachers from Tim Plaehn
Separate conduct from honor. (06:31) Out of the dorm after lights gets a detention; cheating on a test gets a conversation with the honor council chairs about what you did and what you’re going to do about it. Two different problems were getting one response, and Tim says pulling them apart is what made the whole thing work.
Offer restoration before you run detection. (10:57) Tim didn’t put the papers through software — he emailed forty-five students and gave them a way to come back, and twenty-two of them took it that night. A detector would have found the two he already knew about, but he would have missed that over half the class used AI. (As I’ve said before, because DETECTORS DON’T WORK!)
Have them write the letter, then don’t send it. (07:11) Students write to their parents about what they did — and Tim tells them up front it isn’t going anywhere. It’s not the punishment; it’s the part where they have to say it in their own words.
Answer “I know my child’s heart” honestly. (02:58) No parent knows their child’s heart well enough to rule out a mistake, and Tim’s response is the one I keep coming back to: when a kid asks “Dad, don’t you trust me?”, the answer has to be no — not because they’re bad, but because they’re fifteen.
The next time a student owns up to cheating, don’t start with the consequence. Hand them paper and have them write a letter to their parents explaining what they did and why — and tell them before they start that you are not going to send it. Tim’s whole point is that the letter isn’t evidence and it isn’t a punishment. It’s the only step in the process where the student has to put the thing into their own words, to the people whose opinion of them they actually care about. It takes one class period and costs nothing.
The part of this conversation that will divide a faculty room is the fourth line of Tim’s honor code — the one where students agree to report violations. He is candid that it is the piece most often broken, and he doesn’t pretend there is a clean answer. What he offers instead is a way of framing it that doesn’t turn a school into an informant culture: snitching, he says, is what you call it when a community is protecting the person who did wrong instead of the community itself.
I mention the Common Sense Media research in this episode. Here is the current number: 86% of kids ages 9 to 17 use AI, and nearly a quarter use it every day — 81% of 9-to-12-year-olds, 89% of 13-to-15-year-olds, and 92% of 16-and-17-year-olds. More than four in ten say no parent or guardian has ever talked with them about AI safety, and while three-quarters say their school has covered what they can and cannot use AI for, only just over half have been taught how to use it safely. Source: The Common Sense Media Census: AI Use by Tweens and Teens (2026), released June 8, 2026, surveying 1,204 children.
Resources Mentioned in This Episode
The Honor Code: Students, Integrity, and Our Path Forward — Tim’s book, built on four real honor cases he fictionalized to protect the students in them. He explains why at 01:33: at his school, an honor case is nobody else’s business.
Tim Plaehn’s Substack — where he writes about integrity in schools and in the wider culture, and publishes excerpts from the book.
Asheville School — the North Carolina boarding school where Tim teaches American Studies and chaired the honor council for ten years. The campus detail he describes at 08:01 — laptops and backpacks left out in the open because nobody expects them to be taken — is the culture the honor code is protecting.
Editor’s note on the West Point honor code: Tim paraphrases it in conversation. The code itself reads, “A Cadet will not lie, cheat, steal, or tolerate those who do.” Asheville School’s own code — the one whose fourth line Tim is describing — reads, “I will not lie, cheat, or steal, and I will report any violations of the honor code.”
About Tim Plaehn
Tim Plaehn, teacher and honor code advisor
Tim Plaehn is a writer and teacher living in Asheville, North Carolina. Over his thirty-year career in education he has taught in Las Vegas, Hartford, Atlanta, and Asheville and has earned advanced degrees from Harvard’s Graduate School of Education and the Bread Loaf School of English. He has been awarded Clark County’s New Teacher of the Year Award, the Charles N. Carter Leadership Award in Coaching, and the William F. Lewis Faculty Chair for Teaching and Coaching Excellence, among others.
His first screenplay, The Panjiayuan Diary, won the Beijing International Screenwriting Competition, and he followed that up with The Reconstruction of Huck Finn (Over Mark Twain’s Dead Body!), a Nicholl Fellowship semi-finalist in 2017. His play West Asheville was selected to The Barefoot Theatre’s reading series at the Art of Acting Studio in Los Angeles and for a workshop performance at the Cherry Lane Theatre in New York. His short play Jenna Feldman, A One-Woman Show won the audience favorite award at The Hickory Playground’s One-Act Play Festival, and Metallica Is the Last Straw was selected for a performance at Toronto’s InspiraTO Festival. He’s been published in Short Story America and Creative Loafing and has written screenplays on sumo wrestling, golfing across America, and a family reunion in Ireland gone terribly wrong. That last screenplay had a Top 3 finish in The Nantucket Film Festival’s screenplay competition in 2022, prompting Olivia Wingate to sign on as producer. He was also awarded Ireland’s Aran Islands Poetry Fellowship, and he spent a month cold, wet, and happy on Inis Mor.
Other Shows for Teachers and Administrators Facing the Honesty Question
AI and Academic Honesty: The Second Crisis in Schools — philosopher Dr. Christian Miller, author of The Honesty Crisis, on why the internet was the first honesty crisis and AI is the second — and why it runs on both sides of the desk. If this episode landed, that one is the companion.
If your school is still arguing about detection software, send this one to whoever writes the policy.
Episode Transcript
This transcript was generated using AI and has been reviewed by humans for accuracy. Minor errors or artifacts may remain but I worked my best to find any issues with the transcript as I reviewed the show. – Vicki
Click to read the full transcript
Vicki Davis (00:00): Happy Wonderful Classroom Wednesday. This is episode 982, and we’re talking about honesty. We’re talking to a teacher who works with the honor code. He thought that only two papers out of 45 had been written by AI, but when he emailed the whole class, over half. Let’s learn.
Announcer (00:21): This is the 10-Minute Teacher Podcast with your host, Vicki Davis.
Vicki Davis (00:25): Today’s show is sponsored by Experience AI, the free AI literacy program from the Raspberry Pi Foundation, co-founded with Google DeepMind. Stay to the end. I’ll share how you can start teaching AI literacy in your classroom with confidence. No experience required. What if the reason cheating keeps winning in so many schools isn’t that kids have gotten worse? It’s that we’ve been treating honesty like a rule instead of teaching it like a subject. Our guest today has spent 10 years as the head of the honor council at a school in North Carolina, watching real students face real consequences. Tim Plaehn teaches American Studies at Asheville School, holds a master’s from Harvard’s Graduate School of Education, and was awarded Aran Islands Poetry Fellowship. And he has just published The Honor Code: Students, Integrity, and Our Path Forward. Tim, your book is built on four real cases. Pick one of them. The moment a student sat across from you and you realized this wasn’t going to be simple.
Tim Plaehn (01:33): Well, first of all, I fictionalized the stories. One of our big principles at our school is that, this is no one else’s business. It’s really just the chairs of the honor council talking with the student and kind of keeping it in-house. So what I did, I took real cases, I changed the names and I changed the details and I kept the kernels of truth. One of my favorite students who made a terrible decision one day, he walked into another student’s dorm room, took $100 and went back to his room, immediately regretted it and came to me with $100 and said, what do I do now? So to kind of hold his hand through that process of stealing is an immediate expulsion at our school. And yet this kid, by coming forward and talking it through and expressing his regret, we ended up giving him a second chance. We felt like he was teaching, himself his own lesson before we even got to him. And then the story ends during his senior year when he steals from a store and he ends up getting expelled right before graduation.
Vicki Davis (02:43): Oh, my.
Tim Plaehn (02:44): And it’s just a reminder that kids make bad decisions. And we have to be there, not just to bear witness, but to help them process these decisions that they’re making so that they don’t make them as they become adults.
Vicki Davis (02:58): Well, Tim, one of the challenges, and, you know, I’ve been teaching 24 years, is that I believe early on in my career, not always, But for the most part, students and parents would own up to issues. But I’ve seen several cases, you know, in the past 10 years where we had it on film. We knew the student did it. And we’ve even seen this in some very public cases where you see that the person did it. And the parents say, I’m not even looking at the tape, no matter what they said. If they say they didn’t do it, they didn’t do it. That’s not reality. Kids do make mistakes. How do we deal with this dysreality of those who claim their child would never lie? I had a parent one time say, I know my child’s heart. My child would never lie. And I’m like. You don’t know your child’s heart because all humans can make mistakes and can lie.
Tim Plaehn (03:50): It’s so funny. You know, you think of like Hollywood movies or TV shows where the kid says to the parent, Dad, don’t you trust me? The answer has to be no. Like as parents, we have to know that these are teenagers and these are flawed human beings and they’re going to make mistakes and they’re going to try to wriggle out of tough situations. You said you are a parent of three. Yes. And as my kids were growing, I said to myself, I’m always taking the teacher’s side. I’m always taking the school’s side. Did you have a similar commitment?
Vicki Davis (04:21): Well, you know, our commitment was because we had two with learning differences. And so we had advocacy that had to happen. So in front of my child, I was always going to be taking that teacher’s side. I didn’t want my child to feel like, oh, I need to give my parent more ammo. However, there were times where they might take me something and I might privately go to the teacher. Or I would teach my child to advocate for themselves. One of mine could not copy off the board without making mistakes. And his math teacher kept writing tests on the board. And so he would copy down 20% of the problems wrong, as we knew he would. And he would work the wrong problem, but he would work it in the right way. But understanding that authority exists, right and wrong exists, that makes the fabric of a home and a society hold together. When you start conditionalizing truth, we have problems. Yeah.
Tim Plaehn (05:19): You’re in trouble. And that’s one of the reasons I wrote this book is to kind of start a conversation about this and how we all approach this. And I think even when I was frustrated with one of my children’s teachers, it’s all about conversations. It’s all about dialogue. And it’s not coming at the problem necessarily as opposition. And that’s what I found. There have been some cases over the years where. Parents have come after me because I was trying to get their child to tell the whole truth.
Vicki Davis (05:53): Yeah.
Tim Plaehn (05:55): And they wanted to protect their child. But of course, in so doing, they’re not protecting their child in a way that’s going to help them grow and become a responsible, critical thinking adult. I think that challenge is becoming increasingly more pronounced for us as teachers.
Vicki Davis (06:11): Because to have redemption, my goal of any behavioral issue is that we’ll learn from it, we’ll grow from it, we’ll redeem the mistake. But if you can’t admit there’s a mistake, if you can’t admit that you did wrong, you can’t even start that process because then you’re dealing with lying.
Tim Plaehn (06:30): Right. And this is, at Asheville School, what we do is we separate conduct from honor. We found that when they’re together, things get messy, things get overly complicated. Let’s say a student is in another dorm room after lights. That’s a conduct issue. You’ll have detention. But if you cheat on a test, that’s an honor issue. And you’re going to sit down with the chair of the honor council, both the student chair and the faculty chair, and just kind of talk through your thought process, why you did this, what you could have done differently. And then to your point, what is the restorative justice step we’re going to take? You know, is it a quick note to the teacher? Is it, sometimes I would have my offenders write a letter to their parents that we would not send. And I’d tell them, we’re not going to send this to your parents, but what would it feel like for you to have to tell your parents that you did this? How would that impact you? What would you say to them? Just so they can kind of process their misstep emotionally and kind of think through those consequences down the road.
Vicki Davis (07:37): So when you talk about honor, the pledge that you were sworn to uphold ends with, I will report any violations. So you’re asking 15-year-olds to turn in a friend. How can you enforce this without it becoming like a snitch society? I like how you separate conduct from honor, but how do you manage this and encourage to turn in violations and not have that snitch environment?
Tim Plaehn (08:01): It’s a super difficult job. Like the fourth point of our honor code is the one that gives our students the most pause. And without us being able to chase it down, it’s certainly the piece of our honor code that’s most often broken. Students know what other students are up to. Well, I try to frame it, though, is this idea of snitching or tattling that comes from a culture of, that wants to protect those making mistakes against an institution or against a community. And if you want to live in a certain kind of community, you have to face up to certain responsibilities. On our campus, for instance, we don’t put anything away. You’ll see computers and phones and backpacks just strewn about outside rooms, you know, because the expectation in our community is that nothing will be stolen and that we trust each other. Well, part of the creation of that community is that we’ve got some skin in the game, that we all want to see this community in such a light. And so you have to then take responsibility. What I tell my students, because it’s a tough one, it really is. You know, would you tell on your neighbor if he were violating the water ban in your community and, you know, spraying his lawn at night? And they’re like, oh, no, we’d probably let that go. I’m like, okay, what if he was running an international children’s sex ring? Like, what? I’m like, yeah, you’re going to step up. You’re going to do something. That’s a hard step for a teenager to make and figure out when to do that. But to their credit, we have some students who are like, you know, I’m not comfortable telling you this, but I saw Steve cheating on this test. Ultimately, it’s for Steve’s benefit that that student is coming forward to tell us. Ultimately, Steve is going to learn that cheating is only going to get him so far. But I agree with you. It’s a challenge. I like how West Point does it in their honor code. It says, I will not lie, cheat, or steal, nor will I tolerate any violation of the honor code. And that idea that a West Point grad is going to hold his fellow or her fellow cadets up to a certain standard because they’re trying to make our military, you know, operate at a certain level. So you’re not going to tolerate someone who’s taking shortcuts because what’s going to happen down the road when you really need that person to step up?
Vicki Davis (10:36): Honor is important for a reason. None of us want to waste our time. Teachers are optimists. We want the future to be better than the past. And we want to pour into these kids. And when we feel like, what just happened? Did I just pretended to give you an assignment? You just pretended to do it? Then some teachers AI grade it. So then I just pretended to grade it.
Tim Plaehn (10:55): What are we doing?
Vicki Davis (10:56): We’re better than this, right?
Tim Plaehn (10:57): There was an assignment I gave and this was an eight week research paper. They got to choose their topics and dive in. And, you know, we had weeks and weeks of research, note-taking, outlining, drafting. And finally the due date came and it was a Sunday night and I was grading these papers and I came across two that were clearly AI. By the end of the year, I know they’re writing very well. As a teacher, the betrayal you feel. I can’t believe they would do this to me. But I said, you know what? I want to give them a chance. To restore their integrity. So I emailed all 44 of my students and I said, hey, a number of you use ChatGPT on these papers. I want to give you a chance to restore your integrity because you’re headed off to college and it’s going to be all you. So all I want you to do is if you use ChatGPT, just email me back and say, my bad. Then you and I will tackle what comes next. We’ll get through this together. The number of dings that began coming through on my computer as I went away, I had found two out of 45 and it was just ding, ding, ding. By the end of the night, 22 of my students had come forward and said, I did use ChatGPT. It was an eight page paper. It was challenging. They started to run out of things to say on page six and they took that unfortunate step. I love the fact that they came forward. I love the fact that they admitted it. And we got to have a really good talk about what are they going to do in college? Like, How are you going to face this in college with assignment after assignment after assignment?
Vicki Davis (12:25): How did you feel as 22 dings after you sent that email and you’re like, what? Like that shakes you.
Tim Plaehn (12:34): It really did. It really did. It shook them. Which I appreciated. You know, when we unpacked it over the next few days, like they were a little like, golly, all of us? You know, they might have known a few here and there, but not even my students were expecting 22 out of 45. And apparently I was told by another student that one of his friends was too scared to come forward. He was going to ride it out. So apparently it was 23 out of 45. That was a tough, tough April for me.
Vicki Davis (13:08): I appreciate you speaking the truth because I’ll tell you, as I’ve dealt with it in certain things, I’ve had the talk and had one or two and found out it was over half the class. This is not just a Tim or a Vicki thing. This is a everything, everywhere.
Tim Plaehn (13:25): Where it seems so many of my friends at a number of schools, so many of them are focused on AI and AI instructions and AI discussions without the attendant discussions of integrity. What exactly is it we’re after? Companies pushing schools to really invest in AI instruction because this is the future and kids need to know AI. I have no doubt that they’re right, but I think what our kids really need to learn at this stage is how to think critically, how to be honest, how to be persevering. You know, I think those are the skills that are going to provide for them for the rest of their lives.
Vicki Davis (14:12): The current research from Common Sense Media says 84% of kids are now using AI.
Tim Plaehn (14:16): I think that’s one of the biggest challenges for our teenagers is authenticity. Like, Who you are is not your latest Instagram post or who you are is not the influencer you’re following. And so I think ideas like that, ideas like integrity, authenticity, being honest, those are some skills that we’ve really got to dig into. But I think as institutions, we have to figure out a way to bring this idea into every single classroom and kind of speak with one voice about its importance.
Vicki Davis (14:56): So we’ve been talking with Tim Plaehn, American Studies teacher at Asheville School, and he has just published The Honor Code: Students, Integrity, and Our Path Forward. Tim, thanks for coming on the show and thanks for telling your story. Experience AI, co-founded by the Raspberry Pi Foundation and Google DeepMind, sponsored today’s show. Experience AI is a free AI literacy program downloaded more than a million times worldwide. It provides teachers with ready-to-teach lessons with slides, lesson plans, worksheets, and an excellent free AI glossary that I highly recommend downloading first. Within the resources, there are unplugged activities for you to explore with your learners that don’t require a computer at all. Experience AI supports all teachers, regardless of subject area, and doesn’t require any computer science or background knowledge. Get Experience AI free at coolcatteacher.com forward slash experience AI. That’s coolcatteacher.com forward slash experience AI. I’m recommending Experience AI because these lessons help you teach AI literacy with confidence and teach students how these systems actually work. Thank you, Experience AI, for sponsoring today’s show.
Announcer (16:17): Thank you for tuning in to 10 Minute Teacher Podcast. Join us here every weekday and subscribe to the Classroom Matters newsletter. See you later, educator.
Disclosure of Material Connection: This is a sponsored episode and blog post. Experience AI has compensated me to share information about the Experience AI program. However, all opinions expressed are my own. I have personally reviewed these resources and only recommend tools I believe offer genuine value to classroom teachers. My endorsement is limited to the educational products and services discussed in this episode. I am disclosing this in accordance with the Federal Trade Commission’s 16 CFR, Part 255: “Guides Concerning the Use of Endorsements and Testimonials in Advertising.” The sponsor has no impact on the editorial content of this show.
Disclosure of Material Connection: This episode includes some affiliate links. This means that if you choose to buy I will be paid a commission on the affiliate program. However, this is at no additional cost to you. Regardless, I only recommend products or services I believe will be good for my readers and are from companies I can recommend. I am disclosing this in accordance with the Federal Trade Commission’s 16 CFR, Part 255: “Guides Concerning the Use of Endorsements and Testimonials in Advertising.” This company has no impact on the editorial content of the show.
Subscribe to the 10 Minute Teacher Podcast and Cool Cat Teacher Talk anywhere you listen to podcasts.
Generative answers are everywhere. People like them because they are simple and they just “give the answer.” But what if the answer is nuanced. And what if, in the case of the generative tool, it is so eager to give an answer that it makes mistakes that any researcher would understand. Let’s go through an example.
So, first, Generative Adversarial Networks were used from roughly 2017 to 2022 to make deepfake videos. So, one model generates the video, another model evaluates it, and they go back and forth until the video quality becomes so good that it could become undetectable from the real thing. At first I called what I was doing a GAN, but it turns out a GAN is a training architecture, not what I’m doing. The right word is ensembling. Some may call it a “multi-agent debate” or “cross-model verification.” How did I find this out? I am training my AI tools to teach me more about AI and check my word choices and teach me about the models that underlie what I’m doing. It is a helpful method for learning in a domain of knowledge, to be up front with the tools you use about the pursuit of knowledge in that domain. It creates a mini spotlight that the AI tool will use to funnel me more of what I want to learn about.
As I dug into ensemble methods, I found the ROVER Method is one that I’m really using for these transcripts that are more accurate. So, I use Riverside which generates a transcript. Then, I put a video into Adobe Premiere Pro which uses another model to generate a transcript. Finally, I use Auphonic which then uses Whisper to generate a transcript. Then, Claude Cowork has a skill that I built which takes the bio of the guest and the topic to ensure that all acronyms are appropriate for that domain of knowledge and defined in the transcript and I also have fact checking to double check anything said built in. And the transcripts are compared until a final transcript emerges. I’ve found that each model is better at different things. Adobe for timings and finding subtle noises and words. Riverside for getting the names accurate. And Whisper for double-checking hard-to-hear phrases and such. And between them as I run them “against” one another, and have Claude ask me about discrepancies, I can get transcripts that are more accurate, less expensive, and less time consuming than what I’ve ever done before. It is a technique. By running the models against each other, I’m getting better results.
Vocabulary for This Post
Six words you need to read the numbers in this post
Hallucination and confabulation are already on my AI Vocabulary List. The other four are new — I’m adding them today, along with three from Monday’s episode. Want to find out how many of these words you already know? Play Spy the AI, the free vocabulary game I vibe coded from that same list.
Hallucination (also called confabulation)
When an AI states something false as though it were fact. Some researchers prefer confabulation, because the model isn’t seeing something that isn’t there — it’s filling a gap with something that sounds plausible.
Hallucination rate
The share of a model’s answers judged false on a given test. The catch that drives this whole post: every benchmark defines the bottom of that fraction differently. Two numbers both called a “hallucination rate” are often not measuring the same thing at all.
Benchmark
A standardized test used to score AI models. Results are only comparable within the same benchmark, and only as of the date the test was run. Leaderboards change constantly — the announcement post about a benchmark is almost never its current scores.
Generative Adversarial Network (GAN)
A training architecture in which two neural networks are trained together: a generator makes candidates and a discriminator judges them, and the generator improves from that feedback. This is how many early deepfakes were made. Important distinction: comparing several already-trained models to each other is not a GAN — no training is happening. That’s ensembling.
Ensembling (model ensembling)
Running more than one model on the same task and combining the results, because different models are good at different things. For speech-to-text this has a documented name — ROVER, from NIST in 1997. For chatbots checking each other’s answers, researchers call it multi-agent debate.
Adversarial testing
Deliberately feeding a system input designed to make it fail, to find out where it breaks. In the clinical study below, researchers planted one fabricated medical detail in every case. An adversarial score answers “how easily can this be tricked?” — not “how often is it wrong in normal use?” Two very different questions.
Also just added from Monday’s episode:De-identification — removing the details that connect data to a real person before it goes anywhere near an AI tool. API — a doorway that lets one program hand data to another. Interview Prompting — asking the AI to interview you instead of trying to write one perfect prompt. All three come from A.J. Juliani’s data dashboard episode.
We know about AI “hallucination” or as many prefer to say “confabulation” where AI just makes stuff up. But now, AI can cite things so it is supposed to be better, right?
Well, I went through an example that I’ll be using with students because it really shows the nuances of AI and research studies. Simpler is not always better when it means we think we understand and state error as fact.
Now, do not stop at Step 2, or even at Step 3. I need you to follow this chain of reasoning here so we can answer the question: how accurate are “AI Overviews” and is the question even the right question to ask?
Step 1: Google Search
So, first, I’m working to find hallucination rates for AI currently, so I did a simple Google search.
This is a number I update quite frequently, but I was curious as to the accuracy of these numbers in the generative search box. The first thing that bothers me is that whenever I see numbers presented without citations, an alarm bell goes off. See the words “high rates” and “legal research” and “low rates” – perhaps the citation on the second bullet is there, but sometimes it isn’t.
STEP 2: Claude Fact Checking Skill
So, I took a screenshot of the Google search and went to Claude. Now, granted, if I had pasted in the research links, more accurate information would have happened arguably at this step. But I want to demonstrate how we’re fact checking at a conference or event, that we might take a screenshot, so for now, this test is using screenshots.
So, I went into Claude and pasted the screenshot and asked it to “fact-check these numbers from a Google search.” Then, after it came back with errors, I asked it to update the Google graphic with information on what it found. On the right are Claude’s verdicts, and on the left is the original search. But wait, we need to get the ensemble activated here. We’re not done yet. Gemini may not be so bad, and Claude might not be so good. (Again, I didn’t give links, or item 2 would have been a different answer.) I use ChatGPT Pro and Perplexity Pro as part of this process.
Version 1 of fact-checking from Gemini to Claude. Do not cite this one. It is full of errors, as you’ll see!
Note: I have programmed my AI tools to help me teach. Everything I create is in the context of teaching someone, even myself, so you can see the lesson for the student emerge organically from that memory file.
STEP 3: Fact Check with ChatGPT Pro set to “high”
So, now I took the graphic from Claude that is above and I pasted it into ChatGPT, again using the screenshots. Its conclusion, “There are errors on both sides.” So, now this third model is finding errors on both sides of the equation, both Gemini and Claude. Here is the summary it found.
ChatGPT, to summarize, found the following:
Claude got Vectara backwards and found the November 19, 2025 announcement and not the newer announcement.
ChatGPT’s wording is important here, “I would not call most of Google’s individual numbers hallucinations. The more serious issue is they answer different questions…those percentages cannot meaningfully be placed on one common ruler. Here is what ChatGPT states about these.
A Vectara 1.8% means roughly: When the model is handed the source document and told to summarize only that material, how often does the summary contain something unsupported?
The Stanford legal number means: When an older general-purpose LLM is asked precise questions about federal court cases without necessarily being handed authoritative source material, how often is its response inconsistent with the legal facts?
The Mount Sinai clinical number means: If researchers deliberately planted a nonexistent medical fact in a case, how often will the model fall for the trap and elaborate on it?
AA-Omniscience asks yet another question about whether models guess rather than admit they do not know. Its hallucination rate denominator is specifically incorrect / (incorrect + partial + not attempted).
Oh my, so you mean my fact-check tool can be wrong too? Now, we’re getting past hallucination, and we need to be careful about throwing around this word. We’re talking about accuracy here, and mismatched research outputs can make a big mess.
When you look at the cited results above, they do not go together. This is a problem with wanting a “simple answer” in an emerging field like AI lots of studies are being done but have different research questions and methodologies that do not mean they can go together.!
STEP 4: Pasting ChatGPT’s answer back into Claude and asking for a Response
Ok, this is where the apologies start. We all know the drill. We catch AI making mistakes and then it is sickly sorry for what it has done. The two apologies included:
Vectara Leaderboard, it pulled November 2025 instead of May 11, 2026. This is a good catch and precisely how you see how multiple models can help things.
ChatGPT caught a logic error because it implied that o1 was not a reasoning model, but it was, so that comparison can’t demonstrate that reasoning models hallucinate more, only that the newer one scored worse than an older one.
It said it had dismissed suprmind.ai as an “SEO content hub” without looking at the source of the numbers on that page. So, it looked at the source that held the numbers and didn’t realize that it also had a source, revealing a flaw. If something is searchable and findable and holds a number, sometimes AI only goes to that page instead of tracking back to the original sources.
So, then it said that ChatGPT was not right about Gemini 3 Flash at 92%, as it found that Gemini 3 Flash was at 88%, so they disagreed on that model so it is more accurate to leave the ceiling out.
At this point, I think most people would fatigue and say “what is right here” but I’m about to do a big old mammoth update to throw a whole bunch of data into my tool to help determine what is right but the conclusion – by the Ai models themselves – is going to be a powerful one if you can persist. Again, we’re running models against models. And in the end, students need to understand not only are there errors, but sometimes, those errors are there because of mistakes in looking at the wrong information and aren’t just “hallucinations.” That research is nuanced and that human eyeballs are valuable.
STEP 5: Multiple model fact checking
So, I just wanted to be done with this, so I took Claude’s results and put them into both ChatGPT and Perplexity. I specifically told Perplexity that I had used Claude and ChatGPT, and that I had put it in orchestrator mode, so it might need to use other models. I’ve linked the chats above for transparency and so you can see some of the exciting nuance that comes out of these.
Interesting tidbits:
Perplexity and Claude both used the older November 2025 article instead of the newer article. As Claude said, “Perplexity confidently endorsed my wrong verdict using the same bad method that produced it.”
ChatGPT was the tool that found that Llama 2 was the wrong example to use for Google’s ceiling.
Because of “disagreements” between models and the fact they didn’t report the ceiling of some numbers, it is better to leave it out.
Now, if you look at my Claude chat for this, you’ll see a “retrospective,” which is where Claude analyzes what it got wrong. There is a method to my madness here with this and I’ll get there in a moment. But the Perplexity chat says something interesting:
"The teaching point in your footer is the real lesson: the primary sources (live leaderboard, paper abstract, system card) beat any AI summary of them, including one AI's summary of another AI's summary."
So, the AI itself acknowledges that we have a big old mess without consulting the original sources. So, then I took information from both ChatGPT and Perplexity and here’s the current output of AI evaluating the Google AI Overview.
OK, so this is interesting now. I want a graphic evaluating Google, and Claude is interjecting evaluations of itself. This is “mission drift” in action, particularly when you are fact-checking. So, I’m having to ask for a final graphic evaluating the Google generative results on the hallucination rates of AI models.
STEP 6: Re-generate the original graphic for the original purpose of this task
Now, I want you to note a few issues that I do not like about the information above:
The citations are small and listed at the top. Again, we have a graphic and it is hard to fact check. I, thus, asked AI to generate a research box so I can read information and double check the conclusions.
I would really like all models used to be documented somewhere somehow. It is citing the human, for sure, me – Vicki Davis- but it isn’t putting “Created using Claude Cowork” or any other models that I used in the fact-checking process. I think model disclosure will be very helpful in the future, even as we humans are held accountable. Being able to be cognizant of the need to document chats in this way is important. When you see the chat, you’ll see what I mean.
Before I would produce this as “research,” I would need to sit down and read every single study, and I would argue that, with all of this back and forth, reading original source documents is more important than ever. We should be researching slower not faster when we see this happen.
I would really like to be able to generate a link to share this claude check but because I run Claude Cowork on my computer, I’ll have to generate a PDF instead which I will paste below.
When you look at this, you’ll notice how I use Claude to fact-check my writing. Now, you might think – Claude was wrong; why would you use it to fact-check? Well, particularly in AI terminology, I’m learning and need to keep learning and understand various models. I believe that workflow is more important than ever, as are AI techniques, and to teach this, I have to understand and use the proper vocabulary. I live in rural Georgia — to call it the sticks might be an insult to sticks. So I can’t really go to my local coffee shop and hang out with the other AI nerds. I have to watch them on YouTube and read their articles on LinkedIn, so I’ve programmed the AI to help me be more accurate and precise in my AI speech. This is an example of using AI for learning. Also, the ability to produce a PDF of an AI chat is a valuable part of documentation as we look at the process of research.
STEP 7: Generate the Research Used for This Output with Hyperlinks.
I like to generate research as an HTML box I can paste into WordPress and then I can click on the links and review them in a new browser.
Research Citations
Every figure in the graphic above traces to one of the primary sources below, verified September 1, 2026. No aggregator numbers, no launch announcements, and no AI summaries were used as evidence.
Sources used
Grounded summarization
Vectara Hallucination Leaderboard (live repository, updated May 11, 2026). github.com/vectara/hallucination-leaderboard — source of the 1.8%, 3.1%, and 3.3% figures. The 9.6% median across 105 models was computed directly from this table.
Awadallah, A. and Mendelevitch, O. “Introducing the Next Generation of Vectara’s Hallucination Leaderboard.” Vectara, November 19, 2025. Read the announcement — cited for methodology only.
Open-domain factual recall
AA-Omniscience evaluation page, Artificial Analysis (live scores and metric definition). artificialanalysis.ai/evaluations/omniscience
Jackson, D., Keating, W., Cameron, G. and Hill-Smith, M. “AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models.” arXiv:2511.13029, 2025. Read the paper
Reasoning models
OpenAI. “OpenAI o3 and o4-mini System Card,” April 2025, Table 4. Read the system card
Legal research
Dahl, M., Magesh, V., Suzgun, M. and Ho, D. E. “Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models.” Journal of Legal Analysis, 16(1), 2024, pp. 64–93. Read the paper · Stanford RegLab summary
Clinical decision support
“Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support.” Communications Medicine (Nature Portfolio), 2025. Icahn School of Medicine at Mount Sinai. Read the study · PubMed listing
Methods and terminology
Fiscus, J. G. “A Post-Processing System to Yield Reduced Word Error Rates: Recognizer Output Voting Error Reduction (ROVER).” Proceedings of the IEEE Workshop on Automatic Speech Recognition and Understanding, Santa Barbara, CA, 1997, pp. 347–354. NIST record — the documented technique for combining multiple speech-recognition outputs into one composite transcript more accurate than any single system.
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B. and Mordatch, I. “Improving Factuality and Reasoning in Language Models through Multiagent Debate.” arXiv:2305.14325, 2023; published at ICML 2024. Read the paper — multiple model instances debate an answer across rounds, improving factual validity.
Goodfellow, I. et al. “Generative Adversarial Networks.” arXiv:1406.2661, 2014. Read the paper — cited for contrast. A GAN trains a generator and a discriminator together in one loop, with the discriminator’s judgment updating the generator’s weights. It is a training architecture, not a method for comparing finished models.
Sources Rejected, and Why
This list is the more useful one for classroom use. Each source below showed up during the research and was deliberately left out. Every link in this section is marked “nofollow” on purpose — see my note below.
1. Vectara’s launch blog post, quoted as current data. It is a snapshot from November 19, 2025, not the live leaderboard. Quoting its top scores as today’s produced the false claim that no model scores under 2%. The live repository shows 1.8%. The rule: a launch announcement records the day a benchmark was newest. For current numbers, open the artifact the announcement points to.See the announcement
2. AA-Omniscience’s launch article, quoted as current data. The same error on a second benchmark. Its “lowest at 26%” was quoted as the current floor; the live page shows 1%. Two details worth showing students: the article states its own figure inconsistently (28% in one bullet, 26% in another), and it explicitly says “For up to date AA-Omniscience scores, see the AA-Omniscience evaluation page.” The correction was printed right there and still got missed. See the launch article
3. Aggregator and content-hub pages. No methodology, no version history, no date-stamp showing which snapshot a number came from. The subtlety worth teaching: one such page’s numbers were actually correct. It was first dismissed for looking like a content farm rather than for any test of its figures — right conclusion, wrong reasoning. The real test is not whether a page looks credible but whether its numbers trace to a primary source you can open. See the page in question
4. “22% to 94% across 26 models,” attributed to the Stanford AI Index 2026. Could not be verified against the AI Index itself, and the descriptions that could be found identify it as sycophancy-induced hallucination — a different measurement. Two numbers both labeled “hallucination rate” are not necessarily measuring the same thing. See the AI Index report
5. “Gemini 3 Flash at 92%.” Sources disagreed and none was primary: reachable sources said 88%, two AI models later said 91%, the original claim was 92%. The live page publishes lowest scores but not highest, so the question stayed open. The graphic says “above 90%” instead. The rule: when sources conflict and no primary source settles it, report the range you can defend rather than the most quotable figure.See the Gemini 3 Flash analysis
6. Legal AI tool rates (Lexis+ AI 17%, Westlaw 33%). These come from a real follow-up study, but that paper was never opened during this check. Plausible is not the same as verified, so they were left out. See the follow-up study
7. A “1.47% real-world clinical hallucination rate.” Surfaced in passing, never traced to its source, so it was not used. No link — the source was never located, which is exactly why it was dropped.
8. The AI summary’s own citations. Worth pointing out to students: the visible citations under the search result pointed to content-marketing pages rather than to the Vectara leaderboard, the Stanford paper, the OpenAI system card, or the Mount Sinai study — the actual origins of every number it quoted. See one of the cited pages
A caution for anyone reusing this graphic: benchmarks update continuously. These figures were current on September 1, 2026. Check the live pages before quoting them later.
Notes about the research citations above from Vicki: So, do you see what I did there, it said “we” but then when it cited it, it cited “me” – Vicki. I don’t like this subtle use of pronouns. AI is a tool and I’m accountable. However, if I look at the chat, I am using the word we as well — I have to think on that. wow. Look at that.
Additionally, I would like the links to the articles that were rejected. Update: When I looked at adding those links, Claude pointed out something I had forgotten, that the presence of a true link passes SEO on for credibility to those sites, something I’m not really wanting to do so it will mark them as “no follow” links. This is an interesting aspect of citing rejected articles.Also, note that when I do this again, I’m going to create a version of this skill that stops and lets me make the decision as this chat was in “auto mode” for speed. I really had no idea this would be such a hard question to get right in Google generative search, but I’m glad I did this activity and it is one I will be doing with students.
STEP 8: Updating the Fact Checking Skill in Claude Cowork
Now, I’m coming back to the fact-checking skill I’ve built in Claude so you can see why I keep coming back. When I’m done, I ask it to analyze every mistake and then update the fact-checking skill. This is a whole other process because now, I’m working to teach the AI tool how I like to operate and to learn from interacting with other AI tools about the flaws and mistakes in that tool.
This is why skills will become valuable intellectual property for companies (if indeed they can be owned by the company and not harvested by the AI models themselves, which is a whole other topic.)
RETROSPECTIVE
In every chat, I ask Claude Cowork, ChatGPT, or Gemini to conduct a retrospective analysis of what worked, what mistakes were made, and how to prevent those mistakes in the future. It is a learning model. Our purpose in interacting with AI is both to get a job done and to make doing that job better in the future.
Additionally, by doing this sort of orchestration, we can build distrust for the overly simplified generative answers we’re getting from Google right now – or really any tool, for that matter. No tool is always right. Different domains of knowledge have different accuracy rates, and different ways of using AI can yield higher accuracy. Sometimes we need an ensemble in order to do complex tasks.
But here is a big takeaway: When in doubt, go to the original documents to check it out.
When we start using multiple AI tools to check for answers, we see that each tool is flawed and that we must use discernment to understand the nuances. Plus, rushing to get something out can lead to mistakes.
So, I’ve just written this as I’ve gone through the process, when I realized this wasn’t going to be an easy check. I was actually preparing a presentation about research using AI and just wanted to expose flaws, and the rabbit hole went much deeper than I thought!