Can AI Detectors Tell If Students Used ChatGPT? What Teachers Can Do Instead

By GradingPal Team
Published
Updated
14 min read

Can AI detectors really tell if a student used ChatGPT? What the research says on accuracy and false flags, and what teachers can do instead.

A teacher runs a stack of essays through an AI plagiarism checker. One comes back "87% likely AI." It belongs to a quiet student who has never caused a problem.

What now?

This is a real situation in a lot of schools. In a 2025 survey of 806 middle and high school teachers by the Center for Democracy and Technology, 27% said they use AI detection tools regularly, and another 21% said they had tried them. 17% of teachers said students had been accused of using AI improperly when it was never proven.

So plenty of teachers are using these tools. The obvious question is whether the tools deserve that trust. The research, including several studies from the past year, gives a fairly clear answer.

Can AI Detectors Tell If Students Used ChatGPT? What Teachers Can Do Instead

How do AI detectors work?

It helps to know what's going on inside the box, because it explains most of the problems.

A normal plagiarism checker compares a student's essay to things that already exist, like websites, books and other students' papers. If a sentence matches, it shows you the source. You can check it yourself.

An AI detector can't do that. When ChatGPT writes something, it writes new sentences. There's nothing to match against. So when you check for AI plagiarism, the detector does something different: it looks at the style of the writing and guesses.

Most detectors look at two things:

  • How predictable the words are. Tools like ChatGPT tend to pick the most likely next word. So AI writing is usually smooth and predictable. Researchers call this low "perplexity".
  • How much the sentences vary. People tend to mix short and long sentences. AI text is often more even.

If an essay is very predictable and very even, the detector says "probably AI."

You can probably already see the problem. Lots of real students write predictable, even sentences. A student learning English uses the words they know best. A seventh grader follows the paragraph frame you taught them. A careful writer sticks to simple, clear sentences. None of them used AI, but to a detector, they can look the same.

And the result is always a probability, like "87% likely AI." It sounds precise. It isn't a measurement of what happened. It's the tool's best guess about style.

What does the research say about AI detector accuracy?

Researchers have been testing AI detectors since ChatGPT came out. The early studies were harsh. The newer ones, from 2025 and 2026, show the best tools have improved in one important way, but the main problem hasn't gone away. Here's both.

The early studies (2023–2024)

  • OpenAI shut down its own detector. In 2023, the company that makes ChatGPT released a tool to spot AI writing. In its own testing, it caught only 26% of AI-written text and wrongly flagged 9% of human writing. On July 20, 2023, OpenAI took it down "due to its low rate of accuracy."
  • Fourteen tools, none reliable. A 2023 study led by Debora Weber-Wulff, published in the International Journal for Educational Integrity, tested 14 detectors, including Turnitin. All scored below 80% accuracy. When the AI text was run through a paraphrasing tool, accuracy fell to about 26%.
  • Easy to fool. In a 2024 study, Mike Perkins and colleagues found six popular detectors were accurate only 39.5% of the time on plain AI text, and 17.4% after the text was changed with simple techniques designed to fool them.
  • Unfair to English learners. Stanford researchers led by Weixin Liang and James Zou ran 91 real essays by non-native English speakers through seven detectors. On average, the detectors labelled 61% of them as AI-written, and 97.8% were flagged by at least one detector. Essays by U.S. eighth graders were judged correctly more than 90% of the time. When the researchers used ChatGPT to make the vocabulary in the same human essays fancier, the false flags dropped to under 12%. The tools were reacting to simple wording, not to who wrote it.

The newer studies (2025–2026)

  • The best paid tools rarely flag fluent human writing now. In June 2025, Jisc, the UK's national technology body for colleges and universities, reviewed the evidence. It found mainstream paid detectors catch unedited AI text "reasonably well" and wrongly flag human writing about 0–2% of the time. But accuracy drops sharply once AI text is reworded, and, in Jisc's words, "an entire industry now exists to help users circumvent these tools."
  • They still miss a lot, especially mixed writing. A 2026 study by Marijke Van Vlasselaer and colleagues, in the International Journal for Educational Integrity, tested four tools on 160 papers: fully human, fully AI, mixed, and AI text rewritten to sound human. The good news: three tools wrongly flagged none of the 40 human papers, and GPTZero flagged about 5%. The bad news: only one tool, Pangram, reliably caught the fully AI papers. Turnitin, GPTZero and Copyleaks caught between 0% and 30% of them under the study's strict measure.
  • They struggle with science writing and mixed texts. Another 2026 study in the same journal, by M. Hadra, K. Cambridge and M. Mesbah, tested Turnitin and Originality on 192 texts, including real writing by students learning English. Overall accuracy was 61% for Turnitin and 69% for Originality. Both almost completely missed texts that mixed human and AI writing. And both did much worse on science than humanities: Turnitin dropped from 86% to 51%.
  • Even a small error rate adds up. A May 2026 paper by Panagiotis Tsigaris and Jaime Teixeira da Silva, in Next Research, looked at the maths. When most writers aren't using AI, even a detector that's rarely wrong can produce more false accusations than correct ones. Their conclusion: detector findings "are prone to generate more false accusations than correct identifications."

Here's a simple way to see that last point. Imagine a class of 100 essays where 10 were written with AI. Say a detector catches 30% of the AI essays and wrongly flags 2% of the human ones, numbers in the range the studies above found. It flags 3 AI essays, and about 2 honest ones. So 2 of every 5 flagged students did nothing wrong, while 7 of the 10 who used AI aren't flagged at all.


What the experts and companies say now

  • Jisc (2025): "Decisions should never be based solely on AI detection, whether by a tool or a person."
  • The 2026 Van Vlasselaer study, whose best tool performed well: detectors "should not be used as sole evidence in high-stakes decision-making."
  • Turnitin itself, quoted by NPR in December 2025: its AI detection "may not always be accurate … so it should not be used as the sole basis for adverse actions against a student."
  • Mike Perkins, speaking to NPR in 2025: "It's now fairly well established in the academic integrity field that these tools are not fit for purpose."

Some schools have reached the same conclusion. In 2023, Vanderbilt University turned off Turnitin's AI detector. Even at a 1% error rate, it worked out, its 75,000 papers a year could mean about 750 students wrongly flagged.

Here's all of it in one place, newest first:

Can AI Detectors Tell If Students Used ChatGPT?

So what's changed, and what hasn't? The best paid detectors have got much better at not flagging fluent, human writing. That matters, and it's fair to say so. But they still miss a lot of AI writing, especially when it's mixed with the student's own words or reworded, which is exactly how most students who use AI actually use it. Results vary by tool and by subject. Free tools are often much worse. And whether newer tools still treat writing by English learners unfairly hasn't been settled by the recent research, so the Stanford findings remain a real warning.

Two kinds of mistakes, both bad

An AI detector can be wrong in two ways.

  • A false positive is when it flags a student's own work as AI. The student did everything right and is now under suspicion.
  • A false negative is when it misses AI writing. The student who used AI and reworded it a little gets through.

The newest tools have cut down the first kind of mistake on fluent writing. They still make plenty of the second. That's an awkward combination for a teacher: a flag might be right, but a clean result tells you very little, and the students most likely to get past the detector are the ones who know to reword.

Can AI Detectors Tell If Students Used ChatGPT?

Why a flag isn't proof

Even OpenAI says so. In its guidance for educators, published in 2023, the company was asked whether AI detectors work. Its answer: "In short, no." It added that its own detector seemed to unfairly affect "students who had learned or were learning English as a second language and students whose writing was particularly formulaic or concise."

That same guidance busts one popular shortcut. Some teachers paste an essay into ChatGPT and ask, "Did you write this?" OpenAI says ChatGPT has no way of knowing, and "will sometimes make up responses" to exactly that question. So that doesn't work either.

There's a bigger point here, too. When a student copies from a website, you can show them the website. With AI writing, there's no source to point to. A detector score is the only "evidence", and it's a guess. Accusing a student of cheating on the strength of a guess is a serious thing. It can damage their record and their trust in you, and the research says the students most at risk are English learners and students who write in a plain, careful style. CDT's 2023–24 survey found the same pattern in schools: English learners and students with disabilities were more likely to face discipline over AI concerns.

This is also why "AI plagiarism" is a slightly misleading name. AI and plagiarism overlap, but they aren't the same thing. Plagiarism means passing off someone else's existing work as your own. A student using ChatGPT is a different problem, and it needs a different fix.

What to do if you think a student used AI

Sometimes you'll still have a strong feeling a piece of work isn't the student's own. That's fine. You know your students. Here's a fair way to handle it.

  • Don't act on a detector score alone. If you use one, treat it as a reason to look more closely, never as the answer.
  • Look at the work itself. Does it sound like the student's in-class writing? Does it use ideas or facts you never covered? Does it answer the actual question, or something close to it? Does it cite sources that don't exist? Vanderbilt suggested the same thing: compare it to the student's earlier work and look for gaps or errors.
  • Talk to the student, privately and calmly. Go in curious, not accusing. Ask them to explain their main point in their own words. Ask what a word they used means, or why they chose a particular example. Ask how they went about writing it. A student who wrote it can usually talk about it.
Can AI Detectors Tell If Students Used ChatGPT
  • Look at how the work was built, if you can. If they wrote in Google Docs or Microsoft Word online, the version history shows the document growing over time, or appearing all at once.
  • Give them a fair chance to show what they know. A short in-class piece on the same topic, or a five-minute conversation about it, tells you much more than any score.
  • Follow your school's policy, and write down what you did. If it does go further, you'll want a record of the conversation and the work, not just a percentage.

And if you turn out to be wrong, say so. A student who was wrongly suspected and then cleared fairly will remember how you handled it.

What teachers can do instead

If detectors can't be trusted, the answer isn't to give up on honest work. It's to set work where you can see the thinking, so the question "did a chatbot write this?" comes up much less. The researchers behind the Perkins study made the same point: rather than leaning on detection, schools should look at other ways of assessing.

None of these ideas are new. Good teachers have used them for years. They just matter more now.

1. Bring more of the writing into the classroom

The simplest fix is also the oldest. When students write in front of you, you know who wrote it.

  • Use short, timed writes at the start or end of a lesson: one paragraph answering one question.
  • For bigger pieces, do the first draft in class. Students can still revise at home, but you've seen where they started.
  • Paper still works. A handwritten paragraph from Tuesday's lesson is about as authentic as student work gets, and it gives you a baseline for what that student's writing actually sounds like.

2. Grade the steps, not just the final product

Break a big assignment into checkpoints: a topic and question, a plan or outline, a draft, the final piece. Give quick feedback at each step, and give some credit for the steps themselves.

Can AI Detectors Tell If Students Used ChatGPT

This does two jobs. Students get help earlier, when it's most useful. And you can see the work grow. A final essay that looks nothing like the outline from two weeks ago is worth a conversation.


In math and science, this already has a name: show your working. A correct answer with no steps tells you very little. The steps tell you whether the student understands.

3. Make students explain their thinking out loud

Talking about your work is much harder to fake than writing it.

  • Hold short conferences: two minutes per student, one or two questions about their work.
  • Ask students to record a short video explaining their answer, their method or their argument.
  • Add a quick oral check to a big project: "Walk me through your main evidence."

This also helps students who know more than they can write down, including many English learners.

4. Write prompts that only your class can answer

Generic prompts get generic answers, from students and from chatbots. Tie the work to things that happened in your room.

  • Instead of "Explain the causes of World War I," try "Which of the three causes we debated on Thursday do you think mattered most? Use at least one quote from the document packet."
  • Instead of "Write a lab report on photosynthesis," try "Use the data your group recorded on Monday. Explain why your result was different from the group next to you."
  • Instead of "Write a personal narrative," try one tied to a specific moment, place or choice from the student's own life, with sensory details only they would know.

5. Reward what AI is bad at

Build your rubric around the things that show real thinking: using specific evidence from class materials, explaining why, responding to feedback from the last draft, and connecting ideas to the student's own experience. A smooth, general essay that skips all of those won't score well, whoever wrote it.

6. Be clear about the rules, assignment by assignment

Students are often confused about what's allowed. Is it OK to use ChatGPT to brainstorm? To check grammar? To write a first draft? Tell them, for each assignment. A simple three-level system works well:

an AI Detectors Tell If Students Used ChatGPT?

Clear rules take away most of the grey areas. And if a rule is broken, you're having a conversation about something everyone understood, not about a guess.

7. Teach AI, don't just police it

Show students what these tools get wrong. Give the class a ChatGPT answer to a question you've studied together and ask them to find the mistakes, the vague parts and the missing evidence. Students who see how confidently AI can be wrong are less likely to hand their thinking over to it.

Where GradingPal fits

We'll be straight about this: GradingPal does not detect AI writing. We looked at the research above and decided we wouldn't build something that could wrongly accuse a student. If you type a grading instruction asking GradingPal to check whether work was written by AI, it tells you it won't, and why, right there as you type it. Our full position is on our responsible AI page.

What GradingPal can do is take the extra marking out of the approaches above, so they're realistic to use every week.

  • In-class writing on paper. Students write by hand, the way they always have. You scan the whole class set as one PDF. GradingPal splits it into one submission per student, reads the names off the pages, and asks you when it isn't sure who a page belongs to. Then it drafts scores and feedback against your rubric. No student devices needed.
  • Showing the working. For worksheets, problem sets, quizzes and exams, you can mark question by question and give partial credit for correct steps, so the method counts, not just the final answer.
  • Explaining out loud. Students can submit a short video or audio recording explaining their thinking. When you review it, comments are pinned to the exact moment in the recording.
  • Rewarding real thinking. GradingPal marks against your own rubric, one criterion at a time, and shows you the quote from the student's work behind each suggested score. You can add grading instructions in plain words, such as "Only give credit for evidence from the class document packet."
  • Feedback before scores. For checkpoint work, you can share feedback with the score hidden, so students read the comments and revise before they ever see a number.

Every suggested score is a first draft. Nothing reaches a student until you've reviewed it and shared it. And student work is never used to train AI models.


The bottom line

AI detectors promise a quick answer to a hard problem. They've improved, and the best ones now rarely flag fluent human writing. But the research, old and new, still says they can't deliver that answer. They miss AI writing that's been mixed in or reworded, they vary by subject, and a score can't tell you what actually happened.

The better path takes a bit more planning but much less suspicion. Bring more writing into the room. Grade the steps. Ask students to explain. Write prompts only your class can answer. And when something doesn't feel right, talk to the student before you decide anything.

If you'd like help marking the in-class writing, worked solutions and recorded explanations that make this possible, try GradingPal free.

Frequently asked questions

Are AI detectors accurate?

Better than they were, but not reliable enough to decide anything on their own. In 2026 tests, some paid tools rarely flagged human writing, but most missed a lot of AI text, especially mixed or reworded writing, and accuracy dropped on science writing. Researchers, Jisc and Turnitin itself say a score shouldn't be the sole basis for action.

Can Turnitin detect ChatGPT?

Turnitin offers an AI writing indicator. In a 2026 study it flagged none of 40 human-written papers, but caught few fully AI-written ones under the study's strict measure. Another 2026 study found its accuracy fell from 86% on humanities writing to 51% on science writing. Turnitin says its detection shouldn't be the sole basis for action against a student.

Is there a free AI plagiarism checker that works?

Free tools generally do worse. Most of the 14 tools in the 2023 Weber-Wulff study were free, and none were reliable. Jisc's 2025 review found many free or lesser-known detectors performed poorly, and one caught no AI text at all. Every detector, free or paid, is guessing from writing style.

Why would a student's own essay get flagged as AI?

Detectors flag writing that is predictable and even, with common words and similar sentence lengths. That describes a lot of honest student writing, especially from students learning English, younger students, and anyone following a set paragraph structure.

Do AI detectors flag ESL students more often?

Yes, according to research. A 2023 Stanford study found seven detectors flagged 61% of real essays by non-native English writers as AI on average, while getting U.S. eighth graders' essays right more than 90% of the time. Newer studies haven't yet shown whether today's improved tools have fixed this.

Can ChatGPT tell if it wrote something?

No. OpenAI says ChatGPT has no way of knowing whether text was AI-generated, and that it will sometimes make up an answer when asked "did you write this?"

How can I tell if a student used AI?

There's no tool that can tell you for sure. Compare the work to the student's in-class writing, look for ideas or sources that don't fit, check the document's version history if there is one, and talk with the student about their work.

How do I make assignments harder to complete with AI?

Do more writing in class, grade the steps along the way, ask students to explain their thinking out loud, tie prompts to things that happened in your classroom, and build your rubric around specific evidence and reasoning.

Does GradingPal detect AI writing?

No. Detection tools aren't reliable enough, and a false accusation is costly for a student. GradingPal helps you grade the kinds of work where AI is less of a question, such as handwritten in-class writing, worked solutions and recorded explanations.

Get the next guide

One email when we publish something new on grading, feedback and standards. No newsletter filler.

See our privacy policy.

Try It on Something You Have Already Marked

Reading about it only goes so far. Bring a class set you have graded and compare.

Free plan, no credit card. 50 submissions a month, any assignment type.