Blog

AI Detectors: How They Work and Why They Get It Wrong

AI Detectors: How They Work and Why They Get It Wrong

An AI detector is a tool that estimates how likely it is that a piece of text was written by a language model rather than a person. It does this by measuring statistical patterns, mainly how predictable each word is, and returning a probability. That number is a guess about style, not proof of authorship, and independent research shows those guesses go wrong often enough to matter.

If you have ever pasted your own essay into a checker and watched it come back "likely AI", you already know the problem. This guide explains what happens inside a detector, what the accuracy studies found, why certain writers get flagged more than others, and what you can do to protect yourself.

What an AI Detector Actually Measures

A detector never sees who typed the words. All it has is the text, so every method works backwards from patterns in the writing itself.

Language models produce text by repeatedly choosing a likely next word. The result tends to be smooth and statistically safe. Human writing is usually messier: an unexpected word here, a long sentence followed by a short one there. Detectors try to measure that difference.

Perplexity and Burstiness

Perplexity describes how surprised a language model is by a piece of text. Low perplexity means each word was easy to predict, which is what a model's own output looks like. High perplexity means the text took turns the model did not expect.

Burstiness describes how much that predictability varies across sentences. People tend to write in bursts, mixing plain sentences with unusual ones. Model output is often more even. A detector that sees low perplexity and low burstiness leans toward "AI".

The weakness is obvious once you say it out loud. Plenty of human writing is predictable on purpose: lab reports, legal summaries, five-paragraph essays and anything written in a second language with a careful, limited vocabulary.

Trained Classifiers

Most commercial tools add a second layer, a classifier trained on large sets of human and machine text. It learns whatever features separate the two sets it was shown. That makes it better at the models it was trained on and less certain about new models, new topics, short passages and edited text. Turnitin, for example, has said only that its tool looks for patterns common in AI writing, without defining them.

Watermarking: A Different Approach

Watermarking tries to solve the problem at the source. Instead of guessing after the fact, the company that runs the model hides a statistical signal in the text while it is being generated.

Google's SynthID Text works this way. It adjusts the model's word probabilities during generation using a pseudorandom function tied to a private key, so the pattern is invisible to a reader but measurable by a detector that holds the key. Detection is still probabilistic, and the detector returns one of three states: watermarked, not watermarked, or uncertain.

Google lists the limits plainly. The watermark survives cropping, a few changed words and mild paraphrasing, but confidence drops sharply when the text is thoroughly rewritten or translated into another language. It also works less well on factual answers, where the model has little room to vary its wording. And it can only confirm text from a model that applied the watermark, so it says nothing about text from any other model.

How Accurate Are AI Detectors?

This is the question that matters most, and the evidence is less reassuring than the confident percentages on most checker screens.

OpenAI Withdrew Its Own Detector

OpenAI launched an AI text classifier in January 2023. By its own evaluation, it correctly identified 26% of AI-written text as likely AI, while labelling human text as AI 9% of the time. OpenAI also warned that it was unreliable on text under 1,000 characters. On July 20, 2023, the company withdrew the classifier, citing its low rate of accuracy.

If the company that built the model could not reliably detect its output, that tells you something about how hard the problem is.

An Independent Test of 14 Tools

A research team led by Debora Weber-Wulff tested 12 publicly available detectors and two commercial systems widely used in universities, Turnitin and PlagiarismCheck. Their published evaluation concluded that the tools were neither accurate nor reliable. The tools leaned toward calling text human, so a lot of AI text slipped through, and obfuscation techniques such as paraphrasing made their performance significantly worse.

Newer Tools on Newer Models

Detectors have improved on some benchmarks. A 2025 study that tested tools on text from DeepSeek found that QuillBot, Copyleaks and GPTZero each scored above 92% accuracy on paraphrased samples, while a detector built on the older GPT-2 model scored under 4%. That study used 49 samples, so treat it as a snapshot rather than a verdict. The same researchers found that running the text through "humanizing" rewrites lowered the accuracy of GPTZero, Copyleaks and QuillBot, and that GPTZero produced some false positives on human writing.

Accuracy depends heavily on which model wrote the text, how it was edited, and who the human comparison writers are. A single headline percentage hides all three.

Why AI Detectors Flag Human Writing

The most cited study on false positives comes from Stanford researchers led by Weixin Liang. They ran seven popular detectors on 91 essays written by non-native English speakers for the TOEFL exam and on 88 essays written by US eighth graders.

The results, published in the journal Patterns, were stark. The detectors were close to perfect on the US essays. On the TOEFL essays, the average false positive rate was 61.3%. All seven detectors agreed on labelling 19.8% of those human essays as AI, and at least one detector flagged 97.8% of them.

The researchers traced the cause to perplexity. Writers working in a second language often use a narrower, safer vocabulary, which makes their text more predictable. When the team used ChatGPT to enrich the word choice in the TOEFL essays, the false positive rate fell to 11.6%. When they simplified the vocabulary in the US essays, the rate rose from about 5% to about 57%.

That has two uncomfortable implications. Careful, plain writing can look machine-made. And running your text through an AI tool to sound more sophisticated can make a detector more likely to call it human.

Turnitin's AI Detector and University Pushback

Turnitin's AI detector is the one most students meet, because it runs inside the same system that checks for plagiarism. Turnitin has claimed a false positive rate of under 1% at the document level.

Several universities decided that was not good enough. In August 2023, Vanderbilt University disabled the feature, pointing to reported false accusations at other schools, the bias against non-native speakers, and the fact that Turnitin gave no detailed information on how it decides. Michigan State, Northwestern and the University of Texas at Austin also switched it off. The concern at large universities was simple arithmetic: even a 1% error rate across many thousands of submissions means a lot of students wrongly flagged.

Those universities pointed instructors toward clearer AI policies, assignments that are harder to outsource, and teaching responsible AI use, rather than relying on a score.

When an AI Detector Is Actually Useful

None of this makes a detector worthless. It makes it a signal that needs context, used in the right situations:

  • Checking your own draft before you submit. If a section reads as "likely AI", it is often the most generic part of the essay, and rewriting it in your own voice usually makes it better anyway.
  • Spotting copied filler. Editors and teachers can use a high score as a reason to look closer at a passage, then judge it on its content.
  • Comparing versions. Running an early draft and a final draft shows whether editing moved the text toward or away from a generic register.

If you want to try one without signing up, Primo Notes has a free AI detector on its website that you can run on a draft in seconds. Treat its score the way the research suggests you treat any score: as a prompt to reread, not a ruling.

Whether you use GPTZero, Copyleaks, Grammarly's AI detector or Turnitin's, the output is the same kind of thing, a probability built from patterns. A high number on a short passage, on formulaic writing, or on writing in a second language deserves extra doubt.

What to Do If an AI Detector Flags Your Work

A flag is the start of a conversation, and you want evidence ready before it begins. The strongest evidence is a record of how the work came together.

Keep a Record of Your Process

Keep your drafts and their version history. Google Docs and Word both store revision history, and a document that grew over several sessions, with false starts and deleted paragraphs, looks nothing like text pasted in one go. Keep your research notes, outlines and sources too.

Your thinking process is also evidence. If you talk through your argument in a voice note before writing, or record the lecture your essay draws on, you have a dated trail of your own ideas in your own words. Primo Notes turns a recording like that into a written note with a transcript, so the reasoning behind your essay sits next to the essay itself.

Be Ready to Explain the Work

Most academic integrity processes give you a chance to discuss your writing, and a student who can walk through the argument, the sources and the choices is hard to dismiss. Practising that explanation out loud helps, which is the idea behind the Feynman mode in Primo Notes: you explain the material and the app works through the gaps with you.

If you write in a second language, say so, and point to the Stanford findings. If you use AI tools for grammar or brainstorming, check what your course allows and be open about it. Guidance on using an AI writing assistant responsibly is a good place to start drawing that line.

Conclusion

An AI detector measures how predictable text is and how closely it matches patterns from model output, then turns that into a probability. It never knows who wrote the words. OpenAI withdrew its own classifier for low accuracy, an independent test of 14 tools found them neither accurate nor reliable, and Stanford researchers showed that non-native English writers are flagged far more often than native ones. Watermarking is more principled, though it only works for models that apply it and weakens when text is rewritten or translated.

Use a detector as a reason to reread a draft, never as a verdict on its own. Keep your drafts, notes and version history, and be ready to explain your work in your own words, because that evidence carries more weight than any score.