As AI writing tools become routine in and around newsrooms, editors increasingly face a practical question: how do you tell if a piece of text was written by ChatGPT or a similar large language model? The honest answer is that there is no single reliable method, and anyone claiming otherwise is selling something. But there are meaningful signals worth understanding, and knowing their limits is just as important as knowing how to apply them.
What AI text tends to look like
Before reaching for a detection tool, it helps to train your editorial eye. ChatGPT and similar models tend to produce text that is syntactically smooth but tonally flat. Sentences are well-formed, transitions are tidy, and the argument rarely surprises you. Hedges cluster at the start and end of paragraphs. Phrases like "it is important to note" or "in conclusion" appear with a frequency that human writers, especially experienced ones, learn to avoid.
Researchers studying large language model outputs have noted a preference for what is sometimes called "safe" vocabulary: common collocations, generic analogies, and a reluctance to commit to strong claims without qualification. If a submission reads as though every sentence has been sanded down to remove friction, that is worth a second look. So is an unusual uniformity of sentence length across several paragraphs.
That said, these are tendencies, not rules. A careful human writer can produce smooth, hedged prose. A careless one can produce writing that looks more erratic than any AI would generate. Stylistic suspicion is a starting point, not a verdict.
What detection tools actually measure
Tools like GPTZero, Originality.ai, and Turnitin's AI writing detection feature work by measuring statistical properties of text: chiefly "perplexity" (how surprising a word sequence is to a language model) and "burstiness" (how much sentence length varies). Human writing tends to be "burstier" and more perplexing to a model than AI output. Our explainer on the technology behind AI content detection covers these mechanics in more depth.
The critical caveat: these tools produce probability scores, not authorship verdicts. As our facts section states clearly, detectors are probabilistic and cannot prove authorship. A score of 85% "likely AI" means the text shares statistical properties common in AI output. It does not mean the text was written by a machine. Editing, paraphrasing, or even writing in a deliberate style can shift those scores dramatically.
If you want to understand why different tools reach different conclusions on the same text, our piece on the accuracy problem in AI detection explains the underlying disagreements.
The bias problem you cannot ignore
One finding that every editor using these tools must keep in mind comes from a 2023 Stanford study by Liang et al., published in the journal Patterns. The researchers found that AI detectors showed systematic bias against non-native English writers, flagging their text as AI-generated at significantly higher rates than text produced by native speakers. The reason is structural: non-native writers often use simpler, more predictable vocabulary and sentence structures, which look statistically similar to AI output.
In a newsroom context, this is not a minor footnote. If you are evaluating contributions from journalists working in a second or third language, running their copy through a detector and treating the score as evidence of AI use is, at best, methodologically unsound and, at worst, discriminatory. The Liang et al. finding should be prominently in mind whenever a score is cited against a contributor.
A practical editorial approach
Given all of the above, what does a responsible editorial workflow look like? We would suggest treating AI detection as one input among several, never as a standalone conclusion. A reasonable set of steps might include:
- Reading for stylistic signals: unusual flatness, generic transitions, absence of specific detail or sourcing that only a reporter present could provide.
- Running the text through one or more tools and noting the score, while keeping the probabilistic nature firmly in view. Our 2026 comparison of GPTZero, Originality.ai, and Turnitin is useful context for understanding what each tool is actually measuring.
- Checking for verifiable specifics: named sources, direct quotes, datelines, and granular detail that would be difficult to fabricate or synthesise from training data alone.
- Having a direct conversation with the contributor if concerns persist. Ask about their reporting process. Genuine reporters can walk you through their work.
- Documenting your reasoning, whatever conclusion you reach. If a decision is later challenged, a record of the process matters.
It is also worth remembering that the landscape is shifting quickly. The deep technical dive into how AI content detection works makes clear that as models improve, the statistical gap between human and AI writing is likely to narrow further.
What detection cannot do
Detection tools cannot tell you how a piece was used. A journalist who used ChatGPT to draft a structure, then rewrote every sentence, may produce text that scores low on AI probability while still raising editorial questions about disclosure. Conversely, a human writer who favours plain, clear prose may score higher than they should. The question of if a piece meets your publication's standards is ultimately an editorial judgment, not a statistical one.
Our newsrooms need frameworks for AI disclosure and use policy far more than they need perfect detection. The tools are useful as a prompt for further inquiry. They are not a substitute for editorial judgment.
Sources
- Liang, W. et al. (2023). "GPT detectors are biased against non-native English writers." Patterns, Cell Press. (Stanford study)
- GPTZero product documentation, gptzero.me
- Originality.ai product documentation, originality.ai
- Turnitin AI writing detection documentation, turnitin.com