Grammarly is one of the most familiar writing tools in any newsroom. When the company added an AI detection feature to its platform, many editors reasonably assumed it carried the same credibility as its grammar and style suggestions. That assumption deserves a closer look. The question of if Grammarly's AI detector is accurate enough to inform editorial decisions is more complicated than the feature's clean interface suggests.
How Grammarly's AI Detection Works
Like most commercial AI detectors, Grammarly's tool uses statistical and linguistic pattern analysis to estimate the probability that a given piece of text was generated by a large language model. It looks for markers such as low perplexity (predictable word choices), high burstiness uniformity, and sentence structures commonly associated with models like GPT-4. The output is a probability score, not a verdict. For a fuller explanation of the underlying methodology, see our piece on how AI content detection works.
The important word in that description is probabilistic. Grammarly's detector, like every other commercial tool in this category, cannot prove authorship. It can only report that a piece of text shares characteristics with machine-generated writing. That distinction matters enormously when the result is used to challenge a journalist or contributor.
What the Research Tells Us
The most cited piece of independent research on AI detector accuracy is a 2023 Stanford study by Liang et al., published in the journal Patterns. The researchers found that AI detectors disproportionately flag writing by non-native English speakers as machine-generated. The reason is structural: non-native writers often produce text with lower perplexity and more predictable syntax, precisely the features detectors associate with AI output. A foreign correspondent writing in their second or third language could be flagged simply for writing carefully and clearly.
This finding has direct implications for newsrooms with international contributors or reporters working across language contexts. If an editor acts on a Grammarly flag without understanding this bias, they risk a serious misattribution. The accuracy problem across AI detectors is not unique to Grammarly, but the platform's mainstream presence means its results are more likely to be treated as authoritative by editors who are not specialists in detection technology.
Known Limitations of the Tool
Beyond the non-native speaker bias, several other limitations are worth naming explicitly:
- False positives on edited AI text: If a journalist uses an AI draft as a starting point and then substantially rewrites it, the detector may still return a high AI probability score, or it may not, with no reliable consistency.
- False negatives on sophisticated AI output: Newer models, particularly those that have been fine-tuned or prompted to write in a human style, can produce text that scores low on AI probability even though it was machine-generated.
- No published accuracy benchmarks from Grammarly: As of this writing, Grammarly has not publicly released precision and recall figures for its AI detection feature, making independent verification of its claims difficult.
- Inconsistency across versions: Comparisons of multiple detectors on the same text, including Grammarly, reveal significant disagreement. Our analysis of GPTZero, Originality.ai, and Turnitin illustrates how different tools can return contradictory results for identical submissions.
- Sensitivity to paraphrasing tools: Text that has been run through a paraphrasing tool can evade detection, meaning a determined bad actor is not meaningfully constrained by the detector.
How Editors Should Use (and Not Use) This Tool
None of the above means Grammarly's AI detector is worthless. It can be one input among many when an editor is trying to understand a piece of text. The problem is treating it as a conclusive or sufficient basis for an editorial decision.
A practical approach is to treat a high AI probability score the way we would treat any single anomalous data point: as a reason to ask questions, not as evidence of wrongdoing. If a contributor's writing style has changed significantly, if a piece arrives unusually quickly, or if the sourcing feels thin, those contextual signals combined with a detector flag may warrant a conversation. The flag alone does not.
Editorial policies that rely on AI detection scores as proof of policy violation are on shaky ground. The probabilistic nature of these tools means that disciplinary action based solely on a detection result could penalise honest journalists, particularly those writing in English as a second language. This connects to the broader challenge of understanding what detection tools actually measure before building policy around them.
A Note on Platform Trust
Grammarly benefits from accumulated trust built on years of genuinely useful grammar assistance. That trust can create a halo effect around newer, less proven features. The AI detection function is a different kind of product from the grammar checker: it makes probabilistic claims about human behaviour rather than correctable claims about language rules. Those two things should not inherit the same degree of institutional confidence.
As newsrooms continue to develop AI use policies, the question of if any single detector can be used as an authoritative arbiter of authorship deserves honest scrutiny. For a side-by-side view of how Grammarly compares with other tools on the same material, the seven-detector comparison offers useful context.
Sources
- Liang, W., et al. (2023). "GPT Detectors Are Biased Against Non-Native English Writers." Patterns, Cell Press. DOI: 10.1016/j.patter.2023.100779.
- Grammarly product documentation (grammarly.com) on AI detection feature.
- Stanford HAI reporting on large language model detection limitations, 2023.