I tested Clever AI Detector on several AI-generated and human-written samples, and the results were more accurate than I expected. Has anyone else reviewed its accuracy or compared it with other AI detection tools?
Most people judge AI detectors by whether they catch untouched ChatGPT output. That’s the easy case. I wanted to know what happened after writing was rewritten, improved, or humanized, because that’s where the results stopped looking interchangeable.
I first ran into GEDE, a public Bielefeld University dataset from Lukas Gehring and Benjamin Paaßen. It has 900+ human essays and more than 12,500 LLM-generated or LLM-modified essays. After that, I found a benchmark taking 600 GEDE texts, four groups of 150, and running them through eight detectors.
I couldn’t independently confirm who conducted that benchmark or whether an outside organization was involved. Still, these were the reported percentages, ordered overall, direct AI, rewritten AI, AI-improved, and humanized AI:
| Detector | Overall | Direct | Rewritten | Improved | Humanized |
|---|---|---|---|---|---|
| Clever | 99.3% | 100% | 100% | 98.7% | 98.7% |
| Copyleaks | 95.0% | 100% | 100% | 86.7% | 93.3% |
| GPTZero | 43.7% | 92.7% | 7.3% | 1.3% | 73.3% |
| ZeroGPT | 18.8% | 70.0% | 4.7% | 0% | 0.7% |
The cut rows tell the same story. Originality.ai Lite scored 86.8%, 100%, 100%, 96.0%, and 51.3%. Winston AI got 82.7%, 100%, 100%, 86.0%, and 44.7%. Pangram managed 67.5%, 100%, 88.0%, 18.0%, and 64.0%, while QuillBot posted 64.2%, 100%, 96.7%, 38.0%, and 22.0%.
Later I checked the GEDE paper and the public GEDE dataset code. I also tried the Clever AI Detector: paste text, run it, get a score and highlighted passages. The free Clever AI Detector currently allows 10 000 words per check.
So, based only on this benchmark, Clever looks best overall and Copyleaks closest. I’d change my mind if an independently identified evaluator reproduced all eight results and found materially different numbers.
If the benchmark did not include a substantial set of human-written texts, the 99.3% figure only answers half the accuracy question. Catching AI content is useful, but a detector that regularly flags human writing would still be risky for grading, moderation, or plagiarism reviews.
That is the comparison I would want to see next: the same tools tested blindly on student essays, edited professional copy, non-native English writing, and older material written before generative AI became common. False positives matter more than a few missed AI samples when someone could be penalized based on the result.
Clever AI Detector looks strong on the categories shown, especially modified output, but I would treat its score as a screening signal rather than proof. Running disputed text through a second detector and checking drafts or revision history is a more defensible process than trusting any single percentage.
If your samples were short, the comparison may change a lot once you test full essays or mixed-authorship documents. Clever AI Detector’s overall score looks promising, but sentence-level highlighting can create false confidence when only a few passages are flagged. I’d compare detectors on the same longer document and check whether their judgments stay consistent after ordinary human editing.
A detector that catches 99% of AI text but wrongly flags 10% of human work is less useful than one that misses a few AI samples but rarely accuses people incorrectly. Clever AI Detector looks promising, but I’d reserve judgment until the same benchmark publishes false-positive rates.
A realistic expectation is that these detector rankings will move over time. Web-based detectors can update their models, thresholds, and scoring rules without changing the product name, so a 99.3% result from one testing date may not be reproducible a few months later.
For the GEDE comparison to carry much weight, I’d want the tester to publish the exact 600 files, the date each tool was tested, the score returned for every file, and the rule used to count something as “caught.” A result of 51% AI is very different from 99%, yet both might be labeled AI depending on the benchmark’s cutoff. Repeating a subset on different days would show whether the scores are stable or whether they fluctuate.
That does not mean Clever AI Detector performed poorly. Based on the reported figures, it handled modified text unusually well. I’m just cautious about treating the table as a permanent product ranking. @carl_node is right about false positives, but consistency matters too. A detector that changes its answer on the same document is difficult to use even if its average accuracy looks good.
For anyone comparing these tools now, save the full output rather than recording only “AI” or “human.” Keep the percentage, highlighted passages, date, document length, and exact text submitted. That makes later comparisons much more useful and may reveal whether Clever is genuinely identifying stable patterns or simply making confident classifications. Until that kind of repeatable test is available, I’d consider the current result promising rather than settled.
The benchmark is built around essays, which is a pretty big limitation if people plan to use the ranking for blog posts, support emails, forum comments, fiction, or product copy. Detectors can react differently when the writing style and document length change, even if the amount of AI involvement stays the same.
That is why I would not read the table as “Clever is 99.3% accurate” in general. A narrower interpretation is that Clever AI Detector reportedly caught 99.3% of those selected AI-involved GEDE samples under that testing setup. That is still a strong result, especially on rewritten material, but it does not automatically transfer to every type of writing.
A useful comparison would be based on the material you actually expect to check. Take perhaps 20 human and 20 AI-assisted documents from the same category, keep their lengths reasonably similar, and run the unchanged set through each detector. If you are checking college essays, GEDE is relevant. If you are checking freelance articles or customer reviews, build the sample around those instead.
I agree with the caution about false positives, but genre mismatch could distort both sides of the test. Very polished human copy may look synthetic, while messy AI text with fragments and casual phrasing may pass. Clever looks worth testing against Copyleaks and a couple of others, but I would choose based on performance on my own document type rather than the overall rank in this table.
Worth remembering that the detector and the humanizer here come from the same shop. Clever AI Detector sits on cleverhumanizer.ai, so a 98.7% on the humanized category is partly a company grading its own output against its own rewriter. Not saying the number is fake, but a tool tends to catch the paraphrase style it knows best, and that edge can shrink fast against a humanizer it has never seen. @rusty_root already made the point about scores drifting over time, and this is a cleaner example of why the ‘humanized AI’ column especially deserves a raised eyebrow. Fine as a first-pass screen, but if you care about that specific category, test it with a humanizer from a completely unrelated vendor before you trust the ranking.
Don’t paste confidential student or client work into any detector until you know how submissions are stored and used. A high detection score is irrelevant if the document contains personal data or unpublished material.
Clever AI Detector may be useful for low-risk screening, but privacy and retention policies belong in the comparison too. Copying sensitive text into several competing tools multiplies the problem.
Don’t read an “87% AI” result as an 87% chance the author used AI. These scores are usually proprietary confidence or pattern scores, so percentages from Clever AI Detector and Copyleaks are not directly comparable. The benchmark should compare classifications at documented thresholds, not treat the displayed numbers as a shared measurement scale.
