Is the JustDone AI Detector Accurate? What the Tests Show

By Jack Stovell X · Instagram · LinkedIn · 2026-10-10 · Guides

Short answer: nobody has independently checked JustDone's accuracy. Its blog claims 94.1%, its own detector page says 80%, one retest found 59 to 61%, and AFP watched it score a human-written news report as 88% AI. Scores are least reliable on short, mixed or plainly written human text.

What JustDone claims, and where the numbers come from

JustDone's blog post "Best AI Detectors 2026 Tested" ranks seven tools and puts JustDone first, "scoring 94.1% accuracy with the fewest false positives on ESL and hybrid writing across a hands-on test of 15+ tools." The same list gives Turnitin 92.8% and GPTZero 87.3%. The post lists the kinds of sample it used (academic essays, ESL writing, and human, fully AI and hybrid text) but never says how many texts were in each group. Some of the other tools' figures are labelled as declared accuracy, which means they weren't measured in the test at all.

JustDone's detector page gives different numbers: "80% overall detection accuracy across all text types", 98% on academic and scientific texts, and a 10.3% error rate. It calls these independently benchmarked, but the page doesn't say who ran the benchmark or how many texts it used.

So the vendor publishes two headline figures 14 points apart. Plenty of companies rank their own product first, and that alone doesn't prove bad faith. But nobody outside JustDone can check either number.

To its credit, the same detector page says "every AI detection result is a probability, not a verdict" and admits no detector is completely accurate. That advice is sound, and it applies to JustDone's own scores too.

What independent reviewers and journalists report

The outside picture is lower and patchier. None of these tests shares a sample set, the samples are small, and most of the reviewers have something to sell.

The methods and motives differ, but the same failure keeps turning up: human writing scored as AI.

Why detector accuracy figures disagree

The gap between 94.1% and 59% comes mostly from how each test was set up.

  1. Sample choice. Detectors look sharp on raw machine output and wobble on everything else. In one study of 14 tools summarised in our AI detector accuracy reference, average accuracy was 96% on human text and 74% on unmodified AI text, then fell to 42% on lightly edited AI text and 26% on machine-paraphrased text. Whoever picks the samples largely picks the result.
  2. Thresholds. Where you draw the line between "AI" and "human" changes the headline number. A strict cut-off catches more AI and flags more people, and a loose one does the reverse.
  3. Mixed text. Real documents are often partly human, partly machine and partly edited. One score for the whole thing hides the mix, which is how a 59% AI text came back 78% human.
  4. Free versus paid, and run to run. By JustDone's own account, its free tier uses a lighter model. Phrasly's repeat test suggests scores can also change between sessions on identical text.

The mechanics behind these failures (predictability scores, two-model comparisons, trained classifiers and watermarks) are explained in how AI detectors work.

False positives: who gets flagged wrongly

When a detector gets it wrong, a real person takes the blame.

The clearest evidence is general rather than specific to JustDone. In a Stanford study, seven detectors flagged non-native English essays (91 TOEFL essays) as AI 61.3% of the time on average, against 5.2% for essays by US students. Detectors reward varied, unpredictable wording, so plain, careful or formulaic prose looks suspicious to them. That catches second-language writers, people who follow a strict style guide and anyone whose draft has been tidied by a grammar tool. There's more in AI detectors and non-native English writers.

JustDone says it has the fewest false positives on ESL writing. We found no independent test that backs this up, and the Phrasly and AFP results point the other way.

Are Turnitin and JustDone the same?

No. They're separate products from separate companies; JustDone's footer names GM Appdev Limited of Cyprus. JustDone offers an "estimated Turnitin AI score", but its own page says "JustDone is not affiliated with Turnitin, and institutional results may differ." The estimate is JustDone's guess at what another system might say, and a free checker can't tell you what your institution's tool will report.

How to use a detector score sensibly

Treat a score as the start of a review, whether you're checking someone else's work or your own.

  1. Check the length and type of text. By MPG ONE's figures, JustDone is weakest on short passages, and every detector struggles with edited or mixed text.
  2. Rerun it and try a second tool. If the same passage gives different scores, or two detectors disagree, that disagreement is the finding.
  3. Read the text itself. Repeated phrasing, stock transitions and flat rhythm are worth looking at. ScriptGrain's AI detection highlights these writing patterns. Use it to know where to look, and leave the judgement about who wrote the text to a person.
  4. Keep your record. Drafts, notes and version history show how a piece came together in a way no percentage can.

If you've been flagged, don't run your own writing through a humaniser to lower the score. It changes your words and buries the evidence of how you wrote them. Gather your process evidence instead: how to prove you didn't use AI walks through it, and falsely accused of using AI at work covers the workplace version.

Questions

Is JustDone AI detector accurate?

Nobody has checked it independently. JustDone's blog claims 94.1% and its detector page claims 80% overall. One retest found about 59 to 61%, and AFP saw it score a human-written report as 88% AI. Results are weakest on short, mixed and plainly written human text.

Are Turnitin and JustDone the same?

No. JustDone is a separate product, and its detector page says it is not affiliated with Turnitin. Its "estimated Turnitin AI score" is JustDone's own estimate, and your institution's result may differ.

Is 20% AI detection bad?

On its own, no. A score is a probability from a model that makes mistakes, and it can't tell you which sentences, if any, a machine wrote. Turnitin only publishes its under 1% false positive figure for documents scoring above 20%, so even Turnitin treats lower scores as less reliable.

Do AI detectors actually detect AI?

Reasonably well on clean, unedited machine output, and much less well on anything else. In one study of 14 tools, average accuracy fell from 74% on unmodified AI text to 42% on lightly edited AI text, and detectors also flag real human writing.

What should I do if a detector flags my own writing?

Gather your drafts, notes, version history and sources. Explain how you wrote the piece, and ask what the score is being used to show. Don't run your own work through a humaniser to lower the score: it changes your words and buries the evidence of how you wrote them.

More from the Journal

More from the ScriptGrain Journal