Section: Technology & AI Author: IN Reading Length: ~27 min Sources and further reading: 22 items Topics: AI, education, academic integrity, text detection, accountability SEO / Working Title: How to recognise text written by artificial intelligence?
The question is: how to recognise text written by artificial intelligence? It sounds simple until we notice that we are not comparing one thing. Statistical score is not origin. Watermark is not a detector of arbitrary chatbot output. A low FPR in a lab is not a low FPR in a non-native writer. Turnitin Similarity is not Turnitin AI writing. Only when these four things are separated, it is possible to read what the classifiers can do by September 2026, what the shared tests say about them and why the Czechpublic policies of Charles University and CTU do not take the detector as proof of origin.
1. School, supplier and laboratory are not measuring the same thing
Let's imagine a college committee. The teacher will bring a text that seems too smooth. The student claims he was working alone, just getting his tongue fixed. The study department wants a clear rule because cases are increasing and every decision is subject to appeal. The supplier shows the percentage in the dashboard. The lab shows the accuracy table on the shared set. Charles University writes in a public recommendation that a technically reliable method of detectiongenerative AI does not exist [12]. CTU prohibits the use of AI tools to detect student solutions due to false positives and false negatives in the methodological instruction from 2023 [13].
At first glance, it is a dispute whether "we can know AI". In fact, each describes a different object. The supplier measures the flag according to his model and his threshold. The lab measures performance on a specific set of texts, in a specific language, against specific attacks, and with a specific calibration. The university decides whether a number can enter a particular person's sentence. The student bears no consequences for the average accuracy of the instrument. It carries consequences for the conclusion of one text, in one procedure, with onedefense options.
Therefore, everyone can be partly right and the joint decision nevertheless wrong. Turnitin is correct that its tool returns numerical scores and that it processed a large operational volume of submissions in May 2023 [10]. The researchers are right that detectors can be measured and that some signals carry information. UK and CTU are right that a disciplinary verdict is not a laboratory graph. Once the percentage from the dashboard is translated into the sentence "the student cheated", the subject of the measurement has changed. And that's where the most expensive mistake occurs.
The detector measures statistics. It doesn't matter who wrote the text.
— Jiný Kontext
2. One score has at least five layers
One number at the end of the analysis looks simple. But it arises in a chain of layers. The first is the signal. It can be perplexity, classifier, watermark, comparison with saved output or metadata. Each signal answers a different question and each fails differently. Perplexity punishes unusual or, conversely, overly predictable language. The classifier learns the boundary between the sets. Watermark looks for a statistical trace that the generator must have inserted beforehand. Metadata will helponly when they exist and has not been separated from the text.
The second layer is the threshold. A tool can be useful in one setting and dangerous in another. The school doesn't just need to know that the classifier has a nice graph. He needs to know the false alarm in the individual because the punishment falls on the person, not the average. RAID therefore calibrates the true positive rate at a false positive rate equal to 5%. This is more accurate than the unchecked default slider, but still no excuse for a disciplinary automaton [6].
The third layer is the text population. An English essay is not a Czech term paper. The code is not the introduction to the diploma. The text of a native writer is not the text of a person who writes in a second language. A translated paragraph is not evidence of a change of author. Changing the domain may change the order of the tools because the model does not recognise the author. He recognises the similarity to what he has learned as human and machine.
The fourth layer is the adjustment. The text may be paraphrased, translated, redacted or mixed with human work. This article does not describe how to circumvent school rules. It only describes that the modification changes the stat. If the signal changes and the authorship does not, the score cannot be a seal of origin.
The fifth layer is decision. The same result can be an indicative stimulus for the interview, but it must not be independent evidence of fraud. This layering is not a detail for methodologists. Determines whether the number belongs to an educational conversation, to a rule modification, or to a procedure that can harm the student.
3. Similarity statistics are not a seal of origin
The classifier did not see a person at the keyboard. He does not know whether the student wrote the first draft, used the model for the outline, had his English corrected, translated the text back and forth, or took over the finished output. It sees the text and compares its characteristics to what it considers human and machine in its set. This may give rise to suspicion. Authorship cannot arise from this alone.
That's not a pun. The origin of the text is a historical question. Who did what, when, to what extent and with what degree of independence. Statistical scoring is a matter of distribution. How similar is the text to one class versus another. There is a bridge between these questions, but it is not automatic. The bridge consists of knowledge of the process, assignment, interim version, consultation, declaration, defense and the possibility to explain the work.
DetectGPT is a useful example. It is based on the idea that machine-generated text can have a different probabilistic curvature than human text [15]. This is an elegant statistical hypothesis. It is not a file origin record. As the model, language, domain, or text editing changes, so do the signal properties. Sometimes it helps. Sometimes it breaks. The conclusion "resembles the output of the model" is not the same as "this person used AI inappropriately".
The school therefore needs to distinguish three sentences. First: the tool highlighted the text. Second: the text has characters worth talking about. Third: the student broke the rules. The first sentence can be said by the software. The second requires a teacher. The third requires a process.
4. Watermark is not a detector of any chatbot
Watermark in the text sounds like a solution. The generator would select some words with a small statistical shift during creation, and the detector would later verify whether the shift remained in the text. Kirchenbauer and colleagues described a scheme that, according to the 2023 paper, enables detection from as little as a short stretch, in a laboratory setting reported as "as few as" 25 tokens [3]. This is an important research result. It is not a universal text detector from the Internet.
Watermark works only if the generator really used a particular scheme, if the detector knows the necessary configuration and if the text has not undergone a change that breaks the trace. If someone feeds regular output from a public chatbot into the detector, which does not have a publicly verifiable watermark applied and published for review, it does not look for a seal. It looks for statistical similarity. That's a different tool.
As of September 2026, it cannot be inferred from public materials that major consumer chat services give the average user a verifiable text watermark according to this scheme. After withdrawing its classifier, OpenAI talks about provenance in other modalities and broader issues of provenance, but has not returned a public reliable text classifier that alone addresses term paper provenance [1].
This does not mean that watermarks have no value. It follows that their price is different. They can be useful in a closed system that controls both generation and validation. In a public and mixed environment where text travels through editors, translations, copying and human revision, a watermark does not become a court stamp.
5. A low FPR in a lab is not a low FPR in a non-native writer
False positive rate looks like a technical figure. However, in school, it is the probability that an innocent text will be marked as suspicious. And this probability is not a property of the tool itself. It is a feature of the tool on a certain population of texts.
In 2023, Liang and colleagues examined 91 human TOEFL essays and 88 essays by eighth graders in the United States. With seven detectors, including GPTZero, they measured an average false positive rate of 61.22% for TOEFL essays. Eighteen of the 91 TOEFL essays, or 19.78%, marked all seven detectors. At least one detector marked 89 out of 91 essays, or 97.80% [8]. That's a strong number, but it has precise limits. It measures human English TOEFL essays by non-native writers in access to tools from March 15, 2023. It does not measure Czech seminary in 2026.
Still, it's important. It shows that the writer population is part of the measurement. A more limited vocabulary, a simpler sentence structure or a learned school style can look "model" without the text writing the model. To punish such a signal is to punish style, linguistic background or genre. This is exactly the kind of error that institutions must not hide behind the word objectivity.
It does not follow that every detector today makes the same mistake to the same extent. It follows from the fact that it is not enough to take the FPR from an English or supplier test and apply it to Czech students. Whoever wants to use the number in the proceedings must show its validity for the given language, genre, length and population. Without it, a low FPR in a presentation is just a distant figure.
6. Turnitin Similarity is not Turnitin AI writing
Teachers know the similarity check with sources. This has its own problem, but at least it measures another and understandable thing: whether a text matches another text. Turnitin Similarity is not the same as Turnitin AI writing. A match to a source may indicate a copy, a misquote, or a common phrase. The AI writing score asks if a piece of text looks machine-written. To mix these results in one sentence is an object error.
The example is simple. A student can honestly cite a source and have high similarity because the bibliography, legal wording, or assignment is the same. Another student may paraphrase the text so that similarity decreases, but authorship remains problematic. A third student may write alone, but in a genre that is formal, repetitive, and academically polished. Each case needs a different question.
Turnitin itself distinguishes between thresholds and limits in its materials. For the AI writing model, it states a minimum length of 300 words and describes how it handles low markup [11]. In the 2023 operational text, they talk about a document-level false positive rate of less than 1% for documents above 20% AI writing on an internal test of 800,000 academic texts before ChatGPT, and a sentence-level false positive rate of around 4% [10]. These are vendor numbers, with a defined threshold and custom set. They are not the evidence rules of the Czech school.
Therefore, the correct administrative sentence is not "Turnitin showed the percentage, the case is clear". It reads: "In one task, the tool marked the text; we need to find out what it measured, whether the length, language, threshold fit, and whether there is an independent procedural support." That's slower. But a school that punishes faster than it distinguishes does not punish accuracy. Convenience punishes.
7. Fourteen instruments on fifty-four documents are not an operational guarantee
Weber-Wulff and colleagues tested a total of 14 detectors in the winter and spring of 2023, including 12 public ones, as well as Turnitin and PlagiarismCheck. The set had 9 documents in 6 classes, so 54 documents, in English. In a strict binary cut, no tool achieved over 80% accuracy; the best Turnitin had was 76%. He had 79% in the semi-binary cut and 81% in the logarithmic evaluation [5].
This study has limitations. The set is small. The language is English. The tools have changed since 2023. It's not a verdict that a particular version of a particular tool can't do anything useful in 2026. However, it is a good readable demonstration of something else: how easily the detector behaves differently according to text class and evaluation method.
A small shared methodology has the advantage of showing the mechanics. Accuracy is not one thing. A binary result penalizes a different failure than a softer classification. Translated texts, human edited texts and mixed texts are not just noise. These are exactly the situations in which schools live. The student does not need to submit the clean output of the model. It can have its own concept, language proofreading, translation, consultation and citations. It is in these intermediate states that the disciplinary decision is the most difficult.
Therefore, fourteen tools on fifty-four documents are not an operational guarantee. It's a warning against translating a lab chart into a one-man trial. Anyone who wants to use the detector responsibly must know what a mistake looks like on their own texts. Not on the text from the ad.
8. RAID shows crash when changing sampling. It is not a Czech diploma
RAID is a bigger and tougher test. In ACL 2024, Dugan and colleagues described a shared benchmark for robust evaluation of machine-generated text detectors. The dataset contains more than 6 million generations, 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies; evaluates 8 open and 4 closed detectors [6]. The authors observe a true positive rate with a false positive rate of 5%. This is important because without calibration one can look good on a metric that is not acceptable to the school.
At the same time, RAID shows fragility. According to the paper, the repetition penalty reduces accuracy by up to 32 percentage points [6]. For naive thresholds, a false positive rate of up to 1.7% appears for closed-source tools; open metrics are often significantly higher [6]. These numbers have exact limits. It belongs to a given benchmark, models, domains and settings. They are not a direct measurement of the Czech diploma thesis.
That is why they are useful. They don't say "never use any detector". They say: "changing the sampling, domain and adjustment changes the result so much that the default threshold is not proof." RAID-extra also contains Czech and German news and code [7]. This shows that there is a Czech language stress test in another domain. It does not mean that someone has validated the disciplinary use of Czech term papers and final theses.
The public benchmark is a map. It's not a judgment. It will help to find out which signals are worth paying attention to, where tools fall and how they behave when conditions change. The decision about the student must stand elsewhere.
9. The downloaded classifier is proof of the limit. There is no proof that the new tool has no limits
OpenAI released the AI Text Classifier in January 2023. In July of that year, it was withdrawn for low accuracy [1]. On the English challenge set, it reported that the tool correctly labeled 26% of AI texts as "likely AI-written" and falsely labeled 9% of human texts [1]. The threshold was set carefully so that there were not too many false positives. The cost was clear: many AI texts were missed by the tool.
This example tends to be used too quickly. One side will make it proof that detection as a field is worthless. The other will say that the old tool is a thing of the past and there is no need to deal with it. Not even one sentence is enough. It is an audit of one tool of one company, one English set and one period. It is not a measurement of Czech, code, short texts or mixed authorship.
At the same time, it is proof of a limit that cannot be swept away. When even a large provider pulls a public grader because it's not accurate enough, a school can't claim without further support that a different percentage from another service is enough to blame. A new tool may be better. However, it must be shown in the relevant population and at a threshold that corresponds to the consequences.
The useful question is not whether the detectors are getting better. It sure is changing. It sounds whether their error remains acceptable for a particular decision. For an orientation notice, the threshold may be different than for disciplinary proceedings. For an interview with the student, other than for canceling work. The word "suspicion" and the word "evidence" must not share the same slider.
10. Paraphrasing breaks statistics. This text is not a guide to school
Sadasivan and colleagues at work Can AI-Generated Text be Reliably Detected? show that paraphrasing can reduce the performance of both detectors and watermarks [2]. For watermarked output in an experiment on about 100 XSum passages, the accuracy of the detector after PEGASUS paraphrase dropped from 97% to 80% and perplexity increased by 3.5 [2]. For DetectGPT on 200 XSum passages, generated by GPT-2 Medium and paraphrased by the T5 model, the AUROC dropped from 96.5% to 59.8% [2]. Recursively paraphrasing DIPPER, the dossier reports a drop from 82% to 18% AUROC [2][16].
These numbers are not a recipe. They don't tell the student what to do. They don't even tell the school that every redacted text is suspect. They say that the statistical signal can be changed without simply changing the historical origin. Text can go through machine, human, translation and editing. Authorship does not become binary just because the tool returns a binary flag.
AUROC 50% corresponds to chance. When the adjusted result approaches or falls below random resolution, it is wrong to base the penalty on the notion that the score is holding steady. It's even worse to think that a low score automatically clears and a high score automatically convicts. Both errors have the same root: mistaking statistical similarity for ancestry.
Therefore, the paraphrase in the detection article should be mentioned with caution. Not as a walkthrough. As a test of the strength of the claim. If a small change in wording significantly changes the result, the result is not a seal. It is a signal with a limit.
11. Declaration of use is not detection
Some schools are responding to generative AI with a mandatory statement. This is different from detection. The statement does not say that we can retroactively recognise every use of the model. It says that the author should describe his work and take responsibility for how the tool is used.
The Faculty of Humanities of the Charles University, in the dean's measure No. 3/2026, introduces a mandatory statement on the use of generative AI in the final thesis [14]. CTU works with a statement on artificial intelligence in its final theses [20][21]. It's not a black box. It is a procedural element. It forces the author to distinguish between research, outline, language proofreading, text generation and own responsibility.
The AI Act in Article 50 and the related FAQ of the European Commission address transparency obligations for certain types of outputs and texts in the public interest [22]. It does not solve the origin of each seminar paper. It's not a university detector. It is a legal framework for labeling and transparency in another decision-making space.
So the statement can be reasonable. It is not a substitute for proof. It helps to set the rules, improve the conversation between the supervisor and the student, and say in advance what is allowed. However, if the violation is addressed retrospectively, it is not enough to compare the statement with the percentage of the instrument. You need to read the work, follow the process and ask questions about things that the author of the work should be able to defend.
12. Charles University does not accept the detector as evidence. CTU bans it for detection
Czech public policies are more sober on this point than many debates. In its recommendation for theses, Charles University states that there is no technically reliable way to detect generative AI [12]. CTU states in methodological instruction No. 5/2023 that AI tools should not be used to detect student solutions, precisely because of false positives and false negatives [13]. Later CTU materials work more with a statement about the use of AI than with the promotion of the detector to evidence [21].
This does not mean that schools are giving up on academic integrity. It means they are not looking for it in one black box. The supervisor may know the development of the text, consultations, literature, formulation of the problem and the student's ability to defend the procedure. The committee can examine whether the student understands his own statements. The rules can distinguish permitted language assistance from unauthorized takeover of output.
In the context of AI research tools, the National Technical Library notes that the detectors are indicative [19]. That is the exact word. Indicative does not mean useless. It means that the result shows the direction for the next question. No replaces the answer.
In a disciplinary context, this restraint is also humanly important. A false alarm is not just an error in the spreadsheet. It is accusing a person of dishonesty. And such an accusation needs more than a similarity statistic.
13. A flag above twenty percent is not a fraction of fraud
Turnitin reported in May 2023 that it processed 38.5 million submissions in the 7 weeks following the detector preview. Of these, 9.6% scored above 20% AI writing and 3.5% scored between 80 and 100% AI writing [10]. These numbers often boil down to a simple sentence: so many students are cheating. That would be a bad translation.
The unit of measure is Turnitin submissions, including assignments where AI may have been enabled or where use of the tool should have been acknowledged. A score above 20% is a flag by the tool, not a rule violation. A document with a high score may be a case for verification. It is not a writing in itself.
At the same time, the supplier works with its own internal test of 800,000 academic texts before ChatGPT, and for documents above 20% AI writing, it reports a document-level false positive rate of less than 1% [10]. Even this number has a limit. Applies to supplier methodology, specific threshold and set. For low values, Turnitin itself relativized the results and in the documentation states limitations for shorter texts [11].
A simple rule follows from this: the flag is not a share of fraud. It's an administrative signal. If the school uses him for an interview, they must separate him from guilt. If he uses it for punishment, he needs to have another support. Otherwise, the orientation becomes a judgment.
14. The best decision is the one whose false alarm you bear
We often ask the tool how often the AI captures text. With the institution, it is equally important how often it damages the human text. The following diagram is not a tool chart. It is a map of errors that the school must name in advance.
| A kind of failure | What does he look like? | How to test | What will limit the damage |
|---|---|---|---|
| Score as origin | GPTZero or Turnitin will translate the sentence "guilty" | Two detectors on known human text; monitor the discrepancy | Detector outside punishment, statement and defense |
| Laboratory FPR on a different population | The TOEFL essay is marked as AI | Liang: the population of the writer is part of the measurement [8] | Do not read English FPR as Czech; protect non-native writers |
| Similarities like AI writing | Matching the source is read as machine authorship | Separate tools, outputs and conclusions | Two columns, two questions, two rules |
| Default threshold | "99% accuracy" without knowledge of threshold and set | RAID: TPR at FPR 5% [6] | Calibration and prohibition of automatic penalty |
| A short text like a diploma | A paragraph is graded in the same way as a long paper | Minimum 300 word check on Turnitin [11] | Do not judge below the threshold of the instrument |
Illustrative diagram: The detector stands in the first column, the text population in the second and the school's decision in the third. The arrow from the first column to the third does not lead directly. Leads through validation, advocacy and rules.
So the best decision is not the one that closes the case as quickly as possible. It is a decision whose error can be explained, corrected and humanely abducted. For low-risk use, the detector can help line up texts for manual review. The threshold for punishment must be stricter because a false alarm is not just a technical loss. It's a reputational hit.
15. A text whose origin you know will tell more than a percentage of the supplier
The practice test doesn't have to be complicated. It's not supposed to prove that someone cheated. It should show whether the institution understands its own instrument. Take a text whose origin you know: your own paragraph written before ChatGPT, an archived teacher's text, an old student paper with a clear date. Drop it into the two detectors. If the results differ, it is not a brand dispute. It is a warning that the score is not the origin.
How to test the detector on text whose origin you know
The same familiar human text can be translated back and forth. Weber-Wulff and colleagues had a class of machine-translated texts in their set [5]. The translation changes the statistics, but it does not change the person who wrote the original text. If a flag appears after the translation, the school sees the difference between a change of surface and a change of authorship.
The third test is length. Both OpenAI and Turnitin linked their limits to reliability on sufficiently long text; Turnitin states a minimum of 300 words [1][11]. A short paragraph, abstract, or test answer cannot be read as a thesis. It is a different unit of measurement.
Three tests before punishment
The first test is a mismatch between the two tools on known human text. The second is the change in score after translating known human text. The third is the length below the threshold of the instrument. Neither test says who cheated. All three tell whether the score is capable of carrying the weight the institution places on it.
Before disciplinary use
Before punishment must come questions that are not just about the tool. Are there ongoing versions? Can the student explain literature? Does the style fit with the previous work without making it pseudo-science? Was the use of AI in the assignment allowed, restricted or prohibited? Was the rule clear before submission? If the answer to these questions is missing, the percentage will not fill it in.
A disagreement between two instruments on a known text is not a trademark dispute. It is proof that the score is not the origin.
— Jiný Kontext
16. Do not extrapolate the English FPR in Czech
Czech is not just English with diacritics. It has a different inflection, word order, phrases, genre habits and school style. If the detector is trained and tested mainly on English texts, its numbers cannot simply be transferred to a Czech term paper. RAID-extra shows the existence of the Czech news domain in the extended test [7]. This is useful. It is not a validation of Czech university theses.
Liang's study measures non-native English speakers [8]. Weber-Wulff measures English documents in a small set [5]. OpenAI Classifier measured the English challenge set [1]. Turnitin reports its own operational and internal test numbers [10][11]. Each of these sources has value. None of them say: "Here is a reliable false positive rate for Czech diploma theses at the Charles University and CTU in 2025 to 2026."
What would weaken the main thesis? A public, pre-registered test on Czech seminar and diploma texts with a calibrated false positive rate of no more than 1% and a simultaneously stated true positive rate. A test that survives translation and normal human editing. Deployed publicly verifiable watermark for main chatbots. Or the new policy of Czech universities, which would explicitly elevate the detector to evidence. As of the date of the search, there is no such support in the file.
As long as it is missing, it is more accurate to say less. The detector can be an orientation aid. He is not the author of the text. He is not a witness. It's not a proof machine.
17. A hidden cargo is a false alarm that the convict cannot bear
In public debate, the harm of minute AI text is often calculated. A student cheated and the school didn't catch him. This damage exists. Academic integrity is not decoration. But the second damage is less visible: a false alarm. The student wrote the text himself, or used the tool in a permitted way, and the system marked him as suspicious. The teacher begins to look differently. The Commission wants an explanation. Reputation will be damaged before the error is corrected.
This is a hidden account of a comfortable percentage. The software returns a number per second. The institution then has to bear the consequence for weeks or months. If it doesn't have a process that catches the false alarm, the detector is not cheap. It just shifts the prize to the person with less power.
It does not follow that schools should resign. It follows from the fact that they should name more precisely what they are punishing. Unacknowledged acceptance of output, failure to follow the rules for citation of the instrument, failure to defend the work, or misleading statements may be prohibited. These are all human and procedurally verifiable questions. The detector can give a stimulus. He doesn't have the last word.
So the question is not how to recognise text written by artificial intelligence by a single number. It is: what are you allowed to believe from the similarity statistics, for what population of text, with what threshold, and what error will you tolerate for this person? If the answer is "we don't know", that's annoying. But it is more accurate than the certainty that the instrument has never measured.
Related texts in this series
- Deepfake: How to spot a fake video or photo? — the image detector has the same problem as the text detector: the lab is not traffic.
- Is the photo proof? — the authenticity of the file is not the completeness of the plot.
