What Is the Difference Between Plagiarism Detection and AI Detection?

Short answer: plagiarism detection is string matching — it returns the specific passages that match and the specific sources they match, which anyone can open and read. AI detection is statistical inference about authorship — it returns a likelihood with no underlying document behind it. One produces evidence a panel can inspect; the other produces an indicator that can only start an inquiry.

The distinction sounds academic until an appeal reaches a committee, at which point it becomes the whole case. Most institutional policies were written when only the first instrument existed, and they were then extended to cover the second by adding a sentence. That sentence is where the problems begin.

What does plagiarism detection actually do?

It compares a submission against an indexed corpus and reports overlaps. The output is a percentage plus, crucially, an itemised list: which passage, which source, what proportion of the document each source accounts for. Turnitin’s student-facing guidance describes the score as representing “the percentage of your writing that is similar to something found on the internet, in our databases, or in someone else’s paper”, and its Match Overview lists each matching source with its own percentage.

Three properties follow, and all three matter to an integrity office:

  • The output is inspectable. Every match points at a document. A reviewer can open it, read both texts, and form a judgement.
  • The output is contestable in specifics. A respondent can say “that passage is a block quotation, cited on page 14” and be right or wrong in a way that can be checked.
  • The percentage itself means very little. Turnitin states plainly that “Similarity does not mean that your work is plagiarized”, that educators should be considering acceptable forms of similarity “like quotations, citations, and bibliographic material” when reviewing a score, and that “there is no fixed number to receive as a score”. The University of Toronto’s teaching guidance puts it in institutional terms: Similarity Reports “do not indicate whether a student has plagiarized”, they identify sources containing textual similarities, and instructors must use their own judgement.

So the instrument does not detect plagiarism. It locates candidate passages and hands them to a human. That is not a weakness — it is the design, and it is why the output survives an appeal.

A matched passage that a reviewer can open and compare against its source
A similarity match points at a document. The whole evidentiary value sits in that pointer.

What does AI detection do differently?

It estimates the probability that text was machine-generated, using statistical features of the writing itself. There is no source document, because in the relevant sense there is no source — the model did not copy anything.

The consequences run in the opposite direction on all three properties above:

  • There is nothing to open. The output is the finding. A reviewer cannot go behind it.
  • It is contestable only in general terms. A respondent can assert that they wrote it, and can offer drafts and version history, but cannot address the specific reasoning that produced the score, because that reasoning is not expressed as a claim about any particular passage’s origin.
  • Error is asymmetric and invisible. A false similarity match reveals itself the moment someone opens the source. A false authorship inference does not reveal itself at all — it looks exactly like a true one.

That last point is the operative difference for an integrity office, and it is why the two outputs cannot carry the same weight in a procedure. The arithmetic of that exposure — annual submissions multiplied by a stated error rate, expressed as cases — is set out in our analysis of whether AI detection is reliable enough to base a misconduct case on.

How do the two compare side by side?

Property Similarity / plagiarism detection AI / authorship detection
Method Text matching against an indexed corpus Statistical inference from features of the writing
Output Percentage plus an itemised source list A likelihood, sometimes with highlighted spans
Underlying artefact Yes — the matched document No
Can a reviewer verify it independently? Yes, by reading both texts Not from the output alone
How a respondent rebuts it Passage by passage, with citations Only by external evidence such as drafts
Typical false positive Correctly quoted, correctly cited material Non-native phrasing, formulaic academic register, heavily edited prose
Does a false positive announce itself? Yes, on inspection No
Defensible policy role Evidence, reviewed by a human Trigger for inquiry, not a finding

Why does the vocabulary keep collapsing the two?

Because they are sold together and appear in the same interface. A single report can show a similarity percentage and an AI indicator adjacent to each other, in the same visual language, with the same colour treatment. Nothing on the screen tells a marker that one number points at a document and the other does not.

Vendor positioning has also moved. Turnitin’s current line-up spans Turnitin Clarity, Feedback Studio, Gradescope, ExamSoft, Similarity and iThenticate, and Clarity is led with “Don’t just detect AI. Understand it” — bringing writing analytics and AI insights together to provide “the visibility needed to guide responsible AI use”. Read as a signal rather than as marketing, that is the category acknowledging that a score about authorship does not answer the institution’s question, and adding a third thing: evidence about **process**.

Process evidence is a genuinely different category again. It does not infer authorship from the text; it records how the document came to exist. That makes it inspectable in the way similarity matching is, and contestable in specifics, without requiring a corpus match.

A misconduct panel room where evidence has to be produced and examined
The test of an instrument is not its accuracy claim. It is what a panel can do with its output.

What should the policy say differently about each?

Five drafting points that separate a defensible policy from an extended one.

  1. Name the instruments separately. If your regulations say “detection software”, they cover two things that need different treatment. Name similarity screening and authorship inference as distinct categories.
  2. State the evidential weight for each explicitly. A similarity report is evidence to be reviewed. An authorship indicator is a trigger for inquiry. Write both sentences.
  3. Say what a respondent may put forward. Where the allegation rests on authorship rather than matching, the respondent needs a route to demonstrate process — drafts, version history, a viva-style conversation. If the policy does not name that route, the procedure is unfair by construction.
  4. Prohibit a numeric threshold as a decision rule. There is no published threshold for either instrument, and adopting one internally converts a judgement into an automation you cannot defend.
  5. Require the marker to state which instrument they relied on. Cases collapse on appeal when nobody can say afterwards whether the concern started from a matched passage or from a score.

Point three is the one most often missing. It is also the cheapest to fix, and it converts the entire problem from an accuracy debate into a records question — which is the form in which it is actually solvable. The components a policy needs around these clauses are set out in what a university AI policy should include.

Does this change what we should buy?

It changes the question, at least. If the gap you have is “we cannot see how work was produced”, buying a better matcher does not close it, and neither does buying a better classifier. The three categories — similarity screening, authorship inference and process visibility — answer three different questions, and the market’s tendency to bundle them makes it easy to renew a contract that addresses the question you had five years ago. The renewal-side treatment is in our review of what changed in academic integrity platforms in 2026, and the vendor question set is in our procurement question bank.

If you would like to work through which of the three categories your institution is actually short of, request an institutional evaluation.

Frequently asked questions

What is the difference between plagiarism detection and AI detection?

Plagiarism detection matches text against a corpus and returns the matched passages and their sources. AI detection estimates the probability that text was machine-generated and returns no source document.

Can a similarity score prove plagiarism?

No. The vendor states that similarity does not mean work is plagiarised, and university guidance states that reports do not indicate whether a student has plagiarised and that judgement rests with the instructor.

Is there an acceptable similarity threshold?

None is published. Quotations, citations and bibliographies all register as matches, so the same percentage can describe good practice or poor practice.

Why can’t we treat an AI score like a similarity match?

Because there is nothing behind it for a reviewer to open or for a respondent to address. The two outputs are not interchangeable as evidence.

What is a typical false positive for each?

For similarity: correctly quoted and cited material, reference lists, standard methodological phrasing. For authorship inference: non-native phrasing, formulaic register and heavily edited prose.

Which false positive is more dangerous institutionally?

The authorship one, because it does not announce itself. A wrong similarity match is visible the moment someone opens the source.

What is “process visibility”?

Evidence of how a document came to exist — drafting history rather than a text comparison or a classifier score. It is inspectable and contestable in specifics, which is why it behaves like evidence rather than like an indicator.

Should our policy name specific products?

Name categories, not products. Products are withdrawn — one in the standard comparison set was discontinued in 2026 — and a policy tied to a product name ages badly.

Can a student be found in breach on a detection output alone?

That is a matter for your regulations, but a procedure that permits it on an authorship score alone is difficult to defend, because the respondent has nothing specific to answer.

What should a marker record when raising a concern?

Which instrument prompted it, what they read themselves, and what they asked the student. That record is what survives an appeal.

Do these tools work on non-English submissions?

Coverage varies by corpus and by language, and should be asked of the vendor in writing rather than assumed from a marketing page.

Where do we start if our policy conflates the two?

With one amendment naming the instruments separately and stating the evidential weight of each. It is a short change and it resolves most of the procedural exposure.

Bring Tesify to your institution

Scope a departmental pilot: one cohort, one term, and your own measures of what worked.

Request an evaluation We reply within 2 business days

Categories