iThenticate vs Turnitin Similarity for Thesis Screening: Which Corpus Are You Actually Buying?

The comparison table first, and it has one decisive column: corpus. Feature lists for these two products converge; their indexes do not. And for a thesis — a document that is simultaneously student work and pre-publication scholarship — the index is the whole question.

Criterion Turnitin Similarity iThenticate
Stated audience Institutions screening student writing; framed around “plagiarism and student collusion” “Publishers, researchers, and scholars”
Internet corpus “20+ years of internet content” “25+ years of internet content”
Student submissions in the corpus Yes — “a vast collection of student submissions” Not claimed on the product page
Subscription scholarly content “premium publications” “journal articles and subscription content sources from premier publishers”, including via a Crossref partnership
AI writing detection Present in the wider portfolio Claims identification of likely AI-generated content “even if modified by AI paraphrasing and bypassing tools”
Report emphasis Similarity score with colour-coded results Reducing false positives “by focusing on only critical similarity matches”
Natural fit in a graduate school Chapter drafts, coursework, pre-deposit screening Thesis-by-publication, staff outputs, pre-submission to journals

All quoted phrases are from the vendor’s own product pages, read in August 2026. Everything below follows from the third and fourth rows.

Why the corpus is the decision

A doctoral thesis has two overlap risks, and they live in different places.

Risk one: other students’ work. Reused coursework, a collaborator’s undeposited chapter, a purchased document previously sold to someone else, an earlier thesis from a partner institution. None of this appears on the open internet. It appears — if anywhere — in a corpus of student submissions.

Risk two: paywalled scholarship. A literature review built too closely on a subscription-only review article, or a methods section carried over from a paper the candidate co-authored. None of this is on the open internet either. It appears in a corpus of licensed journal content.

The vendor’s descriptions place these two corpora in different products. Turnitin Similarity names student submissions and 20+ years of internet content; iThenticate names 25+ years of internet content plus premium scholarly and subscription sources from publishers, and is positioned for publishers, researchers and scholars rather than for student writing.

So the honest statement is: a graduate school screening theses with one product has a defined blind spot, and it can say in advance which one it is. That is a far more useful procurement finding than a feature comparison, because it is actionable in the regulations rather than in the contract.

Two corpora that overlap but leave a gap neither one covers
The overlap is the open internet. The differences are where thesis-specific risk lives.

Which risk is larger for your programmes?

Four questions that settle it without a trial:

  1. Do your doctoral candidates publish before submission? Where thesis-by-publication is common, overlap with journal literature — including the candidate’s own published papers — is the live risk, and the scholarly corpus matters more.
  2. Do your master’s programmes share cohorts, datasets or industry partners? Where they do, student-to-student overlap is the live risk and the student-submission corpus matters more.
  3. Is your integrity caseload concentrated at taught-master’s level? In most institutions it is, which favours the student-facing product for volume screening.
  4. Do you screen at deposit, or during supervision? Deposit screening is a single high-stakes check; supervision screening is formative and repeated. They imply different licensing shapes.

Question four is the one procurement usually misses. A product licensed for end-of-process screening used formatively across three years of supervision is a different volume of submissions and often a different contract.

Where each product is genuinely weaker

Stated plainly, because a comparison that finds no weaknesses is marketing.

Turnitin Similarity, for thesis work: the corpus is optimised for student writing, which is exactly what a taught-master’s dissertation is and only partly what a doctoral thesis is. A candidate whose chapters draw on subscription-only literature is screened against a smaller relevant universe than the work sits in.

iThenticate, for thesis work: its product page does not claim a student-submission corpus. For contract cheating and cohort collusion — the fastest-growing categories in most integrity offices — that is the corpus that matters, and its absence is not compensated for by a longer internet index.

Both: neither addresses authorship. Similarity screening cannot see text that was generated rather than copied, which is a different instrument entirely and carries a different evidentiary weight — the distinction is set out in our note on how plagiarism detection and AI detection differ. iThenticate’s page does claim identification of likely AI-generated content “even if modified by AI paraphrasing and bypassing tools”, which is a strong claim; before it enters a business case, read our analysis of whether AI detection is reliable enough to base a misconduct case on and ask for the evidence in writing.

The screening points in a graduate school thesis workflow
Decide where in the workflow you screen before deciding what you screen with.

The recommendation

For most graduate schools: the student-facing product for volume screening, with the scholarly-corpus product reserved for thesis-by-publication candidates and staff research outputs. That matches the shape of the caseload, keeps the per-submission cost proportionate, and puts the expensive corpus where the publication risk actually is.

The alternative profile: a research-intensive institution whose doctoral programmes are almost entirely publication-based, and whose taught provision is small. There, the scholarly corpus is the primary need and the student-submission corpus is the secondary one — the inverse of the default.

What both profiles must do: write down which risk they have chosen not to screen for. An unscreened risk that has been identified and accepted is a governance decision. The same risk unidentified is an exposure that surfaces in an appeal, when someone asks why the check that would have caught it was not run.

Six questions to put in writing before signature

  1. Which corpora does this product match against, itemised, and are other customers’ student submissions included?
  2. Are our submissions added to a corpus, and separately, are they used to train or improve models? (Two questions — see why they are not the same permission.)
  3. What is the licensed volume, measured how, and what happens when a thesis is screened repeatedly during supervision?
  4. Which submissions count against the licence — drafts, resubmissions, appendices?
  5. What is the exit position: can we export reports, and what happens to our corpus contributions on termination?
  6. Is the AI-detection claim supported by evidence we can produce to a committee?

The wider written question set is in our procurement question bank, and the market-level changes to check at renewal — including a product withdrawn in 2026 — are in what changed in academic integrity platforms this year.

What screening does not solve

Both products examine work after it has been written. Neither tells a supervisor anything during the eighteen months when the thesis is actually being produced, which is when intervention is cheap and evidence is abundant. Screening and support are different purchases with different value, and an institution that has only the first has visibility only at the end.

To discuss where that gap sits at your institution, request an institutional evaluation.

Frequently asked questions

What is the difference between iThenticate and Turnitin Similarity?

Principally the corpus and the intended population. Similarity names student submissions plus 20+ years of internet content and is framed around student writing; iThenticate names 25+ years of internet content plus premium subscription scholarship and is positioned for publishers, researchers and scholars.

Which is better for screening a doctoral thesis?

Neither, on its own, covers both thesis risks. Choose according to whether your larger exposure is student-to-student overlap or overlap with paywalled literature, and document the choice.

Do we need both?

Usually not. A common split is the student-facing product for volume screening and the scholarly-corpus product for thesis-by-publication candidates and staff outputs.

Are they from the same vendor?

Yes. They are listed as separate products in the same portfolio, which is why “market testing” between them is not a comparison of independent suppliers.

Does either detect AI-generated text?

AI writing detection appears in the portfolio, and the iThenticate page claims identification of likely AI-generated content even when modified by paraphrasing and bypassing tools. Ask for the evidence in writing before relying on it.

Why does the internet corpus differ by five years?

The pages state 20+ and 25+ years respectively. For thesis work the difference matters far less than the presence or absence of the student-submission and subscription-journal corpora.

Can students see the report?

That depends on your configuration and your regulations rather than on the product. Decide it deliberately, because a pre-deposit report changes student behaviour.

Should we screen every thesis?

Most institutions screen at deposit. Whether to screen formatively during supervision is a separate decision with a different licensing shape.

Do appendices and datasets count?

Ask, in writing. Volume definitions differ, and appendix-heavy theses can consume licence allowance quickly.

What about non-English theses?

Corpus coverage varies by language and should be established for your actual languages of submission rather than assumed.

What is the single question that settles the choice?

Which unscreened overlap would embarrass you more: a chapter shared with another student, or a chapter shared with a paywalled article.

What should the regulations say?

Which corpus is screened, at which point in the process, and what the output is permitted to establish.

Bring Tesify to your institution

Scope a departmental pilot: one cohort, one term, and your own measures of what worked.

Request an evaluation We reply within 2 business days

Categories