<
Has the market itself moved on?
This is the part of the story that has changed most since 2023, and it is visible on vendors’ own pages rather than in commentary.
Turnitin’s current product overview leads with Turnitin Clarity, presented with the line “Don’t just detect AI. Understand it,” and described as bringing writing analytics and AI insights together for “the visibility needed to guide responsible AI use and empower authentic learning journeys.” The rest of the line-up is Feedback Studio, Gradescope, ExamSoft, Similarity and iThenticate.
Read that as a market signal. The framing has shifted from adjudicating whether a finished document was machine-written toward giving institutions visibility into how a document was produced. That is a meaningfully different product category, and it is a better fit for what integrity offices actually need — because process evidence is inspectable in a way a probability score is not.
One further change worth recording for anyone maintaining a vendor list: Ouriginal’s service ended on 30 June 2026, per Turnitin’s own page, which now directs former users toward Turnitin Similarity. Comparison content naming Ouriginal as a current option is out of date.
What does a defensible process look like?
The question an integrity office should answer is not “did AI write this,” which may be unanswerable, but “can this student account for this work.” That reframing is legally and pedagogically sturdier, and it does not depend on any vendor.
- Treat any detector output as a trigger, never a finding. Write that into the policy in those words, so a panel cannot treat a score as dispositive.
- Require corroboration before an allegation. Vanderbilt’s own guidance to instructors is a reasonable starting list: compare the work to the student’s previous writing for style, tone and level; look for inaccuracies in sources, arguments and facts, since generative tools may fabricate sources entirely; and talk to the student.
- Make the conversation the primary instrument. A short viva about the submitted work — why this method, where this source came from, what this paragraph means — distinguishes authorship far more reliably than a classifier, and it produces a record a panel can review.
- Never disclose a score as the accusation. “Our software says 82%” invites a fight about the software. “Please talk us through how you produced section 3” invites an account.
- Record the burden explicitly. The institution alleges; the institution proves. A policy that quietly shifts the burden onto a student to disprove a machine output will not survive scrutiny on appeal.
What about designing the problem out?
The durable answer is assessment design, and it is the one every serious source lands on. Vanderbilt recommended reformatting assessment — in-class writing, requiring students to write about specific material discussed in class, focusing on current issues — alongside clear expectations and citation of AI use where it is permitted.
That is not a counsel of despair about technology; it is a recognition that authorship is easiest to evidence when the process is visible. Assessments that generate intermediate artefacts — proposals, annotated bibliographies, drafts, supervision records — give both students and panels something concrete. An assessment whose only artefact is a finished document at a deadline will always be the hardest case.
If you are weighing how to give an institution that process visibility without adding a detection contract, we are happy to talk through what an evaluation would involve at your institution. Request an institutional evaluation.
Frequently asked questions
Can universities detect AI-generated essays?
Tools exist that estimate the likelihood text was machine-generated, but they produce probabilities rather than proof, and their accuracy claims are vendor claims that institutions should test against their own volume and cohort.
Can a detector score alone support a misconduct finding?
It should not. A score is not evidence of intent and provides no inspectable source. Use it to open an inquiry, and corroborate before alleging.
What false positive rate did Turnitin claim?
At launch of its AI detection tool, Turnitin claimed a 1% false positive rate, as recorded by Vanderbilt University in August 2023.
Why did Vanderbilt disable Turnitin’s AI detector?
It cited the false-positive risk at its submission volume, the absence of any explanation of how the tool works, evidence of bias against non-native English speakers, and privacy concerns about third-party detection.
How many false positives would we see?
Multiply your annual submissions by the vendor’s stated rate. At Vanderbilt’s 75,000 papers, a 1% rate implies roughly 750 papers a year.
Are detectors biased against international students?
Vanderbilt cited research finding detectors more likely to label non-native English speakers’ text as AI-written. Any institution with a large international cohort should treat this as a live equality issue rather than a technical footnote.
Is AI detection the same as plagiarism detection?
No. Similarity matching shows you a source document you can inspect; AI detection infers from the text’s own properties and produces no source.
Is Ouriginal still available?
No. Turnitin’s own page states that the Ouriginal service ended on 30 June 2026 and points former users to Turnitin Similarity.
Should we tell students whether we use detection?
Yes. Transparency about what is checked and how findings are used is a basic procedural fairness expectation, and undisclosed screening is difficult to defend on appeal.
What should replace detection in our policy?
A clear permitted-use and disclosure rule, corroboration requirements before any allegation, and assessment design that produces inspectable process evidence — the components we set out in what a university AI policy should include.
How should we pilot an alternative before committing?
Run a bounded departmental pilot with success criteria agreed in advance rather than an institution-wide rollout — the approach described in our guide to running a departmental pilot.
]]>
