Is Student Work Used to Train AI Models? What to Establish Before You Sign

Short answer: it depends on the contract, and the question has to be asked twice — because retention in a matching corpus and use as training data are separate permissions. Many institutions have granted the first, believe they have refused the second, and have nothing in writing that establishes either. The answer lives in the data processing agreement, not on a trust page.

This is the question a data protection officer will ask you and the question a student union will ask you, usually in that order and usually after signature. It is worth being able to answer it from a document.

Why is this two questions rather than one?

Because a submission can be kept for two quite different purposes.

Corpus retention means the vendor stores the submitted document so that future submissions can be matched against it. This is not incidental to similarity screening — it is the mechanism. Turnitin’s own description of Turnitin Similarity states that it compares student work against “a vast collection of student submissions, premium publications, and 20+ years of internet content”, and a student-facing similarity report can list a match type of “Student Paper” alongside internet sources. The retained corpus is the product.

Training use means the text contributes to improving a model. It leaves no trace in any report, it is not visible to the institution, and it is not reversible once done.

An institution can rationally accept the first and refuse the second. What it cannot do is assume that refusing one refuses the other, because the contractual language covering them is different and they usually appear in different clauses.

The consumer market makes the distinction visible in a way the institutional market does not. A free web citation and checking service, Cite This For Me, states beside its paper checker that “the papers you upload will be added to our plagiarism database and will be used internally to improve plagiarism results.” That is a single sentence doing both jobs at once — retention and internal improvement — offered to individual students with no institutional oversight at all. Your students may already be uploading unpublished thesis chapters into services on those terms, which is a separate exposure from anything in your contract and is covered in the AI use your institution cannot see.

Abstract diagram of where submitted work travels between systems
The question is not “is it secure”. It is “where does it go, for how long, and for what purposes”.

Who actually holds the rights you are granting?

This is where institutional intuition is most often wrong.

Copyright. In most jurisdictions the student is the author of their thesis and holds copyright in it. The institution holds whatever licence its deposit terms grant — typically a non-exclusive right to retain, preserve and make available. When the institution then signs a platform contract permitting a third party to retain or process that text, it is granting rights it does not own outright, on the strength of a licence chain that begins with the student. If any link in that chain does not cover the purpose, the grant is unsupported.

US: FERPA rights sit with the student. The US Department of Education states that “when a student turns 18 years old, or enters a postsecondary institution at any age, the rights under FERPA transfer from the parents to the student (‘eligible student’)”. The statute is 20 U.S.C. § 1232g and the regulations are 34 CFR Part 99. The practical consequence for procurement is that consent architectures copied from K-12 vendors — parental consent flows, district-level agreements — are wrong by construction in higher education. The rights holder is the person submitting the work.

EU/UK: two legal questions, not one. Copyright in the thesis and personal data within it are governed separately. A submission can be lawful to retain as a copyright matter and unlawful to process for a new purpose as a data protection matter, or the reverse. Purpose limitation is the clause that bites: a purpose established for similarity screening does not extend to model training simply because the data is already held.

What are the four clauses to establish?

# What to establish What a sufficient answer looks like
1 Is submitted work retained in a corpus? Yes or no, with the retention period, the scope of matching, and whether other customers’ submissions are matched against ours
2 Is submitted work used to train or improve models? A separate answer from clause 1, naming any exception such as “aggregated and de-identified” — and defining it
3 Can a student’s work be removed? A named process, a stated timescale, and what remains after removal
4 What happens at termination? Deletion or return, in what format, by when, and whether the corpus copy is included

Two drafting traps are worth naming, because both are common and neither is malicious.

“Aggregated and de-identified” is not self-defining. A thesis chapter with the author’s name stripped is still the thesis chapter. If the exception is drafted around identifiers rather than around the text itself, it may permit exactly what you meant to prohibit. Ask what is aggregated, and at what granularity.

“We do not sell your data” answers a question you did not ask. Training is not selling. A clean answer on sale tells you nothing about clause 2.

Deposited theses whose copyright remains with their authors
The institution deposits the thesis. The author still owns it.

Where should the answers come from?

From the data processing agreement and the sub-processor list, in writing, signed. Not from a trust page, a security whitepaper, a sales deck or a support article — all four of which change without notice and none of which is a contractual commitment.

Three practical rules:

  1. Ask in writing and keep the reply. A vendor’s willingness to answer clause 2 in writing during a sales cycle is the best available signal of how it will behave during the contract.
  2. Ask about sub-processors by name. If a model sits behind an API operated by a third party, the answer to clause 2 depends on that third party’s terms, not on your counterparty’s intentions.
  3. Date the answer. Terms change at renewal. An answer without a date is not evidence in eighteen months.

The full sequence for running this as a review, with the artefacts each step produces, is in our guide to running a data protection review before deploying an AI writing tool, and the written question set to send is in our procurement question bank.

What should we tell students?

Something specific, and before submission rather than after. A notice that says “your work may be processed by third-party services” describes almost every possible arrangement and therefore informs nobody.

A usable notice names the service, states whether the submission is retained and for how long, states whether it is used to improve or train anything, and gives the route to object or to request removal. Writing that notice is also a good test of your own contract: if you cannot draft four honest sentences from the agreement you signed, the agreement is not specific enough.

It is worth being straightforward about the asymmetry here. The institution chooses the platform; the student supplies the text and holds the copyright in it. That asymmetry is the reason this question generates disproportionate reputational risk relative to its technical complexity, and the reason a clear answer is worth obtaining before anyone asks for it.

How does this interact with your AI policy?

Directly. A policy that requires students to disclose AI use while the institution has not established what happens to their submitted text is asking for a transparency it does not itself provide. The components that keep the two aligned are set out in what a university AI policy should include, and the evidentiary distinctions that policy has to make are in our note on how plagiarism detection and AI detection differ.

If you would like these four clauses answered in writing for an evaluation, request an institutional evaluation.

Frequently asked questions

Is student work used to train AI models?

It depends on the contract, and it is a separate question from whether the work is retained in a matching corpus. Establish both in the data processing agreement.

Is corpus retention the same as training?

No. Retention supports matching future submissions against past ones. Training changes a model. A vendor can do one and not the other.

Do similarity tools keep student submissions?

Retention of student submissions is how student-to-student matching works, and vendor descriptions of the corpus name student submissions explicitly. The terms and retention period are contract questions.

Who owns a student thesis?

Generally the student, as author. The institution holds a licence under its deposit terms, and any platform holds whatever the contract grants.

Who holds FERPA rights at a university?

The student. Rights transfer from the parents when a student turns 18 or enters a postsecondary institution at any age.

Does GDPR consent solve this?

Rarely, and it is usually the wrong instrument in an institutional relationship where consent is difficult to make freely given. Purpose limitation and the processing terms do more work than a consent tick box.

What does “aggregated and de-identified” actually permit?

Whatever the contract defines it to permit. Ask for the definition, because stripping identifiers does not necessarily stop the text itself from being used.

Can a student have their work removed from a corpus?

Only if the contract provides a route. Ask for the process, the timescale, and what remains afterwards.

What happens to the corpus copy if we terminate?

That is clause 4, and it is frequently silent. Silence should be read as retention.

Is “we do not sell your data” a sufficient answer?

No. Training is not selling, and an answer about sale does not address training.

Should students be told before or after submission?

Before, in specific terms naming the service, the retention, the purposes and the objection route.

What if the vendor will not answer in writing?

Record that as the answer. Responsiveness in a sales cycle is the best predictor of responsiveness in a contract.

Bring Tesify to your institution

Scope a departmental pilot: one cohort, one term, and your own measures of what worked.

Request an evaluation We reply within 2 business days

Categories