How to Train Supervisors to Use and Govern AI Writing Tools (2026)
Most institutional AI rollouts train students and brief supervisors. That order is backwards. The supervisor is the person who will be asked, in a one-to-one meeting with no policy document open, whether a candidate can use a tool to restructure a literature review — and whatever they answer becomes the operative rule for that candidate for the next three years. If forty supervisors answer differently, you do not have a policy. You have forty policies, and an appeal waiting to happen.
This guide sets out nine steps to a supervisor training programme that actually changes practice. Each step is self-contained and executable on its own, names the owner, and produces a specific artefact you can put in front of a committee. Timings are given in academic terms, because faculty development schedules do not run in sprints.
Before you start: what this programme is not
It is not a tool demonstration. A vendor walkthrough teaches supervisors which button to press and leaves them exactly as uncertain about what to permit. It is also not an integrity briefing — supervisors have sat through those and they change nothing, because the difficulty is never whether misconduct is wrong. The difficulty is the boundary in a specific case: this candidate, this chapter, this tool, this week.
The programme below trains judgement on boundary cases and gives supervisors a defensible script for the conversation. It assumes your institution has an AI policy in force. If it does not, run the policy work first — the component list in what a university AI policy should include is the prerequisite, because you cannot train people to apply a rule that has not been written.
Step 1 — Establish the mandate and the completion expectation
Owner: Pro-Vice-Chancellor for Research or Dean of the Graduate School. Artefact: a one-page mandate note approved by the research degrees committee.
Decide, before anything is designed, whether completion is expected or required, and what “required” attaches to. In most institutions the enforceable lever is the supervisor approval register: a member of staff cannot be approved or reapproved to supervise research degrees without current training. That is a stronger mechanism than a mandatory-course email and it uses machinery you already run.
Template wording for the mandate note: “From the [autumn 2026] intake, approval to supervise research degrees requires completion of the AI in Supervision module within the preceding three years. Existing approved supervisors complete at first reapproval or by [date], whichever is earlier.”
Step 2 — Audit what supervisors currently believe
Owner: Graduate school academic lead. Artefact: a scenario survey report, anonymised, by faculty.
Do not survey attitudes. Survey rulings. Present eight to twelve short scenarios and ask for a permitted / not permitted / requires disclosure judgement on each. Useful scenarios include: a candidate uses a tool to translate their own draft from Spanish into English; a candidate uses one to generate a first outline of a discussion chapter; a candidate uses one to summarise twenty papers they have not read; a supervisor uses one to draft feedback on a chapter.
The output you want is the spread, not the average. A scenario where 60% of your supervisors say permitted and 40% say misconduct is a scenario your policy has failed to address, and it goes straight into the training as a worked case. This audit takes one academic month and it is the single most persuasive artefact you will produce, because it converts “we should probably train people” into a documented consistency risk.
Step 3 — Write the boundary casebook
Owner: Academic integrity office with two supervisor representatives per faculty. Artefact: a casebook of 15–20 scenarios with an agreed institutional ruling and a one-paragraph rationale for each.
This is the core deliverable of the whole programme. Take the scenarios where your survey found disagreement, put a small cross-faculty group in a room, and force a ruling on each. Where the group cannot agree, that is a finding to escalate to the policy owner rather than a gap to paper over in training.
The casebook must be discipline-aware. A ruling that works in history will be wrong in computer science, where code generation is a normal part of practice, and wrong again in a translation studies programme where the AI tool is the object of study. Write disciplinary carve-outs explicitly rather than leaving them to be inferred. The distinction that does most of the work here — that a thesis is governed differently from coursework — is set out in how AI rules should differ for a doctoral thesis compared with coursework.
Step 4 — Settle the staff-side rules before you teach the student-side ones
Owner: Data protection officer with the graduate school. Artefact: a one-page staff use note, cleared by the DPO, issued with the training.
Supervisors will ask what they themselves may do, and if you cannot answer, the rest of the training loses authority. The answer has a legal component that is not negotiable: putting a candidate’s draft into a personal consumer account discloses their personal data to a company the institution has no contract with. That is a controller-to-controller disclosure with no lawful basis, not a productivity shortcut, and the reasoning is set out fully in whether staff can put student work into an AI tool.
State three things in the note: which tools are institutionally provisioned and therefore covered by a data processing agreement; that no student work goes into anything else; and what the supervisor remains accountable for regardless of tooling, which is the academic judgement in the feedback.
Step 5 — Build the module in three parts, ninety minutes total
Owner: Academic staff development unit. Artefact: a delivered module with facilitator notes, deliverable in person or online.
Ninety minutes is the realistic ceiling for research-active staff, and it is enough if you spend it correctly:
- Twenty minutes — what the tools do and do not do. Not a feature tour. The single concept supervisors need is what the system is grounded on, because that determines whether a generated reference exists. Show one fabricated citation live; it lands harder than any slide.
- Fifty minutes — the casebook, worked in small groups. Give groups four scenarios, collect their rulings, then reveal the institutional ruling and the rationale. Disagreement in the room is the point of the exercise, not a failure of it.
- Twenty minutes — the conversation script. How to open the topic with a new candidate, what to record, and what to do on encountering suspected undisclosed use.
Step 6 — Give supervisors the disclosure conversation script
Owner: Graduate school. Artefact: a supervision agreement clause plus a two-sided prompt card.
The most valuable thing you can hand a supervisor is wording they can use verbatim in the first supervision meeting, because the conversation is awkward and awkward conversations get postponed indefinitely.
Template clause for the supervision agreement: “We have discussed the use of generative AI tools in this project. The candidate will use only institutionally provisioned tools for work on the thesis, will maintain a running record of substantive AI assistance, and will raise any intended new use with the supervisor before adopting it. The supervisor will review this record at each formal progress review.”
Two points make this work. It is a standing agreement rather than a single conversation, so it survives the candidate’s practice changing in year two. And it puts a review point in the progress review, which is a meeting that already happens and already produces a record.
Step 7 — Train the escalation path, not just the judgement
Owner: Academic integrity office. Artefact: a one-page escalation flowchart with named contacts and timescales.
A supervisor who suspects undisclosed AI use needs to know exactly what to do, and the most common institutional failure is that they do the wrong thing in good faith — running the draft through a detector themselves, or confronting the candidate with a score.
Train three rules explicitly. A detector output is a probability, not evidence, and it does not on its own support a finding. The supervisor’s role is to record concerns and refer, not to investigate. And the strongest available evidence is almost always process evidence the supervisor already holds — drafts, meeting notes, the candidate’s demonstrated command of their own material. What survives scrutiny at that point is covered in what evidence stands up when a student appeals an AI misconduct finding.
Step 8 — Deliver by faculty, not centrally, and schedule around the academic calendar
Owner: Faculty associate deans for research. Artefact: a delivery schedule with attendance tracked against the supervisor register.
Central sessions get low attendance and produce generic discussion. Faculty-level delivery lets you use discipline-relevant scenarios and lets the associate dean’s presence do the work that a mandatory-attendance email cannot.
Schedule realistically. Nothing lands in the final six weeks of a teaching term or during the examination period. The workable windows are the induction period at the start of the academic year and the quieter block after examination boards. Plan two full academic terms to cover an entire supervisor body of any size, and accept that the first cohort through will be volunteers who were already interested — that is normal, and their casebook feedback improves the module for everyone after them.
Step 9 — Measure whether it changed anything
Owner: Graduate school with institutional research. Artefact: a re-run scenario survey and a short evaluation report to the research degrees committee.
Satisfaction scores from a training session measure whether people enjoyed ninety minutes. They do not measure consistency, which is the entire objective. Re-run the Step 2 scenario survey six months after delivery to the same population and compare the spread of rulings, not the mean. Narrowing dispersion is the result you are looking for and it is straightforwardly reportable.
Three secondary indicators are worth tracking alongside it: the proportion of supervision agreements containing a completed AI clause; the number of referrals to the integrity office that arrive with process evidence attached rather than a detector score; and the number of policy queries escalated from supervisors, which should rise before it falls as people become confident enough to ask.
Where this sits in a wider rollout
Supervisor training is a component of an implementation programme, not a substitute for one. If you are at pilot stage, fold the casebook work into the pilot design so that the scenarios come from real cases in your own department — the sequence is set out in how to run a departmental pilot of an AI writing tool. If you have cleared the pilot and are scaling, supervisor training belongs in the same wave as provisioning, not after it, for the simple reason that supervisors will be asked about the tool the week students get access to it.
The two questions you will be asked in every session
“Am I now responsible for policing this?” No, and say so early and clearly. The supervisor’s duty is to set expectations, record what is agreed, and refer concerns. Investigation sits with the integrity office. Supervisors who believe they have been handed an enforcement role disengage from the training entirely.
“What if I do not use these tools myself?” Then they need the casebook more than anyone, because their candidates are using them regardless. Frame it as supervising a method the supervisor does not personally use — which is a familiar situation in any discipline where techniques move faster than the supervisory generation.
Provisioning that makes the training enforceable
Every rule in this programme depends on there being an institutionally provisioned tool to point candidates towards. A policy that permits AI assistance in principle while providing nothing in practice pushes candidates onto consumer accounts the institution cannot see, cannot audit and has no data agreement with — which leaves the supervisor’s disclosure conversation with nowhere to land.
Tesify for Institutions gives graduate schools a provisioned academic writing environment with SSO, institutional data agreements, and a per-candidate record of assistance that a supervisor can actually review at a progress meeting. If you are designing supervisor training this year, book an institutional demo or a free departmental pilot and build the module around a tool your candidates will genuinely have.
Frequently asked questions
Should supervisor AI training be mandatory?
Attach it to the supervisor approval register rather than issuing a mandatory-attendance instruction. Requiring current training for approval or reapproval to supervise research degrees uses governance machinery the institution already operates, gives a natural three-year refresh cycle, and avoids the compliance-email approach that faculty routinely ignore.
How long should the training module be?
Ninety minutes is the realistic ceiling for research-active staff and is sufficient if the time is spent on boundary cases rather than a tool demonstration. Allocate roughly twenty minutes to how the tools work, fifty to working the institutional casebook in small groups, and twenty to the disclosure conversation script.
What is a boundary casebook and why does it matter more than policy text?
It is a set of 15 to 20 realistic scenarios with an agreed institutional ruling and rationale for each. Policy text states principles; supervisors face specific cases. The casebook is what converts a principle into a consistent answer, and building it exposes the scenarios your policy never actually resolved.
Can supervisors use AI tools to draft feedback on student work?
Only within an institutionally provisioned tool covered by a data processing agreement, and never in a personal consumer account. Pasting a candidate’s draft into a private account discloses their personal data to a company with no contractual relationship to the institution. The academic judgement in the feedback remains the supervisor’s responsibility in all cases.
Should supervisors run drafts through AI detection tools themselves?
No, and the training should say so explicitly. A detector output is a probability with no examinable artefact behind it, and a supervisor confronting a candidate with a score creates a procedural problem for the institution. The supervisor’s role is to record concerns and refer them to the integrity office through the documented escalation path.
How do we handle disciplines where AI use is normal practice?
Write the disciplinary carve-outs into the casebook explicitly rather than leaving supervisors to infer them. Code generation in computer science, machine translation in a translation studies programme, and AI methods that are themselves the object of study all need named rulings, or faculties will simply conclude that the institutional policy does not apply to them.
How do we measure whether supervisor training worked?
Re-run the scenario survey six months after delivery and compare the dispersion of rulings, not satisfaction scores. Narrowing disagreement across the supervisor body is the objective. Track alongside it the share of supervision agreements with a completed AI clause and the proportion of integrity referrals arriving with process evidence rather than a detector score.
How long does it take to train an entire supervisor body?
Plan two full academic terms for a body of any size, delivered at faculty level rather than centrally. Viable windows are the induction period at the start of the year and the block after examination boards; nothing lands in the final weeks of a teaching term. Expect the first cohort to be self-selecting volunteers, and use their feedback to improve the casebook.
