<
Step 6: Decide the integration depth you actually need
Owner: IT lead. Artefact: an integration decision note.
Full single sign-on and VLE integration make a pilot feel institutional and add weeks of security review before anything can start. For a bounded pilot, a lighter arrangement is often defensible and much faster — provided you record explicitly that the pilot configuration is not the production configuration, so nobody treats a smooth pilot as evidence that integration will be smooth.
Step 7: Publish the rules to students before day one
Owner: head of department. Artefact: a paragraph in the assessment brief.
Students in the pilot must know what is permitted, what must be disclosed, and that participation carries no integrity risk if the rules are followed. Without this, cautious students opt out and your uptake data measures anxiety rather than usefulness. Align the wording with the permission categories in your institutional policy — the components are set out in what a university AI policy should include.
Step 8: Train the staff, and count the training
Owner: academic developer. Artefact: attendance record and a short competence check.
Under-trained staff produce a failed pilot that looks like a failed product. Record who was trained and who was not — the comparison between them is one of the more informative results a pilot generates. For EU institutions this also has a regulatory dimension, since Article 4 of Regulation (EU) 2024/1689 requires deployers to take measures ensuring a sufficient level of AI literacy among staff operating AI systems on their behalf.
Step 9: Instrument for the mid-point, not just the end
Owner: departmental administrator. Artefact: a mid-point note.
Schedule one checkpoint halfway. Its purpose is not to judge the tool but to catch the failure that invalidates the whole exercise — nobody has logged in, one seminar group never received access, the survey link was broken. A pilot that fails silently for six weeks yields nothing, and you will not get the term back.
Step 10: Repeat the surveys verbatim
Owner: departmental administrator. Artefact: matched pre/post datasets.
Same questions, same order, same scale. Add new questions at the end if you must; never edit an existing one. Report the response rate, and report it again in the decision paper — an enthusiastic result from a 12% response rate is a result about enthusiasts.
Step 11: Write the decision paper against the criteria sheet
Owner: sponsor. Artefact: a three-page paper.
Structure it as: the problem sentence; the baseline; what happened; each success criterion marked met or not met; the data protection position; cost at pilot scale and at institutional scale; and a recommendation with its main risk stated.
Three pages, tied to criteria agreed before the evidence existed, is a document a procurement committee and an auditor can both use. A slide deck of positive quotations is not.
A realistic term-length timeline
| Period | Activity | Owner |
|---|---|---|
| Term before | Problem statement, baseline capture, department selection | Sponsor, administrator |
| Weeks −6 to −1 | Criteria sheet signed; DPIA; vendor questions; integration note; staff training | Sponsor, DPO, IT, developer |
| Week 0 | Rules published to students; access issued | Head of department |
| Weeks 1–5 | Run; usage monitoring only | Administrator |
| Week 6 | Mid-point check | Administrator |
| Weeks 7–11 | Run | Administrator |
| Week 12 | Post surveys, repeated verbatim | Administrator |
| Weeks 13–14 | Decision paper and review meeting | Sponsor |
If a departmental pilot is the step you are trying to get through procurement, we can supply the pilot brief, the criteria template and the data protection documentation as a package. Request an institutional evaluation.
Frequently asked questions
How long should a pilot run?
One full academic term, so that the assessment cycle the tool is meant to affect actually occurs within the window.
How many students should be involved?
Enough that a change would be visible against your baseline, and concentrated in one department so that confounding factors are limited. Breadth across the institution is a rollout, not a pilot.
Do we need a DPIA for a pilot?
Assess it properly rather than assuming a pilot is exempt. The ICO publishes criteria for when a DPIA is required and examples of high-risk processing; note that its guidance is currently under review following the Data (Use and Access) Act.
Who should own the pilot?
An academic sponsor with authority over assessment, not IT. IT owns integration; the academic question is whether the intervention works.
What if the results are mixed?
That is the normal outcome and the criteria sheet is what makes it actionable — it tells you in advance which mixed results count as pass.
Should the vendor help run it?
They can supply training and support, but the evaluation and the survey instrument must be yours, or the result is not independent.
How do we prevent the pilot expanding mid-term?
Write the boundary into the criteria sheet. Uncontrolled expansion destroys the comparison with your baseline.
What is the most common reason pilots fail to inform a decision?
No baseline. Without it, every observation after launch is uninterpretable.
Should we measure integrity cases during the pilot?
Yes, but interpret them carefully — a single term in one department is a small denominator, and case counts are affected by reporting behaviour as much as by conduct.
How do we compare vendors during one pilot?
Generally you cannot, within one term and one department. Shortlist on the written question set first, then pilot the leading candidate — see our procurement question bank.
What should we do about detection tooling during the pilot?
Leave existing arrangements unchanged so you are testing one variable, and ensure no allegation rests on a detector score alone — the reasoning is in our analysis of AI detection reliability.
]]>
