How to Run a Departmental Pilot of an AI Writing Tool (2026)

<![CDATA[

Most institutional pilots are demonstrations. A cohort is given a tool, everyone reports that it was well received, and a purchase follows. That process cannot produce a negative result, which means it produces no information — and it leaves the person signing the contract with nothing to show a finance committee eighteen months later.

A pilot worth running is one that could tell you not to buy. Here is how to structure one inside a single department and a single term, with the owner and the artefact named at each step.

Step 1: Write the problem statement before you name a product

Owner: the sponsoring associate dean. Artefact: one sentence.

The sentence must contain a number and no vendor name. “Our writing centre turns away roughly a third of postgraduate requests in weeks 6 to 10.” “Supervisors report spending the majority of first-draft feedback time on formatting and referencing rather than argument.” “Our integrity caseload rose by X cases year on year and Y% concern unattributed AI use.”

If you cannot write that sentence, stop. A pilot without a stated problem will be evaluated on enthusiasm, and enthusiasm is not a procurement criterion.

Step 2: Capture the baseline — you cannot recover it later

Owner: departmental administrator. Artefact: a one-page baseline record, dated.

This is the step pilots skip and the omission that makes them uninterpretable. Once the tool is live, the “before” state is gone. Capture, for the term preceding the pilot:

  • Volume measures — writing centre appointments requested versus delivered, waiting times, supervision meeting counts.
  • Caseload measures — integrity cases opened, upheld, appealed, and average days to resolution.
  • A short student survey using questions you will repeat verbatim afterwards.
  • A short staff survey, same rule.

On the surveys, one methodological warning worth heeding: change the wording between waves and your comparison is worthless. The reason national surveys can report that direct inclusion of AI-generated text in assessed work moved from 3% to 8% to 12% across three years is that the instrument stayed stable. Yours must too.

Step 3: Choose a department that can produce a real signal

Owner: sponsor. Artefact: a short rationale in the pilot brief.

Two failure modes bracket this choice. Pick your most enthusiastic early-adopter department and a positive result tells you nothing about the institution. Pick your most resistant and a negative result tells you nothing either.

Choose on structural criteria instead: enough students to see an effect, a coherent assessment pattern, a head of department who will attend the review, and at least one sceptic willing to participate honestly. The sceptic is not a risk to manage; they are your only source of disconfirming evidence.

Step 4: Agree success criteria — and stop criteria — in advance

Owner: sponsor with the head of department. Artefact: a signed one-page criteria sheet.

Write both columns before launch, and have them signed. The stop column is the one that makes this a pilot rather than a rollout with extra steps.

Adopt if… Stop if…
The baseline measure moves in the intended direction by an agreed margin The measure does not move, or moves in the wrong direction
Staff report time reallocated to substantive feedback Staff report net additional workload
Uptake is broad rather than concentrated in a keen minority Uptake stalls below an agreed threshold
No unresolved data protection or accessibility issue Any unresolved data protection or accessibility issue
Integrity cases do not rise attributably Integrity cases rise attributably

Step 5: Run the data protection review in parallel, not afterwards

Owner: data protection officer. Artefact: a completed DPIA or equivalent.

Start this the week you start recruiting participants. In the UK, the ICO publishes detailed guidance covering what a DPIA is, when one is needed, how to carry it out, whether you must consult the ICO, and examples of processing likely to result in high risk. Note for anyone working from an older internal template: the ICO states that its DPIA guidance is under review following changes made by the Data (Use and Access) Act and may be subject to change — so verify against the current page rather than a copy in your quality system.

US institutions have a different frame with one detail that regularly trips up procurement. Under FERPA — the statute is at 20 U.S.C. § 1232g — rights over education records transfer from parents to the student when the student turns 18 or enters a postsecondary institution at any age. In higher education the rights holder is therefore the student, whatever their age, and vendor arrangements must be built on that basis.

The full question set to put to the vendor is covered separately in our procurement question bank; treat the DPIA and the vendor questionnaire as one workstream.

Wall planner mapping a pilot across an academic term
One term, one department, and a decision date fixed before launch.

Step 6: Decide the integration depth you actually need

Owner: IT lead. Artefact: an integration decision note.

Full single sign-on and VLE integration make a pilot feel institutional and add weeks of security review before anything can start. For a bounded pilot, a lighter arrangement is often defensible and much faster — provided you record explicitly that the pilot configuration is not the production configuration, so nobody treats a smooth pilot as evidence that integration will be smooth.

Step 7: Publish the rules to students before day one

Owner: head of department. Artefact: a paragraph in the assessment brief.

Students in the pilot must know what is permitted, what must be disclosed, and that participation carries no integrity risk if the rules are followed. Without this, cautious students opt out and your uptake data measures anxiety rather than usefulness. Align the wording with the permission categories in your institutional policy — the components are set out in what a university AI policy should include.

Step 8: Train the staff, and count the training

Owner: academic developer. Artefact: attendance record and a short competence check.

Under-trained staff produce a failed pilot that looks like a failed product. Record who was trained and who was not — the comparison between them is one of the more informative results a pilot generates. For EU institutions this also has a regulatory dimension, since Article 4 of Regulation (EU) 2024/1689 requires deployers to take measures ensuring a sufficient level of AI literacy among staff operating AI systems on their behalf.

Step 9: Instrument for the mid-point, not just the end

Owner: departmental administrator. Artefact: a mid-point note.

Schedule one checkpoint halfway. Its purpose is not to judge the tool but to catch the failure that invalidates the whole exercise — nobody has logged in, one seminar group never received access, the survey link was broken. A pilot that fails silently for six weeks yields nothing, and you will not get the term back.

Step 10: Repeat the surveys verbatim

Owner: departmental administrator. Artefact: matched pre/post datasets.

Same questions, same order, same scale. Add new questions at the end if you must; never edit an existing one. Report the response rate, and report it again in the decision paper — an enthusiastic result from a 12% response rate is a result about enthusiasts.

Step 11: Write the decision paper against the criteria sheet

Owner: sponsor. Artefact: a three-page paper.

Structure it as: the problem sentence; the baseline; what happened; each success criterion marked met or not met; the data protection position; cost at pilot scale and at institutional scale; and a recommendation with its main risk stated.

Three pages, tied to criteria agreed before the evidence existed, is a document a procurement committee and an auditor can both use. A slide deck of positive quotations is not.

A realistic term-length timeline

Period Activity Owner
Term before Problem statement, baseline capture, department selection Sponsor, administrator
Weeks −6 to −1 Criteria sheet signed; DPIA; vendor questions; integration note; staff training Sponsor, DPO, IT, developer
Week 0 Rules published to students; access issued Head of department
Weeks 1–5 Run; usage monitoring only Administrator
Week 6 Mid-point check Administrator
Weeks 7–11 Run Administrator
Week 12 Post surveys, repeated verbatim Administrator
Weeks 13–14 Decision paper and review meeting Sponsor

If a departmental pilot is the step you are trying to get through procurement, we can supply the pilot brief, the criteria template and the data protection documentation as a package. Request an institutional evaluation.

Frequently asked questions

How long should a pilot run?

One full academic term, so that the assessment cycle the tool is meant to affect actually occurs within the window.

How many students should be involved?

Enough that a change would be visible against your baseline, and concentrated in one department so that confounding factors are limited. Breadth across the institution is a rollout, not a pilot.

Do we need a DPIA for a pilot?

Assess it properly rather than assuming a pilot is exempt. The ICO publishes criteria for when a DPIA is required and examples of high-risk processing; note that its guidance is currently under review following the Data (Use and Access) Act.

Who should own the pilot?

An academic sponsor with authority over assessment, not IT. IT owns integration; the academic question is whether the intervention works.

What if the results are mixed?

That is the normal outcome and the criteria sheet is what makes it actionable — it tells you in advance which mixed results count as pass.

Should the vendor help run it?

They can supply training and support, but the evaluation and the survey instrument must be yours, or the result is not independent.

How do we prevent the pilot expanding mid-term?

Write the boundary into the criteria sheet. Uncontrolled expansion destroys the comparison with your baseline.

What is the most common reason pilots fail to inform a decision?

No baseline. Without it, every observation after launch is uninterpretable.

Should we measure integrity cases during the pilot?

Yes, but interpret them carefully — a single term in one department is a small denominator, and case counts are affected by reporting behaviour as much as by conduct.

How do we compare vendors during one pilot?

Generally you cannot, within one term and one department. Shortlist on the written question set first, then pilot the leading candidate — see our procurement question bank.

What should we do about detection tooling during the pilot?

Leave existing arrangements unchanged so you are testing one variable, and ensure no allegation rests on a detector score alone — the reasoning is in our analysis of AI detection reliability.

]]>

Bring Tesify to your institution

Scope a departmental pilot: one cohort, one term, and your own measures of what worked.

Request an evaluation We reply within 2 business days

Categories