How to Measure Whether a Writing Support Intervention Actually Worked (2026)

Most writing support evaluations report two things: how many people came, and how satisfied they were. Neither is an impact measure, and a finance committee that has been asked for a renewal will eventually say so.

This guide sets out how to produce a finding that holds. Seven steps, each with an owner and an artefact — and one methodological problem that has to be dealt with first, because everything else depends on it.

The problem underneath every evaluation of this kind

Students who use writing support are not like students who do not, and they were not like them beforehand.

Three reasons, all of them ordinary:

  • Attenders tend to be more engaged. The act of booking is itself evidence of investment in the work.
  • Attenders are often more worried about the piece, which can mean they were doing less well to begin with.
  • A share of them were told to come by a supervisor who already had a concern.

Notice that these push in opposite directions. The first inflates any apparent effect; the second and third deflate it. You cannot even predict the sign of the bias, let alone its size. That is why “attenders scored higher than non-attenders” is not a finding — it is a description of who chose to attend.

This is the same class of problem that makes sector-level figures hard to interpret, and it is why the numbers you inherit need reading carefully before they are reused — a habit we set out in where graduate completion and attrition data comes from.

Two parallel lines that begin together and separate over time
The only comparison worth making is between groups that started in the same place.

The seven steps

Step 1 — Write the question as a testable claim

Owner: the service lead with institutional research.
Artefact: a one-sentence claim naming the population, the intervention, the outcome and the window.

“Does the writing centre help?” cannot be answered. “Among first-year doctoral candidates in the faculty, does attending three or more structured sessions in term one reduce the proportion requiring major corrections at first submission?” can be. Write it before you look at any data, because a question written afterwards will be written to fit what you found.

Step 2 — Choose outcomes at two distances

Owner: institutional research.
Artefact: a short outcome table — measure, source system, distance, expected direction.

Pick one proximal outcome, close enough to the intervention that it can plausibly move: a rated feature of a submitted draft, revision behaviour between drafts, on-time submission of a chapter. Then one distal outcome the institution actually cares about: progression, time to submission, corrections category.

You need both. A proximal-only result invites the question of whether anything downstream changed. A distal-only result is too far from the intervention to attribute, and too slow to inform a decision this year.

Say in advance which direction each measure should move. A measure that can be read as a success whichever way it goes is not a measure.

Step 3 — Build a comparison group from a constraint you already have

Owner: the service lead with institutional research.
Artefact: a stated design with its assumption written down.

You will not be permitted to randomise access to help, and you should not ask. Use structure that already exists:

  • Waiting list. If demand exceeds capacity — and it usually does — the people waiting are a delayed-access comparison group, arguably the cleanest available. Nobody is denied support; some receive it later, which was going to happen anyway.
  • Staged rollout. If the service reaches departments in sequence, later departments are a comparison for earlier ones over the same period. This supports a difference-in-differences reading, whose assumption is that both groups would have moved in parallel without the intervention.
  • Eligibility threshold. Where access is triggered by a cut-off, students just either side of it are near-identical, which supports a regression discontinuity comparison. Powerful, and it only speaks about students near the threshold.
  • Matched comparison. The fallback. Match attenders to non-attenders on prior attainment, discipline, entry route, mode and language background. Weaker than the others, because you can only match on what you can observe, and motivation is not in the student record.

Whichever you pick, write its assumption in the protocol. Every one of these designs is valid only under a condition, and a committee is entitled to see it.

Step 4 — Calculate the minimum detectable effect before you start

Owner: institutional research.
Artefact: a one-line power statement in the protocol.

This step is skipped almost universally and it determines whether the exercise is worth running. Take the sample you will realistically have, and work out the smallest effect the study could detect. If a writing centre seeing 180 students a year can only detect a change large enough to be implausible, the evaluation will return a null result that says nothing about the intervention and everything about the sample.

Better to know that in week one. It usually points to pooling across cohorts, extending the window, or choosing a more sensitive proximal outcome instead of a rare distal one.

Step 5 — Clear the data route before you collect anything

Owner: the data protection office with the service lead.
Artefact: a documented lawful basis, a minimised field list and a retention period.

Impact evaluation means linking service attendance records to student outcome records, which is a new use of personal data and needs to be handled as one. Decide the lawful basis, limit the fields to those named in step 2, agree who holds the linked dataset and for how long, and separate identifiable data from the analysis file.

Run it through the same sequence you already use for deployments rather than inventing a parallel one — the method is in how to run a data protection review before deploying an AI writing tool. An evaluation that has to be paused for a retrospective review loses a cohort.

A booking diary and terminal at a university support service desk
Attendance records are the exposure variable. If they are incomplete, nothing downstream is recoverable.

Step 6 — Record dose, not just attendance

Owner: the service, at the point of delivery.
Artefact: a session log with date, duration, type and stage of work.

“Attended: yes” throws away most of the information. One drop-in session and eight structured appointments are different interventions, and collapsing them guarantees a diluted result.

Recording dose also gives you the most persuasive evidence available without randomisation: a dose-response relationship. If more sessions are associated with more change, in an orderly way, that pattern is much harder to explain by selection alone. It is not proof, and you should not present it as such, but it moves a committee.

Set this up before the intervention starts. Dose cannot be reconstructed afterwards.

Step 7 — Report the finding with its limitation attached

Owner: institutional research with the service lead.
Artefact: a two-page report — question, design, result, limitation, decision implication.

Four rules for the write-up:

  1. Report the effect size, not just significance. A statistically significant change too small to matter operationally is not a reason to fund anything.
  2. Report the null results. An evaluation reporting only what moved is a marketing document, and readers who notice will discount everything in it.
  3. State the design’s assumption in the body. Not a footnote. The assumption is the finding’s warranty.
  4. Separate the satisfaction data. Report it as service quality, clearly labelled, in its own section.

A report that names its own weaknesses is more persuasive than one that does not, because the reader stops looking for the ones you hid.

Printed result tables under review in a planning meeting
The finding a committee accepts is the one whose limits are stated before they are found.

Four measures to stop reporting as impact

  • Attendance counts. A utilisation measure. Useful for capacity planning, silent on effect.
  • Satisfaction scores. A service-quality measure. A demanding intervention can score poorly and work best.
  • Self-reported confidence. Real, but it moves for reasons unrelated to capability, and it is measured on people who chose to be there.
  • Testimonials. Genuinely valuable for communicating the service. Not evidence, and presenting them as evidence damages the case they support.

None of these should be dropped. They should be labelled correctly and reported in a different section from the impact finding.

What this changes about capacity decisions

The reason to do this properly is that the alternative argument is unwinnable. “Demand exceeds capacity” is a statement about demand, and the standard response is that demand for a free service always exceeds capacity. An effect estimate with a comparison group behind it answers a different and better question: what does the marginal appointment produce, and for whom?

That also identifies where support is worth concentrating, which is usually a more realistic outcome than an overall expansion — the underlying capacity problem is described in the writing centre waiting list you cannot staff your way out of.

If a new intervention is being introduced, evaluate it from the first cohort rather than after it is embedded. A bounded pilot with criteria agreed in advance is the natural vehicle, and its design is in how to run a departmental pilot.

If you would like to see what usage and progress data a supported writing environment produces for an evaluation of this kind, request an institutional evaluation and we will map it against the outcomes you have chosen.

Frequently asked questions

Why can we not simply compare attenders with non-attenders?

Because the groups differed before the intervention. Attenders are more engaged, more worried, or referred by a supervisor with a concern — and those factors push the result in opposite directions, so the bias cannot even be signed.

Is satisfaction an impact measure?

No. It measures experience. Report it, label it as service quality, and keep it out of the impact section.

What if we cannot randomise?

Use an existing constraint. A waiting list gives delayed access, a staged rollout gives difference-in-differences, an eligibility cut-off gives regression discontinuity. None requires withholding support.

Is a waiting list ethical to use as a comparison group?

Yes, where the wait already exists because demand exceeds capacity. You are describing what happens, not creating it. Do not lengthen a wait to improve a design.

Which outcomes should we choose?

One proximal outcome close enough to move, and one distal outcome the institution cares about. Name the expected direction of each in advance.

How many students do we need?

Enough to detect a plausible effect. Calculate the minimum detectable effect for your realistic sample before starting, and redesign if it is implausibly large.

Why record session dose?

Because collapsing one drop-in and eight appointments into “attended” dilutes any effect, and because a dose-response pattern is the most persuasive non-randomised evidence available.

Do we need a data protection review?

Yes. Linking attendance records to outcome records is a new use of personal data, with a lawful basis, a minimised field list and a retention period to settle first.

Should we report results that show no effect?

Always. Selective reporting is detectable and it discredits the results you do report.

How long should an evaluation run?

Long enough for the distal outcome to occur, which for doctoral work may be years. Report the proximal outcome at the end of the first cycle and the distal one when it arrives, rather than waiting silently.

Who should own the analysis?

Institutional research rather than the service being evaluated. Both should design it together; the service should not be the sole author of its own effect estimate.

What is the single most common mistake?

Choosing outcomes after seeing the data. It guarantees a positive finding and destroys the credibility of every subsequent evaluation the service produces.