With three mechanisms, in this order: criteria specific enough that two independent markers apply them identically, calibration before marking rather than moderation after it, and a sampled second read with a documented reconciliation route. Double-marking everything is the expensive answer, and it is not the most effective one.
Why is extended written work harder to mark consistently than anything else?
Because it is a single long artefact judged holistically, and because the thing being judged is not one construct but three or four bundled together: the quality of the argument, the handling of evidence, the presentation of the prose, and originality. Each marker weights that bundle slightly differently, and nothing in a typical criterion grid forces them not to.
An examination with discrete questions constrains disagreement to each answer. A dissertation offers no such constraint. A marker forms an impression in the first ten pages and reads the remaining eighty through it, which is a well-known feature of holistic judgement rather than a failing of any individual.

What actually causes two markers to land a grade band apart?
Four causes, in rough order of frequency. Only the last one is fixed by second marking.
- They are weighting different constructs. One marker is answering is the argument sound and the other is answering is this a well-made document. Both are legitimate readings of a vague criterion set, and the resulting marks can differ by a full band while both markers are behaving correctly.
- They hold different implicit standards of written English. One treats intelligibility as the bar; another treats fluency as the bar. Unstated, this becomes an equality problem rather than a marking problem, because it lands unevenly across cohorts.
- They have different tolerance for claim strength. One marker reads an unhedged conclusion as confidence, another reads it as overclaiming. Neither is wrong; the criteria simply never said which reading applies.
- They applied the same criterion to the same evidence and disagreed. This is genuine marker variance, it is the smallest of the four, and it is the only one a second read reliably catches.
The practical consequence is uncomfortable: the most common causes of inconsistency are upstream of marking, so a department that responds to inconsistency by adding second markers is buying the least effective remedy at the highest price.
Which mechanism should we actually use?
| Mechanism | Staff cost | What it fixes | Evidence it leaves for an appeal |
|---|---|---|---|
| Criteria specification | One-off drafting | Causes 1, 2 and 3 | The published criteria themselves |
| Calibration before marking | Two hours per cohort | Causes 1 and 3, and partly 4 | Attendance record and the agreed exemplars |
| Sampled second read plus reconciliation | Moderate, proportional to sample | Cause 4, and detects drift | Strongest: two records and a reasoned reconciliation |
| Blanket double marking | Highest | Cause 4 only | Two records, often without reconciliation reasoning |
| Automated writing evaluation | Licence cost | Surface features only | A score with no reasoning behind it |
The last row deserves a direct answer, because it is the one procurement asks about. An automated writing evaluation tool measures surface features reliably and says nothing about whether an argument holds. It produces a number with nothing behind it to examine, which is the same structural problem that makes detector output weak as evidence — set out in whether AI detection is reliable enough to base a case on. Use it as a workload aid for feedback on surface features if you wish. Do not let it near a summative mark on extended work.
What does calibration before marking actually involve?
Two hours, once per cohort, and it is the highest-return item on this page.

- Select three scripts from a previous cohort, chosen to sit at the boundaries rather than in the middle of bands.
- Every marker reads and marks all three independently, in advance.
- The group meets and compares. The purpose is not to agree marks but to surface why the marks differ, which almost always turns out to be cause 1 or cause 3.
- The group writes one paragraph per script explaining the agreed mark. These become the exemplars, and they are the artefact that makes the exercise reusable.
- Anything the discussion revealed that the criteria do not say goes into the criteria for next cycle.
Step 5 is the one departments skip, and it is the reason some run the same calibration argument every year.
What should the criteria say about the writing itself?
Only what a marker can apply identically twice. That test rules out most of what people want to write down, and the discipline it imposes is the point. The full drafting method, with template clauses, is in how to write a departmental writing standard for theses.
The short version for a criterion grid:
- Workable: the argument is stated and can be followed; claims are proportionate to the evidence presented; sources are attributed consistently in the named style; structure follows the published chapter expectations.
- Not workable: the writing is elegant, engaging, fluent, or of a high standard of English. Every one of those invites a different reading from each marker.
Note the second workable item. Claims are proportionate to the evidence presented is a markable version of the hedging question, because both markers are looking at the same two things and comparing them. Appropriately cautious is not markable, because caution has no referent.
Can we mark written English separately from content?
You can, and if you do, three conditions apply.
- It must be a named, weighted criterion published in advance, not a silent adjustment applied inside a holistic judgement.
- It must be capped at a weight the department can defend, because at a high weight it becomes a proxy measure of prior educational advantage rather than of the research.
- It must never be a route to a misconduct finding. Unusual prose is not evidence of anything, and treating it as such is exactly the reasoning failure examined in what evidence stands up when a student appeals an AI misconduct finding.
Monitor the outcome by cohort. If a language criterion produces systematically lower marks for one group, you now hold the evidence of a differential effect, and you are better off knowing.
What evidence does a moderation process have to leave behind?
Enough for someone who was not in the room to reconstruct the decision a year later. In practice, four items per moderated script:
- The first marker’s mark and written reasoning against each criterion.
- The second reader’s mark and reasoning, recorded before seeing the first.
- The reconciled outcome, with the reason it landed where it did.
- Who reconciled it, and on what date.
Blind second marking matters here. A second reader who sees the first mark anchors to it, which produces agreement without producing reliability — and agreement produced by anchoring will not survive an appeal that asks how the second mark was reached.

How large should the moderation sample be?
Set it by risk rather than by a flat percentage. A defensible sampling rule covers, at minimum: every script at a classification boundary, every fail, every script where the marker recorded uncertainty, a random draw across the remaining distribution, and every script from a first-time marker. That rule is auditable, which a flat ten per cent is not.
Does any of this apply to theses and vivas?
Partly, and the differences matter. A thesis is examined rather than marked, usually by two examiners of whom one is external, and the outcome is a recommendation rather than a number. The mechanisms that transfer are the criteria specification, the requirement for independent judgement recorded before discussion, and the written reasoning. Calibration does not transfer, because examiners are appointed per candidate.
What can drift instead is the corrections threshold: two examiner pairs in the same department can require materially different amounts of rework for equivalent theses, and almost nobody measures it. The measurement design for that kind of question is in how to measure whether a writing support intervention actually worked.
Can markers use AI tools to help mark written work?
Not in a personal account, and not as the source of the judgement. Pasting a student’s chapter into a consumer chatbot discloses personal data to a company the institution has no contract with, which is analysed in whether your staff can put student work into an AI tool. Separately from the data question, a mark has to be reasoned by the person who signs it, because a mark you cannot explain is a mark you cannot defend on appeal.
Where the upstream fix sits
Most marking inconsistency on extended work is inherited from inconsistency in what candidates were told to produce. A published standard, a chapter structure table and a consistent reference style remove a large share of the surface variation that markers then have to judge around.
Tesify for Institutions holds those conventions inside the drafting environment, so what arrives for marking varies on argument rather than on presentation. That does not replace calibration, and we would not claim it does. It reduces the amount of noise your moderation process has to work through.
Request an institutional evaluation and we will scope a free departmental pilot against a programme you name, with the marking-consistency measures agreed before it starts.
Frequently asked questions
Is double marking the best way to get consistency?
No. It is the most expensive mechanism and it addresses only genuine marker variance, which is the smallest of the four causes of disagreement. Criteria specification and pre-marking calibration address the larger causes at a fraction of the cost.
What is the difference between moderation and double marking?
Double marking means every script is marked twice. Moderation means a sample is second-read to confirm that standards are being applied consistently, with a documented route for reconciling differences. Moderation is the proportionate mechanism for most extended written work.
What is calibration and when does it happen?
Calibration happens before marking begins. Every marker independently marks the same three boundary scripts from a previous cohort, the group meets to surface why their marks differ, and the outcome is a written paragraph per script that becomes a reusable exemplar.
Why do two markers disagree on the same dissertation?
Most often because they are weighting different constructs, one reading for the soundness of the argument and the other for the quality of the document. Other frequent causes are different implicit standards of written English and different tolerance for the strength of claims.
Should the second marker see the first mark?
No. A second reader who sees the first mark anchors to it, producing agreement without producing reliability. Record the second judgement before the two are compared, then reconcile with written reasoning.
How big should the moderation sample be?
Set it by risk rather than a flat percentage: every classification boundary, every fail, every script where the marker recorded uncertainty, a random draw across the rest, and everything from a first-time marker. That rule is auditable in a way that a flat ten per cent is not.
Can we include written English as a marking criterion?
Yes, on three conditions: it is named, weighted and published in advance; the weight is low enough that it does not become a proxy for prior educational advantage; and it never provides a route to a misconduct finding. Monitor its effect by cohort.
Which writing criteria are actually markable?
Ones where two markers look at the same two things and compare them, such as claims are proportionate to the evidence presented, or sources are attributed consistently in the named style. Elegant, engaging and fluent are not markable, because each marker supplies a different referent.
Can automated writing evaluation help with marking consistency?
Only on surface features. It measures those reliably and says nothing about whether an argument holds, and it produces a score with no reasoning behind it to examine. It has a place in formative feedback and no place in a summative mark on extended work.
Does moderation apply to doctoral theses?
Partly. A thesis is examined rather than marked and examiners are appointed per candidate, so calibration does not transfer. What does transfer is criteria specification, independent judgement recorded before discussion, and written reasoning for the recommendation.
What records does an appeal need?
Enough for someone who was not present to reconstruct the decision a year later: the first mark and its reasoning against each criterion, the second reader’s independent mark and reasoning, the reconciled outcome with its reason, and who reconciled it on what date.
Can markers use an AI tool to help them mark?
Not through a personal account, which discloses student personal data to a company the institution has no contract with, and not as the source of the judgement. A mark has to be reasoned by the person who signs it, because a mark that cannot be explained cannot be defended.
