,

What Evidence Stands Up When a Student Appeals an AI Misconduct Finding? (2026)

What Evidence Stands Up When a Student Appeals an AI Misconduct Finding? (2026)

Process evidence survives appeal; detector output generally does not. Appeal bodies examine whether the finding rests on something the student could see and answer. Drafting history, version records, a documented authorship discussion, and the student’s demonstrated command of their own material all meet that test. A detection probability does not, because there is no underlying artefact for the student to contest.

This distinction decides most AI misconduct appeals, and institutions lose them for a narrow and repetitive set of reasons. This article sets out what an appeal body actually looks at, which evidence types hold and which collapse, the procedural defects that overturn a finding regardless of its merits, and what happens when a case leaves the institution for an external reviewer.

What is an appeal body actually reviewing?

Not, in most systems, whether the student used AI. Internal appeal and external review are usually confined to whether the institution followed its own procedure, whether the finding was one a reasonable panel could reach on the evidence before it, and whether the penalty was proportionate. That is a narrower question than guilt, and it is a question about your paperwork.

The practical consequence is that a case can be substantively correct and still be overturned. If the panel relied on material the student was never shown, or applied a standard of proof it never stated, or reached its conclusion before hearing the student, the finding fails on process and the underlying question is never reached. Institutions consistently underestimate how often this is the actual cause of a lost appeal.

Why does a detector score fail under scrutiny?

Because it is a probability with nothing examinable behind it. A similarity report names a source: the panel can open it, the student can explain it, and the match either exists or it does not. An AI-likelihood score names nothing. It is a number produced by a model whose reasoning cannot be inspected, and neither the panel nor the student can interrogate what generated it.

This is the operative difference between the two categories of tool, and it is examined at length in what separates plagiarism detection from AI detection. There is also a base-rate problem that panels routinely misread: even a low false-positive rate produces a substantial absolute number of wrongly flagged students once applied across an institutional submission volume, which is worked through in whether AI detection is reliable enough to base a misconduct case on.

A further exposure is worth stating plainly. Detection tools have documented elevated false-positive rates on text written by non-native English speakers. An institution whose findings are disproportionately drawn from its international cohort is carrying an equality risk in addition to an evidential one, and an external reviewer will notice the pattern before you do.

Which evidence types actually hold?

Ranked by how well they survive scrutiny, strongest first.

Evidence type What it shows Holds on appeal?
Drafting and version history How the document came into existence over time Strong — contemporaneous, student-visible, hard to reconstruct after the fact
Authorship discussion record The student’s command of their own argument, sources and choices Strong — the student participated and can respond
Supervision record What was agreed, what was submitted, what changed between meetings Strong — routine business record, not created for the case
Fabricated or unverifiable references Citations to sources that do not exist Strong — objectively checkable, and the student can be asked to produce them
Internal inconsistency in the text Content that contradicts the student’s own data or method Moderate — needs expert articulation, not assertion
Stylistic change from prior work A shift in register or fluency Weak alone — has innocent explanations, including support the institution provided
AI detection score A model’s probability estimate Weak — no examinable artefact; not sufficient on its own

Why is drafting history the strongest evidence available?

Because it is contemporaneous and it was not created for the purpose of the case. A document that developed over eleven weeks, with sections written out of order, passages deleted and restored, and reference entries added and corrected, has a shape that reads as composition. A document that appeared in three sittings with almost no revision has a different shape.

Two disciplines are required to use it properly. Neither pattern proves anything on its own — plenty of legitimate writers draft elsewhere and paste in, and a panel that treats a thin history as conclusive has simply substituted one unexaminable inference for another. And the student must be shown the record and given a genuine opportunity to explain it. Its evidential strength comes precisely from being contestable, so an institution that presents it without disclosure destroys the thing that made it useful.

What is an authorship discussion, and why does it survive appeal?

It is a structured conversation in which the student is asked about their own work: why this framework rather than an alternative, what a particular source actually argues, how a limitation in the method was handled, what a specific paragraph means. It is not an interrogation about whether they used a tool.

It holds up on appeal for three reasons. The student was present and participated. The questions are recorded and the answers can be reviewed by a body that was not in the room. And it tests the thing the institution actually cares about — whether the candidate can stand behind the intellectual content of the submission — rather than a proxy for it.

The failure modes are procedural. Convene it with notice, tell the student what it is for, allow the support or accompaniment your regulations provide, keep a record, and separate the person conducting it from the person who will decide the case. An authorship discussion sprung on a student in a corridor produces nothing usable and creates an additional ground of appeal.

What standard of proof applies?

Almost universally the balance of probabilities, not a criminal standard — but the standard has to be stated in the regulations and applied consistently, and this is where institutions get caught. Panels frequently apply an unstated higher standard when they are uncomfortable, and an unstated lower one when the case looks obvious. Either produces inconsistency an appeal body can identify.

Two related points belong in the regulations rather than in panel practice. The burden sits with the institution: a student cannot be required to prove they wrote their own work, and a finding that rests on the student’s failure to produce exculpatory material is defective. And where the regulations attach different standards or penalties to different levels of severity, the panel must record which it applied. The component checklist in what a university AI policy should include treats burden of proof as a mandatory element for exactly this reason.

Which procedural defects overturn findings most often?

  1. Undisclosed evidence. The panel saw a report, a detector output, or a staff email the student was never given. This is the single most common fatal defect.
  2. Predetermination. The person who raised the concern also decided the case, or the record shows the conclusion preceded the hearing.
  3. No stated standard of proof. The decision letter records a finding without recording the test applied.
  4. Inadequate reasons. A letter that asserts a conclusion without explaining which evidence was accepted, which was rejected, and why cannot be defended on review.
  5. Policy applied retrospectively. The submission predates the rule relied on, or predates the version of the rule quoted.
  6. Disproportionate penalty. A penalty outside the published range, or the same penalty applied to materially different cases without recorded reasoning.
  7. Unreasonable delay. A case that takes many months, particularly one that blocks progression or conferral, attracts scrutiny independent of the merits.

Note how few of these concern AI at all. They are ordinary administrative-fairness failures, and they surface more in AI cases only because the volume of AI cases is higher and the underlying evidence is weaker.

What does a defensible decision letter contain?

Five elements, and their absence is what an external reviewer looks for first. The allegation as put to the student, in the words used at the time. The evidence relied on, item by item, with confirmation that each was disclosed. The standard of proof applied, named explicitly. The findings of fact, with the panel’s reasoning on any point the student contested — including why the student’s explanation was not accepted. And the penalty, with the regulation it derives from and a note on proportionality.

A letter that says “the panel considered the evidence and found the allegation proven” is not a decision letter. It is a result, and on review it is close to indefensible because it gives no basis on which the finding could be assessed as reasonable.

What happens when a case leaves the institution?

Every one of the five markets has an external route, and each takes procedural regularity as its primary object of review.

  • England and Wales: the Office of the Independent Adjudicator for Higher Education reviews completed cases once internal procedures are exhausted and a Completion of Procedures letter is issued. Its published casework consistently emphasises evidence disclosure and adequacy of reasons.
  • United States: for public institutions, procedural due process obligations attach; for private institutions the framework is typically contractual, resting on whether the institution followed its own published procedure. In both, the published procedure is the standard you are held to.
  • Germany: examination decisions are administrative acts subject to Widerspruch and then to the administrative courts, with a well-developed Prüfungsrecht jurisprudence on procedural regularity and on the limits of reviewing academic judgement.
  • Netherlands: the College van Beroep voor de Examens hears examination appeals internally, with onward appeal to the CBHO route for higher education disputes.
  • Spain: internal reclamación procedures under the university’s own regulations, with public universities ultimately subject to the contencioso-administrativo jurisdiction.

The common thread is that none of these bodies is well placed to re-decide whether a student used a language model, and none tries to. They assess whether the institution ran a fair process and reached a conclusion open to it. That is why the evidential and procedural work has to be done in the first instance.

How should an institution reduce its appeal exposure?

Four measures, in order of effect. Remove detector scores as a standalone basis for a finding, in the regulations rather than in guidance. Make the authorship discussion the standard investigative step, with a documented protocol. Train supervisors on the escalation path so that referrals arrive with process evidence rather than a screenshot — the programme in how to train supervisors to use and govern AI writing tools covers the escalation module directly. And audit your decision letters annually against the five-element test above.

There is also a volume argument. Panel capacity is fixed by the number of academics willing to sit on panels, and a caseload built largely on weak-evidence referrals consumes that capacity while producing findings that do not survive review — the structural version of this problem is set out in why a growing integrity caseload cannot be met with more panel time. Reducing what arrives is a stronger lever than processing it faster.

Making the evidence exist before you need it

The strongest evidence in any of these cases is a record of how the work was made, and most institutions do not have one because the writing happened in consumer tools they cannot see.

Tesify for Institutions provides a provisioned academic writing environment where drafting history, reference verification and disclosed assistance are recorded as a routine part of the work — visible to the student throughout, and available to a supervisor at a progress review rather than reconstructed under investigation. It does not replace an integrity contract and it does not claim to detect anything; it makes the process evidence exist. If your appeal exposure is rising, request an institutional demo or a free departmental pilot and see what a term of real drafting records looks like.

Frequently asked questions

Can a university base an AI misconduct finding on a detection score alone?

It should not, and findings that rest solely on a detector output are the most vulnerable on appeal. A score is a probability with no examinable artefact behind it, so the student cannot meaningfully contest it and the panel cannot assess what produced it. Most vendors state that their output requires human review rather than constituting evidence.

What standard of proof applies in an academic misconduct case?

Almost always the balance of probabilities. The critical requirement is that the standard is stated in the regulations, named in the decision letter, and applied consistently across cases. Panels that silently apply a higher standard when uncomfortable and a lower one when the case looks obvious create the inconsistency an appeal body is designed to catch.

Does the burden of proof ever shift to the student?

No. The institution bears the burden of establishing the allegation. A student cannot be required to prove that they wrote their own work, and a finding grounded in a student’s failure to produce exculpatory drafts or notes is defective. The student’s inability to explain their own argument is evidence; the absence of records they were never asked to keep is not.

What is an authorship discussion?

A structured, recorded conversation in which a student is asked about the substance of their own submission — why a framework was chosen, what a cited source argues, how a methodological limitation was handled. It tests command of the material rather than tool use, and it survives appeal because the student participated and the record can be reviewed independently.

Is drafting history admissible in a misconduct case?

Yes, and it is generally the strongest evidence available, because it is contemporaneous and was not created for the case. It must be disclosed to the student with a genuine chance to explain it, and it should not be treated as conclusive on its own — legitimate writers sometimes draft in another application and paste the result in.

Which procedural failures most often overturn findings?

Evidence the panel saw but the student did not; the same person raising and deciding the case; no stated standard of proof; a decision letter that asserts a conclusion without reasons; retrospective application of a policy; and disproportionate or inconsistent penalties. Very few of these are specific to AI — they are ordinary administrative-fairness defects appearing at higher volume.

Are international students at greater risk of a wrongful AI finding?

Detection tools have documented elevated false-positive rates on text produced by non-native English speakers. An institution whose findings are disproportionately drawn from its international cohort carries an equality exposure alongside the evidential one, and the pattern is visible in case data before it becomes visible in a complaint.

What must a defensible decision letter contain?

Five elements: the allegation as put to the student; each item of evidence relied on with confirmation it was disclosed; the standard of proof named explicitly; findings of fact with reasoning on contested points, including why the student’s explanation was rejected; and the penalty with its regulatory basis. A letter stating only that the allegation was found proven cannot be defended on review.

Where does a student go after the internal appeal is exhausted?

In England and Wales, the Office of the Independent Adjudicator, following a Completion of Procedures letter. In Germany, Widerspruch and then the administrative courts. In the Netherlands, the College van Beroep voor de Examens and the onward CBHO route. In Spain, internal reclamación and then the contencioso-administrativo jurisdiction. In the US, due process claims against public institutions or contractual claims against private ones.

Should supervisors run their own detection checks before referring a case?

No. Ad hoc checks by individual staff produce material of no evidential value, frequently outside any institutional data agreement, and confronting a student with a score before referral creates a further procedural ground of appeal. The supervisor’s role is to record specific, articulable concerns and refer them through the documented escalation route.

How long should an AI misconduct case take?

Set and publish a target, and track against it. Unreasonable delay is a ground of complaint in its own right, independent of the merits, and it carries disproportionate consequences where a case blocks progression, conferral, a visa position or a job offer. Cases held open across a vacation period are a common source of upheld complaints.