Updated Sep 07, 2026
On 22 May 2026, Cambridge reported an OpRaise study of 761 undergraduate psychology essays from 125 volunteers at Cambridge, Nottingham and Manchester Metropolitan. It tested Claude Opus 4.6, GPT-5.4 and Gemini 3 Flash: the model versions used in that study, not a permanent assessment of every AI system. Cambridge announcement.
The essays covered 50 modules and 87 assignments from 2022–2025. Cambridge’s sample consisted of invigilated work, Manchester Metropolitan’s of coursework, while Nottingham’s was mixed. Differences between their results cannot therefore be read as an institution ranking or isolated institutional effect.
The report’s combined-model results matched human-assigned degree bands for 63% of Cambridge essays, 53% at Nottingham and 35% at Manchester Metropolitan. These are agreement figures against routine moderated grades, not independently established ground truth.
Models tended towards the middle of the grade distribution and were more sensitive than human marks to measured linguistic features. However, repeat marking by the same model was highly consistent. Repeatability and agreement with human judgement are different properties.
Best-performing prompts were selected on a 20% calibration subset of 153 essays, then applied to the full corpus, including that subset. The headline results are therefore not wholly independent hold-out estimates. The study also covers one discipline and volunteered submissions; it does not establish performance across all university assessment. OpRaise report, results and methodology.
Cambridge describes staff and student concerns about trust and human relationships, alongside possible uses such as identifying work needing further review. These suggestions are not proof that an AI-assisted workflow improves grades, fairness or learning. The announcement argues against making these systems primary markers on this evidence. Cambridge account of findings and implications.
Our practical suggestion is to test the precise use being proposed. A consistency check, formative feedback aid and final marking decision have different requirements. Specify the model version, assessment type, comparison process and route for human review.
Student survey comments can inform an evaluation of perceived usefulness or trust. They cannot validate the accuracy of marks or replace assessment expertise. The governance checklist supports planning that comment review, while the grading system itself needs separate validation.
Review note, 7 September 2026: checked the announcement and report’s results and methods, distinguished repeatability from accuracy, added assessment-context and calibration limits, and removed claims that possible supporting uses were already validated.
Request a walkthrough
See all-comment coverage, sector benchmarks, and reporting designed for OfS quality and NSS requirements.
UK-hosted · No public LLM APIs · Same-day turnaround
Research, regulation, and insight on student voice. Every Friday. Prefer audio? Listen to the podcast.
© Student Voice Systems Limited, All rights reserved.