Updated Sep 07, 2026
Before judging whether an algorithm is fair, specify what you want the judgement to mean. Equal outcomes, equal error rates and comparable probability estimates are different requirements.
Sahil Verma and Julia Rubin's 2018 paper, Fairness Definitions Explained, illustrates these distinctions using a logistic-regression classifier and the German Credit Dataset: 1,000 applicants described by 20 attributes. This is a structured credit-data example, not sentiment analysis of applicant descriptions or a study of students.
| Measure | What is compared across groups? |
|---|---|
| Statistical parity | The proportion receiving a positive prediction. |
| Predictive parity | The proportion actually positive among those predicted positive. |
| Equalised odds | Both false-positive and true-positive rates. |
| Conditional use accuracy equality | Both positive and negative predictive values. |
| Calibration | Among people assigned the same score, the proportion actually positive. Well-calibration additionally requires that proportion to equal the score. |
Their example meets some criteria and fails others. A claim of fairness therefore depends on the selected definition. The full paper explains additional measures and their assumptions.
Our editorial recommendation is to write down the decision being evaluated before selecting a metric. If a system flags comments for review, distinguish the proportion flagged from the proportion incorrectly flagged. Ask what a missed case means and what happens after a flag.
Then discuss whose experience the evaluation should represent, how the comparison groups are defined and whether the evaluation material is suitable for that purpose. Those choices need their own justification. A mathematical measure cannot establish that the chosen task or decision process is appropriate, and this credit example does not validate an educational application.
Correction, 7 September 2026: The previous summary confused conditional use accuracy with error-rate criteria and omitted the score condition from calibration. It also described a text-analysis application that was not part of the dataset or study. These errors and the associated FAQ claims have been removed.
Request a walkthrough
See all-comment coverage, sector benchmarks, and reporting designed for OfS quality and NSS requirements.
UK-hosted · No public LLM APIs · Same-day turnaround
Research, regulation, and insight on student voice. Every Friday. Prefer audio? Listen to the podcast.
© Student Voice Systems Limited, All rights reserved.