Updated Aug 14, 2026
An average evaluation score can show that something shifted. It cannot tell a lecturer what changed, why it changed, or what to do next. Lisa Levelt, Nouchka T. Tick and Maarten J. van der Smagt's Journal of Further and Higher Education paper, "Beyond numbers: the merit of routine student evaluation for starting academic teacher development", shows why universities need open comments and structured discussion alongside ratings. For UK teams using student voice to improve teaching, the message is practical: evaluation data becomes more useful when it supports reflection rather than ranking.
Student evaluations of teaching carry an awkward mix of purposes. Universities use them for quality assurance, staff development, and sometimes high-stakes performance decisions, even though ratings can reflect likeability, expected grades, gender, race, and other factors beyond teaching quality. Generic end-of-course questions can make the problem worse by mixing teaching skill, course design, and student experience into one score.
Levelt, Tick and van der Smagt examined a different approach. Early-career, teaching-only Psychology staff at a Dutch university designed the 17-item routine student evaluation of teaching, or 17-RSET, to inform their own development. The survey ran anonymously twice each year, midway through each semester. It combined 14 five-point rating items on teaching skills and learning experience with three open-ended questions. Results included individual ratings, averages, subscale scores, open comments, and graphs showing change over time. Supervisors were expected to discuss the evidence with teachers during annual development reviews. This bottom-up design complements evidence that teaching evaluation surveys improve when staff and students help shape them.
The study combined 5,050 student evaluations of 42 teachers, collected between 2015 and 2022, with interviews with 16 teachers. Each teacher had between one and eight evaluation occasions. The authors tested whether the rating scales were reliable and structurally valid, whether scores captured development over time, and how teachers understood and used the results.
Aggregated ratings were more defensible than isolated scores, but the measurement model was not perfect. The subscale scores aggregated across students showed acceptable reliability and structural validity for feedback purposes. However, high correlations between subscales suggested that a simpler structure might fit the data better. For university teams, the lesson is to validate the level at which results will be used and avoid treating every subscale label as a distinct truth.
Average scores changed very little across the whole group. The overall 17-RSET score did not rise significantly over time. One subscale increased by only 0.03 points per semester, while four individual items increased by 0.04 to 0.06 points. Changes differed significantly between teachers, and those with lower starting scores tended to change more, although consistently high ratings may have created a ceiling effect. A flat institutional average can therefore conceal meaningful individual trajectories.
Open-ended feedback supplied the context that ratings lacked. Twelve of the 16 interviewed teachers said the open questions gave them new insights, ten focused on those comments when interpreting their results, and six specifically valued the combination of comments and ratings. Comments helped explain whether a lower score reflected feedback on writing, lesson timing, approachability, or another concrete teaching behaviour. That is far more actionable than asking staff to infer a cause from a number.
One teacher described the wider value of the process:
"we show that we want to develop ourselves and that we are attentive to their voices"
Most teachers used the feedback, but use varied. Eleven teachers reported changing their teaching and three discussed results with students. Teachers rated the survey's insightfulness at 7.28 out of 10 and its general usefulness at 7.13. Yet the average self-rated impact on development was lower, at 5.66 out of 10, and teachers said the process mattered most early in their careers. Useful evidence can prompt action, but it does not prove that every score change was caused by the evaluation.
Developmental framing and support shaped how staff received the evidence. Thirteen teachers reported some tension around being evaluated, with average stress rated at 5.38 out of 10. Comparison with colleagues and possible summative use added anxiety. Six teachers said supervisors barely used the results, while six described constructive discussions that identified strengths and areas for development. The same survey can support reflection or provoke defensiveness depending on what the institution does after the results arrive.
UK universities should report ratings and comments as one evidence package. Show score distributions and trends, then add a concise account of recurring open-text themes with representative source comments available for checking. This prevents a small numerical movement from being mistaken for a diagnosis and gives teaching teams a concrete starting point for action.
Institutions should also compare staff with their own prior evidence before comparing them with colleagues. The study found substantial differences between individual trajectories, while colleague comparisons contributed to stress. Longitudinal self-comparison, accompanied by clear caveats about response numbers and ceiling effects, gives early-career staff a fairer view of development and keeps attention on improvement.
The post-survey process needs as much design as the questionnaire. Programme leaders or educational developers should help staff interpret repeated themes, distinguish teaching issues from course or institutional issues, and agree one or two changes to test. This aligns with evidence that student evaluations help teaching improve when staff can discuss them, giving teams a route from feedback to an accountable development plan.
Finally, universities should protect the formative purpose in policy and practice. Document who can see raw comments, how harmful language is handled, when evaluation data can inform performance decisions, and how analysis remains traceable. A student comment analysis governance checklist can help teams set those boundaries before sensitive feedback reaches staff. Clear governance makes the process safer for academics and more credible for students.
Q: How can a university apply these findings to module evaluations?
A: Keep a stable set of behaviour-focused rating items, add a small number of open questions, and collect feedback early enough for staff to respond. Report trends and recurring comment themes together, then build a facilitated review into the cycle. The benefit is a shorter path from student evidence to a specific teaching change.
Q: What methodological cautions matter when interpreting this study?
A: The study covers early-career, teaching-only Psychology staff at one Dutch university, so its findings should not be treated as a universal effect estimate. Anonymous responses could not be linked across semesters, high ratings may have limited measurable improvement, and the researchers were connected to the process they evaluated. UK teams should test the approach locally and validate any scale before using it for consequential decisions.
Q: What does the paper imply for student voice more broadly?
A: Student voice should not be reduced to a score or a comment dump. Ratings can show where to look, while comments explain what students experienced and structured dialogue helps staff decide what to change. Treating those elements as complementary makes feedback more likely to improve teaching and less likely to become a blunt performance measure.
[Paper Source]: Lisa Levelt, Nouchka T. Tick and Maarten J. van der Smagt "Beyond numbers: the merit of routine student evaluation for starting academic teacher development" DOI: 10.1080/0309877x.2025.2604705
Request a walkthrough
See all-comment coverage, sector benchmarks, and reporting designed for OfS quality and NSS requirements.
UK-hosted · No public LLM APIs · Same-day turnaround
Research, regulation, and insight on student voice. Every Friday. Prefer audio? Listen to the podcast.
© Student Voice Systems Limited, All rights reserved.