The Student Voice Weekly / Episode 25

Where your student comments actually go

14 August 2026 · 5 min 59 sec

This week, the episode discusses 83 students show what grades miss about learning. Student narratives reveal development that attainment measures cannot show.

Audio file: MP3 · 5.5 MB · direct download

Student Voice Weekly episode 25 artwork with Dr Stuart Grey

This week, Dr Stuart Grey steps away from the usual research and sector-news format to talk directly about data security and provenance.

Prompted by Student Voice's recent move to serve Irish institutions from an EU data centre, Stuart explains how regional residency and model provenance work together. UK customer data remains entirely in the UK, while Irish customer data is processed in the EU. Categorisation and sentiment analysis use deterministic machine-learning models; locally run LLMs are used only for summarisation.

In This Episode

  • Why student comments need to be treated as sensitive evidence rather than ordinary spreadsheet text.
  • Why UK customer data remains entirely resident in the UK.
  • What EU data residency provides for Irish institutions-and what it does not answer by itself.
  • Why categorisation and sentiment analysis use deterministic machine-learning models.
  • Why LLMs are used only for summarisation and run locally on controlled hardware.
  • What data provenance means in practical terms.
  • Why model, taxonomy, prompt, and processing versions matter for credible trends.
  • Five questions to ask before uploading real student data to an analysis tool.

Practical Resources

Practical Takeaway

Pick one student comment and map its full journey from upload to analysis, reporting, retention, and deletion. Any point that cannot be clearly explained deserves attention before real student data is processed.

About This Recording

This episode was recorded by Dr Stuart Grey. The transcript was prepared from the final recording and lightly corrected for names, technical terminology, and readability.

Subscribe

Subscribe to The Student Voice Weekly.

Transcript

Hi, and welcome to Student Voice Weekly. I'm Dr Stuart Grey, founder of Student Voice AI.

So I'm doing something a little bit different this week. Rather than talk through a research paper or piece of sector news, I'll talk about data security, and particularly data provenance.

The reason this has been on my mind is that we've recently started serving customers in Ireland from an EU data centre. And that sounds like quite a technical change, but it gets to a much bigger question universities should be asking of any system that handles student comments: where does the data actually go?

So student comments are not just anonymous lines in a spreadsheet. They can contain names, details about disability or mental health, accounts of difficult relationships, complaints about individual members of staff, or an experience that identifies somebody simply because the size of their cohort is very small.

So when a university uploads those comments for analysis, doing things properly really matters.

Data residency is part of that. For our Irish customers, the data is resident in the European Union and processed in our EU environment.

And, of course, for all our UK customers, data is entirely resident in the UK, and it is stored and processed in the UK. It does not move into the EU environment simply because we now support institutions in both places. They're very separate.

But the location of the server is only the start.

You also need to know who can access the data, what gets logged, how long individual files are kept, where backups sit, and what deletion actually means. You need to know whether the text quietly leaves that environment at any point in the process.

And that last question has become much more important with the advent of generative AI.

So it's very easy to take a file of comments, put them into a general-purpose AI tool like ChatGPT, Claude or Gemini, and get a plausible-looking summary back.

The practical problem is that those comments may then be passing through another company's systems. There may be different retention terms, subprocessors or model settings involved. And the person receiving the summary may have very little record of how it was produced.

So our process separates two different jobs.

All of the categorisation and sentiment analysis is done using deterministic machine-learning models, and LLMs are not used for either of those tasks. That means that the analytical foundation is stable and reproducible. With the same data and model version settings, you get exactly the same result.

We only use LLMs for summarisation, after the categories and sentiment analysis have already been produced. Those LLMs for summarisation run locally on our own hardware, inside a controlled processing environment. We do not send student comments out to public LLM services.

And local does not have to mean using an old, weak or outdated model. You can now run genuinely capable open-source language models on local hardware. So we get the value of a good summary without handing the underlying student data to an external AI provider.

The key thing is that deterministic models produce the analysis, while the local LLM helps communicate what that analysis says. These roles are deliberately separate.

So that brings me to data provenance.

Provenance is really just the ability to reconstruct what happened to a piece of data.

What was the original source file? Which version of the data was analysed? Was anything redacted? Which deterministic models and category definitions were used? What quality checks took place? If a summary was generated, which local LLM and prompt produced it? And can a claim in a report be traced back to the comments that support it?

That matters because a polished summary can look authoritative, even when nobody can explain its history.

If a report says students are concerned about assessment, support or the delivery of teaching, somebody should be able to look behind that statement. They should be able to see the relevant evidence in the verbatim comments, understand how it was all grouped together, and understand how that summary was generated.

The data is not the finish line here. But you can't have a useful conversation about the data if you don't have a route back to what the students actually said in those comments.

This also makes year-on-year work credible, because all of our output is entirely self-contained. Whenever a model changes, we reprocess everything and then regenerate reports. That means that you're not comparing different models for different years. You have one model across all years.

So if you're reviewing a student feedback tool or survey analysis tool, I would ask these five questions.

Firstly, where exactly will the raw comments be stored and processed?

Does any of the text go to a third-party AI service or another subprocessor?

Can you tell us which models, category definitions, summarisation prompts, and other processing rules were used for this particular run?

Can we trace a reported theme or summary back to the underlying evidence?

And when the work is finished, what is retained, for how long, and how is deletion verified?

You do not need every person in a student experience or surveys team to become a security engineer. But the institution should have clear answers to those questions. Those answers should be written down before real student data is uploaded.

For me, that's the difference between using AI because it is convenient versus using it responsibly.

So keeping UK customer data entirely in the UK, and Irish customer data in the EU, is one practical part of that responsibility. Using deterministic machine-learning models for categorisation and sentiment is another. And running summarisation LLMs locally means the comments remain inside that controlled environment.

Keeping the route from source comment to final conclusion visible is also what joins the whole thing together.

So the takeaway this week is simple. Think about one student comment and map its journey-its entire journey through the entire analysis pipeline. If you're not confident with any piece, ask your provider what's going on.

So thank you for listening to a slightly different episode of Student Voice Weekly. If you have any questions or comments, please get in touch. You can sign up for the newsletter and the podcast at studentvoice.ai.

Thank you very much.

The Student Voice Weekly

Research, regulation, and insight on student voice. Every Friday. Prefer audio? Listen to the podcast.

© Student Voice Systems Limited, All rights reserved.