Book a demo or get in touch

Name and email are required. Your role and institution are optional.

You can also email info@studentvoice.ai.

Using machine learning for automated language analysis

By David Griffin

Updated Sep 07, 2026

A useful comparison between two collections of comments needs a description of their differences. Ruiqi Zhong, Charlie Snell, Dan Klein and Jacob Steinhardt explore an automated approach in Describing Differences between Text Distributions with Natural Language.

This review uses the May 2022 version of the paper. A GPT-3 model proposes descriptions from samples of each collection. A separate verifier, based on UnifiedQA, evaluates and ranks the candidates. The researchers fine-tune the components using synthetic examples screened by people.

What the 76% result means

The benchmark contains 54 binary text-classification tasks. For the strongest system, at least one of its five highest-ranked descriptions was judged to match or be close to a human reference description on 41 tasks: approximately 76%.

This is a description-matching result across tasks, not the percentage of individual texts classified correctly. It comes from the fine-tuned Davinci proposer with verifier-based ranking. The authors acknowledge ambiguity in language, incomplete descriptions and the small, manually evaluated benchmark. See the methods and evaluation.

What to test before using this with student feedback

As an editorial application, treat a proposed difference as something to investigate. Ask to see examples from both collections and examine whether sampling, question wording or the mix of respondents could explain it. A description such as “more comments mention assessment” needs a denominator and a clear definition of what was counted.

This paper does not establish that the method is faster or more accurate than manual analysis in higher education. Nor does a difference between collections establish why it occurred. For a survey comparison, specify those limits alongside the proposed interpretation and connect it to the original comments. Our NSS open-text methodology guide sets out questions about coverage and comparisons.

Correction, 7 September 2026: This summary now includes the verifier and explains the top-five basis of the 76% result. Unsupported claims about educational efficiency and unrestricted applications have been removed.

Request a walkthrough

Book a free Student Voice Analytics demo

See all-comment coverage, sector benchmarks, and reporting designed for OfS quality and NSS requirements.

  • All-comment coverage with HE-tuned taxonomy and sentiment.
  • Versioned outputs with TEF-ready reporting.
  • Benchmarks and BI-ready exports for boards and Senate.
Book a free demo Prefer email? info@studentvoice.ai

UK-hosted · No public LLM APIs · Same-day turnaround

Related Entries

The Student Voice Weekly

Research, regulation, and insight on student voice. Every Friday. Prefer audio? Listen to the podcast.

© Student Voice Systems Limited, All rights reserved.