Case study · Research → roadmap

A year of NPS free-text, turned into four shaped pitches

Student NPS was negative — −19.51% in Q1 2024. Thousands of open-text responses explained why, unreadably. I cleaned the year, tagged every comment, built the dashboard, and converted the four biggest themes into pitches carrying a combined +17.5 point expected gain.

My role
Product Manager — analysis, dashboard, PRD
Dataset
3,476 cleaned responses · 11 institutions
Output
4 shaped pitches for Design & Engineering
Expected gain
+17.5 points across 4 pitches
The problem

The answer was already in the data — nobody could read it

Student NPS and satisfaction feedback, eleven institutions across AU, NZ and the UK.

The score

Measurable. Tracked quarterly, reportable to the business.

The reasons

Thousands of free-text comments — too many to read, too messy to prioritise against.

Student NPS by quarter, 2024
Net Promoter Score = % promoters − % detractors. Zero is the line where promoters and detractors cancel out.
+15 0 −15 −19.51% +13.03% +24.27% +15.00% Q1 Q2 Q3 Q4 n=41 n=284 n=342 n=200
Full year 2024 closed at +16.38% across 867 exam-survey responses. Q1's small sample (n=41) makes it volatile — but it is the real reading, and it is the negative number the roadmap had to answer.
The problem

A score tells you whether something is wrong. Never what to build. All the value sat in the free text — and the free text was unusable raw.

Method

Clean it, tag it, then look at it

Hygiene first

Duplicates and blanks would have skewed every theme count. They came out before anything was tagged.

Collected
A full year

Student NPS and satisfaction responses across eleven institutions, split by assignment and exam.

Removed
Duplicates & blanks

Stripped before any counting, so theme volumes reflect distinct students rather than repeat submissions.

Analysed
3,476

The clean denominator every number in this case study is calculated against.

L1–L4

Four independent labels, not a hierarchy. One student can praise the interface, struggle with the learning curve, hit a device problem and complain about spellcheck in a single sentence. Forcing that into one category throws away three-quarters of the signal.

Formatting Improvements Positive Experience Learning Curve Easy To Use Prefer Word Spelling Device Bad UI Referencing One word response

"One word response" exists deliberately — a bucket for noise, so answers like "good" could be excluded from theme counts rather than quietly inflating them.

Spreadsheet showing student free-text responses tagged across four independent label columns L1 to L4
1 2 3
1L1–L4 as parallel slots. Most rows use one; the ones that matter most use three or four.
2The raw free-text response, left untouched — the label is added beside it, never instead of it, so any theme count can be audited back to the sentence that produced it.
3This row alone carries Learning Curve, Device, Spelling and Bad UI — four separate product problems in one comment. Single-label tagging would have counted it once and lost three.

Live 2024 response data. Submission tokens redacted; responses are anonymous by collection.

I started tagging manually. Accurate, and far too slow — the responses were too varied to compress by hand. My manager pointed out I was taking the long route.

First attempt
Manual categorisation

Reading and tagging by hand. Accurate, but the throughput made a full year of data impractical to finish in any useful timeframe.

What I switched to
Generative AI for the bucketing, Excel for the view

Used AI to apply the taxonomy at volume, then built the dashboard to visualise the distribution — cutting the analysis from a grind into something repeatable.

Output

Four themes, four shaped pitches

Deliverable

Not a report — a PRD with a section per theme, each carrying an explicit ask for Design and for Engineering. Pickup-ready, no second round of interpretation.

ThemeRespondentsExpected gainWhat I asked for
Spellchecker
103
+5%
Context-aware correction regardless of typed or pasted text, plus grammar and auto-caps, and academic vocabulary support.
Editor formatting
100
+5%
Font choice and sizing, indents, line and double spacing — readability controls students expected from a writing tool.
Referencing
85
+4%
Direct insertion from reference managers without losing formatting, and better in-text referencing in the workspace.
Tables
64
+3.5%
Row and column management closer to a word processor — inline editing, resizing, cleaner paste handling.
Why it landed

Not four opinions competing on rhetoric — four ranked volumes, each with a named owner and a defined ask. Spellchecker, the largest, is in build.

The forward case

What shipping all four would move

NPS was negative when the analysis began. Each pitch carries its own expected gain, sized from the students who raised that theme.

+17.5
points of expected NPS gain,
combined across the four pitches
Spellchecker+5.0
Editor formatting+5.0
Referencing+4.0
Tables+3.5
Spellchecker · +5.0 pts · 103 respondents Editor formatting · +5.0 pts · 100 respondents Referencing · +4.0 pts · 85 respondents Tables · +3.5 pts · 64 respondents
Caveat

Estimated from respondent volume, straight from the PRD. Modelled, not measured — the case for investment. The measured movement is the trajectory above.

Reflection

What I'd tighten

4 pitches shaped & handed off Spellchecker in build Dashboard now a standing reference Post-ship NPS movement not yet attributable

The lesson I actually took from this

Students never reported “an editor problem.” They reported font size, then indentation, then double spacing, then typefaces — dozens of individually-small complaints.

As tickets

Four low-priority items nobody schedules.

As one category

A hundred students describing one coherent gap — an Editor Improvements workstream with one design solution.

Bucketing symptoms up into a cause, then scoping against the cause, is what turned this into roadmap items instead of a backlog of noise.

And one on method

Hand-tagging was accurate and far too slow. Generative AI for bucketing plus Excel to visualise made a full year tractable. The useful part isn't the tooling — it's dropping work I'd already invested in rather than defending it.