What Aerospace Taught Me About Microsoft Purview
I recently led an evaluation of Microsoft Purview for a housing association trying to get serious about data governance: tenant records, safeguarding files, legal documents, commercially sensitive board papers. Going in, I assumed AI would do most of the hard work. Purview can scan documents, recognise sensitive content, and recommend classifications. Given what large language models can now do, it seemed reasonable to expect governance was about to become largely automated.
One thing nagged at me throughout. I'd seen document classification work well before, just in a completely different industry, under completely different rules, and for reasons that had very little to do with how good the technology was.
A different discipline
Earlier in my career I worked in aerospace, where every document carried a classification. Not because a system decided it. Because the engineer writing it did, as an unremarkable part of the job. Much of that material was export controlled, military sensitive, or years of intellectual property. Misclassify it and you weren't looking at an awkward conversation, you were looking at export control breaches or serious commercial loss. Pausing to pick the right label wasn't bureaucracy. It was obviously worth the ten seconds it cost, and nobody questioned it.
A housing association is a different world. The information still matters, but the stakes and the arithmetic aren't the same.
Where AI helps, and where it stops
Purview is capable, and this isn't an argument against it or against AI. It's good at spotting recognised sensitive data like National Insurance numbers, applying default labels, flagging documents for review, and cutting manual overhead.
In our case, though, out of the box classification wasn't accurate enough to rely on without significant tuning: training custom classifiers on our own documents, adjusting exact data match rules, running iterative tests to cut false positives and false negatives. Even after that investment, it still wasn't fully reliable. That's not a knock on Purview specifically, it's what happens when pattern matching meets human judgement calls.
The harder problem was never the pattern matching.
It was business context.
A board paper about a restructure, a commercially sensitive tender, a safeguarding report with nothing a classifier would technically flag yet real harm if it leaked. A human understands instantly why these need protecting. An algorithm trained to spot payment card numbers doesn't.
So the real question isn't whether the AI is accurate enough yet. It's whether we ever expected it to make judgement calls that were never pattern recognition problems to begin with. I don't think it can, and I don't think better models quietly close that gap.
Why the difference is proportionate, not a maturity gap
In aerospace, mandatory manual classification is rational because one leak could mean major commercial loss or a national security incident. The effort is trivially justified against the downside.
In a housing association, asking every member of staff to classify every document, every day, is a substantial recurring cost. Weighed against a lower likelihood of catastrophic consequences, that cost doesn't clear the same bar. That isn't lower governance maturity, it's a proportionate response to a different risk profile. How much effort an organisation can justify is shaped almost entirely by the consequences of getting it wrong, and those vary by sector, regulator, and what's actually in the documents.
Defence and similarly high consequence sectors will reasonably keep mandatory manual classification. Regulated sectors like healthcare or finance will likely blend manual labelling where stakes are highest with automation elsewhere. Most other organisations will get most of the value from sensible defaults plus AI assisted recommendations, without asking staff to carry the full weight themselves. None of this is correct in the abstract; it depends on risk profile, regulatory exposure, culture, and how bad it would actually be to get it wrong.
The question that actually mattered
I started this evaluation asking whether AI could classify documents accurately enough. Looking back, that was never really what I was testing.
I was evaluating how much governance our organisation was actually prepared to buy, in time, in friction, in staff attention, not in software cost. The technology was never the limiting factor. The economics were.
Purview is making governance more tractable than ever, and that's worth taking seriously. But deciding what's worth protecting, and building a culture where that judgement gets made consistently, is still a human responsibility. Technology can help govern information. It can't decide what that information is worth.