top of page

Data Governance: The Difference Between a Library and a Landfill

  • Writer: E.C. Scherer
    E.C. Scherer
  • 4 days ago
  • 6 min read

Every company wants two things from its customer data: protection and value. Most chase the value first and treat protection like a tax they'll pay later, if legal makes them.


The pitch for the data usually comes from marketing or product, not security. Better personalization. Sharper segmentation. A model that predicts churn before it happens. Somebody in a strategy meeting points at the loyalty program's five years of behavioral history and says we should be doing more with this.


They're right. And almost none of them can, because nobody trusts the data enough to build on it.


Trust isn't a feeling here. It's a specific, checkable property. Can you say where a field came from? Can you say whether it's still accurate? Can you say who's allowed to touch it and why? If the answer to any of those is a shrug, the data doesn't get used for the thing marketing wants. It gets used for nothing, or it gets used anyway by someone who skips the question entirely, which is worse.


Unclassified data isn't neutral. It's dead weight with legal exposure attached. It costs money to store, creates risk just by existing, and produces zero business value because nobody can vouch for it. That's the asset-versus-liability question in one sentence: data that's protected and classified is a resource. Data that's neither is a bill you haven't received yet.


Why classification comes before value, not alongside it

This is where security and the business usually talk past each other. Security frames classification as a control. The business hears "compliance project" and files it under things that slow them down.


But classification is metadata, and metadata is what makes data usable by anything downstream, a model, an analyst, a report to the board. A sensitivity label doesn't just say "protect this." It says what the data is, which makes it possible to answer the next question: what can we responsibly build with it.


Take the behavioral profile sitting in the loyalty platform. Before it's labeled, nobody can tell you which fields are payment-adjacent, which are regulated location data, which are just click history nobody cares about. So the whole dataset gets treated like the most sensitive thing in it, or worse, treated like none of it, depending on who's asking. Either way, the analytics team either can't touch it or touches all of it without knowing what they're carrying. Neither produces a usable model.


Once it's labeled, the picture changes. Now the analytics team knows exactly which fields they can feed into a segmentation model without a legal review, and which ones need one first. The data doesn't get less valuable because it's classified. It gets valuable for the first time, because someone can finally build on it with a straight answer to "are we allowed to do this."


What this looks like when it works

We've watched this play out the same way more than once. An organization comes in wanting a security assessment, and underneath it is a business team that's been asking for better customer insights for two years and getting nothing. The data team wants to build. Legal keeps saying no, or worse, keeps saying nothing, because nobody can tell them what's actually in the dataset they'd be approving.


Classification breaks that stalemate. Not because it satisfies legal on paper, but because it gives legal something real to evaluate. A dataset with sensitivity labels and a documented lineage is a request legal can say yes to. A dataset that's just "the loyalty data" is a request they can only say no to, because no is the only answer that doesn't require someone to guess.


The foundation work looks unglamorous from the business side. Labeling schemes, DLP policies, retention rules, none of it produces a dashboard anyone gets excited about. But it's the same work that eventually lets the business build the dashboard. Protection and monetization aren't competing budget lines. They're sequential steps on the same project, and most organizations only fund the first half.


A catalog, not a stack of exports

Classification answers what a piece of data is. It doesn't answer where the real copy lives. That's a separate problem, and it's the one that quietly kills more analytics projects than bad labeling does.


Every campaign pulls a CSV. Every vendor integration syncs a copy. Every analyst who gets tired of asking for access exports what they need into a spreadsheet they control. None of this is malicious. It's just what happens when there's no obvious place to point people instead. Six months in, there are four versions of the same customer record living in four systems, and the one everyone's building on might be the stalest of the four.


This is what Purview's Data Governance solution, Unified Catalog and Data Map, is

built to fix, and it's a different piece of the platform than labeling or DLP. Data Map is the plumbing. It scans your data sources and builds the technical inventory of what exists across the estate, the layer nobody outside IT ever needs to look at directly. Unified Catalog sits on top of that and turns the inventory into something a business user can act on: governance domains with an accountable owner, and data products inside them that data stewards curate so consumers can find the real thing instead of guessing.


Here's the detail worth sitting with. Everything in Data Map and Unified Catalog is metadata. None of the roles or permissions in either one grant access to the underlying data itself. That's not a gap, it's the design. The catalog answers "does this exist, and am I allowed to use it." Sensitivity labels and DLP still answer "can this specific person actually touch it." Governance and protection are two different jobs. You need both running at the same time for either one to pay off.


That's also where the library stops being a metaphor. A governance domain has an owner. A data product has a steward keeping the description, the lineage, and the sensitivity current. When someone in marketing needs the behavioral data, they're not emailing around asking who still has last quarter's export. They're looking in one place that's supposed to be current, owned, and already scoped for what it's allowed to be used for.


Where this gets harder than the pitch suggests

Classification isn't a switch. A labeling scheme built for compliance alone tends to be too granular for anyone doing analytics to use, five sensitivity tiers with three sublabels each, technically accurate and operationally useless. If the goal includes enabling value, not just avoiding fines, the scheme has to be designed with the downstream use case in mind from the start, not bolted on after the fact.


There's also a timing problem nobody likes to hear. Classification doesn't produce insight next quarter. It produces the conditions under which insight becomes possible, which is a harder thing to put in a board deck than a revenue number. Organizations that fund this work expecting an immediate return usually give up right before it would have paid off.


The catalog has the same problem in a different shape. A governance domain without a real owner turns into exactly the sprawl it was supposed to prevent, data products nobody's updating, descriptions that stop matching what's actually in the source. Standing up Unified Catalog gets you the structure. It doesn't get you the stewardship. Someone still has to show up and keep the entries honest, and that's a people commitment, not a configuration.


And some data won't clear the bar no matter how well it's labeled. Data collected without a clear legal basis, or sitting with a vendor nobody can account for, doesn't become an asset once you classify it. Classification tells you what you have. It doesn't retroactively make bad collection practices good ones.


The actual choice

Every dataset in the building is already answering the asset-or-liability question, whether anyone asked it or not. The behavioral profile you've built over five years is either something the business can point to and defend, or something it's hoping nobody asks about. Classification is what decides which one it is. The catalog is what keeps that answer findable next year, not just true on the day someone audited it.


That's not a compliance initiative. It's the first step of the same project that gets you the personalization model, the segmentation, the churn prediction, everything the business actually wanted when this started. Protection isn't the toll you pay before you're allowed to create value. It's the same work, looked at from the other direction.

Mandatory Brain Break

A North American beaver feeding among willow saplings at a pond's edge at dusk.
Capture by @WanderingElias on Instagram

North American beaver — Castor canadensis

We caught this one at the edge of the water in the last light of the evening, tucked into a stand of willow saplings with a half-eaten stump in front of it. It stayed low and mostly still, working through what was left of the shoot in its paws while keeping the open water in view behind it. Beavers spend most of their visible time exactly like this, close to cover, close to an exit, rarely out in the open for long.


Turns out not every infrastructure project starts with a blueprint.

Build the catalog.

Comments


©2026 by E.C. Scherer

bottom of page