6/25/2026 · Ben Weller

Where Your Data Actually Lives

Where Your Data Actually Lives

You ask your ops lead for last month's revenue number. She sends you one. You ask your bookkeeper. He sends you a different one. Both are confident. Both are pulling from a real, current source. Both numbers exist somewhere in the business right now.

Call it a location problem: revenue is fine, the number just lives in more than one place, with no agreed rule for which version wins when they disagree. Most small companies have plenty of data. What they lack is a single authoritative source for each thing that matters.

How it happens

Nobody decides to create a location problem. It accumulates.

Year one, you run everything out of QuickBooks. Year two, your sales lead starts keeping a tracker in Sheets because QuickBooks doesn't surface the pipeline view she needs. Year three, your ops manager is maintaining his own version because the Sheets tracker doesn't reflect returns. By year four, you have three sources and a standing argument about which one is right.

Each version made sense when it was created. The sources were never rationalized: nobody ever said, out loud and in writing, which one is the canonical source and what happens to the others.

What "canonical" actually means

Canonical means simple: when this number and any other number disagree, this one wins. That status comes from a designation, not from recency or from who happens to maintain it. Someone names the source of record, in writing, and every other version becomes explicitly subordinate to it.

Naming a canonical source is a decision, made once and documented, not a technical setting you configure. No tool will make the choice for you.

Take any piece of data your business runs on: revenue, headcount, pipeline, inventory, customer records, whatever is load-bearing in your specific operation. Either it has a named canonical source or it doesn't. Without a named source, every report, every meeting, and every decision downstream of that data is contaminated by the ambiguity, even when nobody notices.

The audit you can do right now

Take the five numbers that matter most to how you run the business. For each one, answer three questions:

Where does this number come from? Not where you check it. Where it originates. If you're not sure, that's already data.

Who else has a version of this number? List every place a version of it exists: tools, spreadsheets, email threads, people's heads.

If your version and someone else's version disagree, what happens? If the answer is a conversation, that conversation is the tell. It means no source has been named, and the same disagreement will resurface next quarter.

For most companies running between fifteen and seventy-five people, this exercise surfaces two or three places where the answer to question three is "we have a conversation." Those are your location problems. They are also, almost always, the places where operational drag is highest, because every decision that depends on that data requires a synchronization step before it can move forward.

What you do with what you find

The work is to name a winner and write it down.

For each contested source: name the winner. Write it down. Tell the people who maintain the other versions what changed and why. Then comes the step that gets skipped most often: deprecate the losing versions. Deprecating means more than archiving. It means making them unavailable as decision inputs, or at minimum, labeling them clearly as non-authoritative.

The canonical source is the one everyone agrees counts. A spreadsheet that everyone treats as the source of record is better than a purpose-built tool that half the team ignores.

The second-order problem

There's a reason this matters beyond operational tidiness.

When your data sources are contested, you can't trust your own reporting. When you can't trust your reporting, you make decisions based on feel. When you make decisions based on feel, you can't learn from them, because you can't reconstruct what you actually knew at the time.

Small companies often blame their tools when this happens. The data is messy, the system is wrong, they need a better platform. That's sometimes true. Usually the harder, cheaper fix is being avoided: naming which source counts and retiring the rest.

That decision is free. It takes about an afternoon and someone with the authority to make it stick.


No comments here. Ben reads email: ben@deployedskills.com.

All field notes