When two dashboards disagree, the data is usually fine. The definitions are not, and an LLM on top of the warehouse will pick one and sound sure.
Fix who owns each metric and write the definition down. Do that before anyone chats with the warehouse.
Two dashboards, one word, two answers
Picture two teams in the same weekly meeting. Both bring a number for active customers. The numbers do not match.
Finance counts a customer as active if they paid an invoice this quarter. Product counts a customer as active if they logged in this month. Both teams are correct by their own rule.
Nobody wrote either rule down. Nobody has the authority to pick one. So the meeting turns into a debate about whose pipeline is broken.
The pipeline is not broken. The word is.
The cost is more than one lost hour. Leaders stop trusting both numbers. Analysts keep private spreadsheets to defend their version, and each new report adds one more definition to the pile.
Definitional drift is not data drift
In an earlier post I wrote about data drift. Your model did not get worse. The data moved, and the outputs followed.
Definitional drift is different. The data stays the same. The meaning of the metric changes under it.
It happens in small, reasonable steps:
- An analyst excludes test accounts.
- A new region launches with a different billing cycle.
- Someone treats refunds as negative revenue in one report and ignores them in another.
- A team copies a query and adds one filter for a one-off ask.
Each change makes sense at the time. None of them gets written down. After a year, one metric name points to several calculations.
Your monitors will not catch this. Freshness checks pass. Row counts look normal. A catalog will list every version of the metric and still not tell you which one the business means.
Why the model sounds sure and is still wrong
Now put an LLM on top of that warehouse and let people ask questions in plain English.
The model reads the schema and maybe some column descriptions. It finds a column called active and writes valid SQL. It returns a clean number in confident, fluent prose.
The model did not make a mistake in the usual sense. It made a choice the business never made. It picked one of several definitions and did not tell you.
Ask the same question with different words and you may get a different definition. Neither answer will mention the other.
That is worse than two dashboards that disagree. At least the dashboards show the conflict. A chat answer hides the conflict behind good grammar.
The model cannot own a definition that nobody owns. It can only repeat the ambiguity faster and with more confidence.
Ownership comes before definitions
Most teams try to fix this with a glossary project. They gather every metric and write a description for each one. Then they publish a page. The page goes stale in a quarter.
A definition without an owner is a suggestion. Somebody has to decide what the word means and defend that decision when the next request comes in.
A common pattern I see is an owner field that names a team, such as Data Team or Analytics. A team is not an owner. An owner is one named person who can approve a change and settle a dispute.
Usually that person is the business leader who steers on the metric. The owner does not need to write SQL. The owner decides what counts and approves every change to that rule.
A data partner turns that decision into tested logic and keeps the calculation honest. One metric, one owner, one written definition.
What a written metric contract contains
I call this a metric contract because it binds the people who produce the number to the people who use it. It does not need a new tool. A short document in version control is enough to start.
Each contract should hold:
- Name and meaning: the metric name and one plain sentence that says what it measures.
- Owner: one named person, plus the data partner who maintains the logic.
- Calculation: grain, filters, inclusions, exclusions, time window, and time zone.
- Source: the certified table or model it comes from, and nothing else.
- Near misses: related metrics with different names, and how they differ.
- Do not use if: known caveats and cases where the number misleads.
- Change log: who approved each change, what changed, and the effective date.
- Tests: checks that confirm the calculation still matches the written rule.
When two teams need different rules, do not vote. Give them two metrics with two names. Active payer and active user can both be true. They cannot both be called active customer.
Before anyone chats with the warehouse
Text-to-SQL and chat-with-your-data demos look great on a clean sample schema. Your warehouse is not a clean sample schema. Do this work first:
- List the metrics leadership already argues about. Keep it short.
- Name one owner for each metric. Get their agreement in writing.
- Write the contract for each metric and resolve conflicts by renaming, not by voting.
- Publish the definitions in a governed semantic layer or certified views.
- Point the LLM at those certified definitions only, not at raw schemas.
- Make the model name the metric and the contract version it used in every answer.
- Make the model say it does not know when a question maps to no defined metric.
- Log those misses and send them to the owner as candidates for new contracts.
That last step turns the chat tool into a feedback loop. The questions people ask show you which definitions are missing.
What I would do this week
Pick the metric that caused the most recent argument in a leadership meeting. Find out how many calculations hide behind that one name.
Name an owner and write the contract on one page. Retire or rename every calculation that does not match it.
Then, and only then, let a model answer questions about it. The model can only be as sure as your definitions are clear.
