Data Engineering

Your Warehouse Knows Numbers. Does It Know Meaning?

Thiru Arunachalam, Founder & CEO, WALT
October 6, 2026
All posts

Every enterprise has a system that swears it knows who your customers are. Most have three. Each one is right about its own records and silent about everyone else's.

Ask a simple question: how many customers do we have? Salesforce counts the parent company as one account. Your ERP invoices its US subsidiary separately, and your support desk holds a third record for a plant in Texas.

Investor Jamin Ball has a useful lens for this. He thinks about systems of record as "where does the truth live." For customers, the honest answer is often: in three places at once.

The last decade's fix was the warehouse. In Ball's words:

“The pitch was that if you poured all the data into one place and layered semantic models and metrics definitions on top, you would finally get a single source of truth that analytics, dashboards and downstream tools could all agree on.”

He adds that, in practice, the vision "got part of the way there". We got the ‘one place’. We never got the agreement. Centralizing everything takes years, and by the time it lands, the business has moved.

For years, a human closed the ‘meaning’ gap. A senior analyst knew which "customer" the CFO meant and wrote the SQL to match. However, your AI agent has no such analyst sitting beside it, so it counts rows and ships the answer.

Tomorrow's system of record has to hold meaning itself: which records are the same company, which definition each team uses, and who approved it. That's the data context graph.

Grupo Mag or Mag Bakeries USA: one customer or two?

Let's make this concrete. The companies below are made up for this article. Any resemblance to your CRM, however painful, is purely accidental.

Say you sell packaging to food manufacturers. Your systems hold four records that all smell like the same buyer:

1. CRM: An account called "Grupo Mag," owned by your enterprise team in Mexico City, with one global deal attached.

2. ERP: A bill-to entity called "Mag Bakeries USA, Inc." with its own tax ID, its own payment terms, and its own invoices.

3. Support desk: A queue labeled "MAG BAKERIES TX," full of tickets from a plant manager in Texas.

4. Contracts folder: A master agreement, signed by the parent, that covers every subsidiary.

So which of these records is "the customer"? The honest answer depends on who's asking.

Your warehouse has no way of knowing who's asking, so it counts rows and hopes for the best.

"Customer" means something different on every floor

Software architect Martin Fowler called this out more than a decade ago. He wrote that he sees confusion recur "time and time again" with words like "Customer" and "Product," because each part of a business uses them a little differently.

Walk the Grupo Mag example through your org chart and you'll see it:

- Finance: A customer is a legal entity you invoice. Mag Bakeries USA counts. The parent may never see a bill.

- Sales: A customer is a relationship with a quota attached. Grupo Mag is one logo, one account executive, one number.

- Customer success: A customer is whoever opens tickets. That plant manager in Texas may have never heard of Grupo Mag.

- Risk and compliance: A customer is whatever sits at the top of the ownership chain.

Every one of those definitions is correct. Mash them into one query with no context and you get a number nobody can defend in a board meeting.

That's the data context problem.

Why your AI agent guesses (with a straight face)

A large language model looking at your warehouse sees a table called customers, a column called customer_id, and a very tempting COUNT(DISTINCT). It has no idea that three of those IDs roll up to one parent. It counts them anyway, formats the answer beautifully, and heads to the next question.

When these LLMs meet messy, real-world enterprise data, you get answers you can't fully trust. And the stakes are real. Gartner expects at least 50% of agentic AI projects to be canceled by the end of 2027, citing escalating costs and unclear business value as two of the top reasons why.

Then there's consistency. Ask an LLM that writes its own SQL the same question twice, and it can pick a different join the second time. Your CFO will notice before you do.

None of this comes cheap and costs organizations, as Gartner puts it, millions on average. I'd bet a healthy slice of that is really poor data meaning: good rows, wrong interpretation.

How to solve the data context problem: Settle what "customer" means before anyone asks

Most teams try to fix this at query time. They stuff a note into the agent's system prompt: "Remember, Grupo Mag owns Mag Bakeries USA." That's a sticky note on a jet engine.

A context graph moves the decision earlier. By the time anyone types a question, the graph already holds three things:

- Identity: Grupo Mag, Mag Bakeries USA, and the Texas plant are linked as parent and child records, with the evidence behind every link.

- Meaning by audience: "Customer" resolves to the billed legal entity for finance and to the parent account for sales. Same word, right meaning, right person.

- Governance: Which definition is approved, who approved it, and which version is live today.

So when your CFO asks how many new customers you added in Q3, WALT’s Data Context Graph doesn't guess. The LLM does one job: figure out what the CFO meant. WALT's deterministic engine then builds the SQL from the graph, so it's the same question, same answer, every time.

How WALT builds that context map

Step 1: Start from real questions

WALT anchors on the 20 to 50 questions your business actually asks, so the entities that matter get resolved first.

Step 2: Read what your team already knows

WALT introspects the values inside columns, reads your dbt models and query logs, and finds the joins your analysts have been making by hand for years. That's tribal knowledge, finally written down.

Step 3: Match at scale, judge at the edges

Pattern matching and vector similarity shrink millions of records into a short list of likely pairs. An agent reviews only the ambiguous ones, like whether "Joe's" and "Joe's Sausage Stand" are the same business.

Step 4: Put it where humans can edit it

Every link and definition lands in human-readable YAML. Your data steward can override any of it, and every change is versioned.

Step 5: Keep it fresh

WALT re-checks the graph every night against new data and new questions. When Grupo Mag buys another bakery next spring, the graph catches up without anyone filing a ticket.

Each team’s “customer” lands on a different record. A context graph links them before any query runs.

One word, four teams, four different records. The graph settles the mapping once, so the SQL never has to guess.

This October, bring us your messiest customer

Every data team has one. The account that shows up three times in the CRM, twice in billing, and once in a spreadsheet someone's manager keeps "just in case."

Book a demo and bring it. We'll run WALT on it live and show you the evidence behind every link and the SQL behind every answer. Worst case, you leave with a sharper map of your mess than any spreadsheet has given you.

Your warehouses already know the numbers. Let's teach them what they mean.

Book a demo and bring your messiest number.

Sources