Why Deterministic Beats Probabilistic AI When the Board Asks Twice
.webp)
It's 7:40 on a Thursday, the board deck goes out at nine, and your CFO asks the AI agent for quarterly net revenue retention. It says 112%.
She asks the exact same question again, just to be safe. Now it says 108%.
Which one goes on the slide? Neither. You now hold two numbers and zero trust, and somebody on your data team is about to lose a morning rebuilding it by hand.
That morning comes down to one design choice. The AI agent your CFO asked was probabilistic, so each run was a fresh guess at the query. A deterministic system would have handed her the same number both times, plus the exact path behind it.
Probabilistic systems are very well-read guessers. An LLM writes SQL the way it writes anything: one likely word after another. Ask twice, and a small shift in those guesses can change a JOIN, a filter, or which revenue column it picks. A deterministic system behaves more like a compiler. The same input goes through the same fixed rules and produces the same output, whether you run it once or a thousand times. You can also trace any answer back through each rule that produced it.

Each approach has a job. LLMs are great with language, so let them draft, summarize, and brainstorm. When a number will be read aloud in a boardroom or signed by a CFO, it needs the compiler.
Temperature zero does not make an LLM repeatable
LLMs predict the most likely next word. Most teams assume they can switch that behavior off. Set the temperature to zero, and the model should give the same answer every time.
The research disagrees:
- 1,000 runs, 80 different answers: In September 2025, Thinking Machines Lab sent one prompt to Qwen3-235B 1,000 times at temperature zero. They got 80 unique completions. The main cause was server load, because batch sizes shift with traffic and the math shifts with them.
- Accuracy that swings between runs: A study of five LLMs across eight tasks ran each model 10 times under "deterministic" settings. Accuracy varied by up to 15% between runs, with a gap of up to 70% between best and worst performance. None of the models delivered repeatable accuracy across all tasks.
That's the difference between "AI writes SQL" and "AI understands the question, software writes the SQL." The one step that involves a model is also the narrowest. It turns words into a structured intent, and everything downstream of that intent is compiled.
In production, this design delivers 95% query accuracy (F1 score) across enterprise datasets and close to 100% consistency. Same question, same answer, every time.
Messy enterprise data makes the problem worse
Now add your warehouse to the picture. Clean demo data makes text-to-SQL look solved. Real enterprise schemas tell a different story.
Every enterprise I talk to has some version of revenue, revenue_final, and revenue_corrected sitting side by side. An LLM writing SQL will pick one. Tomorrow it might pick another.
A bigger model won't close that gap. Context closes it: what each table means, which revenue is the real one, and how your business defines an active customer. That's the data context problem, and it sits underneath every wrong number.
The mechanism: Interpret, compile, run the same plan
We flipped this problem on its head. The LLM handles only language, the one job it does well. Everything that decides your number runs through software that behaves the same way every time:
1. The LLM interprets the question: Its only job is intent. "What was net revenue retention for enterprise accounts last quarter?" becomes a structured request with a metric, a segment, and a time window.
2. A symbolic engine compiles the plan: Walt's deterministic inference engine builds the query from the Data Context Graph (ontology, knowledge graph, semantic layer) and predefined analytical patterns. No probability picks the JOIN. No model guesses your column names.
3. The same plan runs every time: Same intent, same plan, same SQL, same answer. Each step traces back to the exact definition and pattern that produced it.
That's the difference between "AI writes SQL" and "AI understands the question, software writes the SQL." The one step that involves a model is also the narrowest. It turns words into a structured intent, and everything downstream of that intent is compiled.
In production, this design delivers 95% query accuracy (F1 score) across enterprise datasets and close to 100% consistency. Same question, same answer, every time.
With determinism, you get:
- An answer you can defend: When a director asks where a number came from, you can show the question, the compiled SQL, and the source table behind it.
- A trail compliance can follow: Every answer carries lineage from the question asked, to the SQL generated, to the data source accessed. Nobody has to reconstruct it after the fact.
- Data that stays put: The Data Context Graph stores context about your data. The data itself never leaves your warehouse or your security perimeter.
Where probabilistic still earns its place
None of this makes LLMs the villain. I'd happily use one to draft the board memo, summarize call notes, or brainstorm why churn spiked. That's language work, and they're great at it.
My rule is simple: if a number is to be quoted, signed, or acted on, run the plan compiled by a symbolic engine (like WALT’s deterministic engine). Let the model write the words around it.
Ready to test deterministic against probabilistic on your board metric?
Friends say that when a problem and I walk into a room, only one of us walks out. So bring me the number that keeps moving at Gartner IT Symposium/Xpo™, from October 19-22, 2026 at Walt Disney World Swan & Dolphin in Orlando, Florida. Pick that one metric your board asks about every quarter. Ask Walt as many times as you like, and watch the plan it compiles, line by line.
Sources
He, Horace and Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference", Thinking Machines Lab: Connectionism, Sep 2025.
Atil et al., arXiv: Non-Determinism of "Deterministic" LLM Settings




.webp)

.webp)