LLM · 19 Aug 2026
LLMs on your warehouse, not a chatbot in the way
Atalaya Digital
Finance will not ask ChatGPT for revenue if they already got burned by three dashboards. They will ask if they can trust the number. An LLM does not fix that. A model layer does. The model on top is optional.
Retrieval is a join
We do not dump the schema into a 128k window and pray. The catalog is a table: name, grain sentence, owner, PII flag. The LLM sees three objects that match the question, not four hundred.
If fct_orders is "one row per order per plant," that sentence goes in the prompt every time. Hidden grain is how you get double-counted VAT.
SQL you can refuse
Generated SQL hits a read replica or a warehouse user that cannot DROP. We lint it: no SELECT * on raw, no unscoped patient grain, no query without a date filter on a fact.
If the lint fails, we do not run it. We send the SQL back to a human. That is slower than a demo. It is faster than explaining a leaked extract.
Grounding beats personality
The answer is a number plus the query plus the watermark. "Looks like revenue is up" with no SQL is a slide. We show the table, the filter, and when the partition last landed.
You can wrap that in a Next.js chat if you want. The product is still the warehouse.
Cost is an SLO
A 70B call on every dashboard load is a hobby. We cache the plan for a repeated question. We use a small model to classify intent and a bigger one only when the catalog match is weak.
If the board already answers the question, do not open the LLM. Agents and chat are for the questions you have not productized yet. That list should shrink, not grow.