Just point the LLM at the data?
Every retailer I talk to wants the same thing from AI: one place to ask a question and get an answer that spans promotions, pricing, assortment, and supply chain. And the first idea on the whiteboard is almost always the same: give a large language model access to all the databases and let it figure things out.
It doesn’t work. Not because the models aren’t smart enough, but because the meaning of the data isn’t in the data. It lives in the applications that were built around it, and in years of client-specific configuration that never made it into a table.
I’ve spent my career implementing enterprise planning systems both across merchandising and supply chain at Tier-1 and Super Tier-1 retailers, first at Predictix, IBM and Manhattan Associates and now with PromoAI at Cognira. That experience makes me fairly confident about one thing: in the age of agents, the semantic layer is not a nice-to-have. It is absolutely essential.
Thirty definitions of margin
Ask five retailers how they calculate margin and you’ll get thirty answers. Is it before or after vendor funding? Does it include scan-backs, off-invoice allowances, shrink, freight? Is it calculated at the item, the promoted group, or the category including halo and cannibalization?
The same goes for nearly every term a merchant uses. What counts as a new customer versus a lapsed one. What baseline means. When a promotion is “live” versus “approved.” Which stores are in a zone this quarter.
None of these answers are written down in one place. They’re encoded in how an application was configured during implementation, in business rules that were negotiated with the client, and in calculation logic that someone wrote years ago. Then there’s the second problem: even when you know the definition, you still need to know which table, which column, and which version of the data it lives in.
This is exactly where the naive approach breaks down. Give a model too little context and it fills the gaps with guesses. It finds the column that looks like margin and runs with it. Give it too much (a schema with thousands of tables, years of naming conventions, half-deprecated view) and it gets lost in the noise, joining the wrong tables or mixing definitions. Either way the error rate goes up, and the answers still sound confident.
Enterprise applications have always solved this the same way. You sit with the client, you codify their definitions, you configure the system, and from then on every screen and report answers with that understanding built in. That work is most of what an implementation actually is.
The most valuable answers aren’t in any table
There’s a deeper problem with the “point the LLM at the warehouse” idea. Many of the numbers merchants actually care about don’t exist as stored data at all.
Take a simple question: “Should I run this promotion again next month?” Answering it requires a baseline forecast, an estimate of incremental lift, the cannibalization of neighboring items, the halo on the basket, and the vendor funding on the table. Those are outputs of models that the application runs, often on demand, with parameters specific to that client and that category.
An LLM reading raw sales history can’t reconstruct that. At best it produces a plausible-sounding approximation. At worst it confidently tells a category manager that a promotion made money when it didn’t.
This is why the application matters more in an agentic world, not less. The application isn’t just a user interface sitting on a database. It’s where the forecasting, the optimization, and the decision logic live. Building that from scratch takes years. Even with an LLM writing the code, it would take quarters.
Answering is easy. Acting is hard.
Most of today’s demos are about answering questions. The real value, and the real risk, comes when agents start doing things: creating a promotion, changing a price, approving an ad, committing vendor funds.
When an agent answers a question with the wrong margin definition, it’s embarrassing. When it acts on the wrong definition, it costs money. And every one of those actions is governed by rules the application already enforces:
- Business rules. Price ladders, minimum margins, promotion frequency limits, vendor contract terms.
- Workflow. Who has to approve what, in which order, before an offer goes live.
- Permissions. Which categories, banners, and stores this user is allowed to see or change.
- Audit. A record of who changed what, when, and why, which matters a lot more when “who” is an agent.
Here’s the part people miss: many of these rules don’t live in the database either. They’re spread across the application: in API validation, in service code, in checks the user interface runs before a merchant can hit save. Nobody wrote them down in one place because nobody had to; the application was the only way in. An LLM that goes straight to the data doesn’t just skip those rules. It doesn’t know they exist.
A central agent that bypasses the application to write directly into data has to rebuild all of that, for every domain, for every client. The application has already done it. It should own the action, and it should be accountable for the answer.
“But we’re building an enterprise semantic layer”
The strongest counterargument comes from the data platform world. Snowflake, Databricks, dbt, and ontology-style platforms all promise a unified semantic layer: define your metrics once, centrally, and every tool and agent uses them.
I think that’s genuinely useful for reporting. If the question is “what were sales by region last quarter,” a well-governed central metric layer is exactly right.
Where it struggles is decision logic. A central layer can define margin. It can’t easily own a promotion lift model, an elasticity curve, or the rules for which offers can stack. Those evolve with the application, are tuned per client, and often need to be computed rather than looked up. Pulling all of that into one enterprise model means rebuilding each application’s brain in a second place and keeping the two in sync forever.
And across organizations, it’s not a realistic goal at all. Every retailer is its own semantic universe. Nobody will ship one model that works for all of them.
Each application owns its semantics and its agent
The model I believe in is federated. Each enterprise application owns the semantics of its domain, and exposes them through its own agent that understands the domain deeply.
At Cognira, that agent is Cora. When a merchant asks Cora a question in plain language, Cora isn’t guessing at table names. It works through PromoAI’s semantic layer: it knows what lift, baseline, and promoted margin mean for that specific client, which models produce them, and which rules apply. It understands the intent behind the question, gets the answer from the right place, and responds in the merchant’s terms.
That’s the same thing enterprise applications have always done: codify the client’s definitions and answer with that understanding built in. The difference is the interface. Instead of a screen, it’s a conversation. And increasingly, the one asking won’t be a person. It will be another agent.
This is the real shift. An enterprise-level agent that spans promotions, pricing, and assortment doesn’t need to understand all of those domains itself. It needs to know which agent to ask.
What this means for retailers
If you’re a retailer planning your AI strategy, I’d suggest three things.
First, judge enterprise applications by their semantic layer, not just their screens. Ask a vendor how their system knows what your margin means, and whether an agent can ask it a question and get a trustworthy answer and whether your own agent will be able to work with theirs.
Second, don’t assume a central data layer will replace domain logic. Invest in it for reporting and governance. But let the systems that own forecasting, optimization, and execution keep owning them.
Third, design your enterprise agent as an orchestrator, not an oracle. Its job is to understand the merchant’s question, route it to the domains that can answer, and bring the pieces together.
Which raises the obvious next question: how should these agents actually talk to each other? Should each application expose tools through MCP, or should agents collaborate as peers through A2A? That’s the subject of my next piece.
By Bahadir Ustaoglu
— Cofounder / Chief Operating Officer, Cognira