AI-ready data
FeaturesLong read

How Agents Actually Consume Data Once It Reaches Them

Agents need semantic layers to avoid confident wrong answers.

Columnist · · 10 min read
Features · October 7, 2026 · 10 min read · 2,319 words

An agent acts on whatever data it is handed, and it acts with full confidence whether that data is right or wrong. That single fact is the reason data infrastructure built for human analysts cannot simply be handed over to agents and expected to work. A human analyst pulling a revenue number brings a mental checklist built over years: which tables are stale, which columns have names nobody trusts, who to ask when a figure looks suspiciously round. Sadalage and Chandrasekaran at Thoughtworks describe the gap: when data feels wrong, a human double-checks it, while an agent confidently acts on it anyway. That gap is structural, produced by model quality and prompt wording having nothing to do with it.

Every piece of implicit labor an analyst used to perform now has to live inside the data itself. Sanity-checking a total, knowing the fiscal year starts in February, understanding that "revenue" already has returns stripped out: none of that knowledge transfers to an agent unless someone encodes it into the data layer first. An agent has no colleague to ask and no instinct for a number that looks off. It takes the input and produces an output, and the output carries the same tone of certainty whether the underlying number is accurate or three weeks old.

That is what makes the failure mode so consequential. A stale or partial answer looks exactly as confident as a correct one. Human judgment used to sit between bad data and bad decisions, catching errors before they turned into actions. Once an agent is the one making the call, that buffer disappears, and whatever is wrong with the data flows straight into whatever the agent does next.

The three mechanisms by which agents retrieve and use numerical data

Agents get numbers into their reasoning through three distinct architectural paths, and the path in play determines what the data layer has to supply for the answer to be correct. These are not three flavors of the same thing. Each represents a different tradeoff between flexibility and determinism, and each leans on the layer beneath it in a different way.

The first mechanism is grounded text-to-SQL with semantic injection. An agent receives a question in plain language, semantic metadata gets injected into its context window (table descriptions, metric definitions, relationships between entities), and the model writes SQL on the spot. Asking this kind of agent for gross margin forces it to figure out, in real time, which tables hold cost data, which hold revenue, and how the two join together. The failure mode here is a confident wrong number: the model produces SQL that is syntactically valid and answers a question slightly different from the one that was asked.

The second mechanism is compiled named-metric queries. Instead of writing SQL, the agent selects a named metric, monthly recurring revenue, active users, gross margin, and the engine behind it emits SQL that has already been written and validated ahead of time. The model never touches the query itself. Asking for gross margin here gets the agent the same calculation every single time, because the calculation was fixed before the question was ever asked. The tradeoff is reach: this mechanism only answers questions the metric catalog was built to answer. An agent that needs a shape the catalog doesn't cover has to fall back to text-to-SQL or escalate to a human.

The third mechanism is the API-wrapped data product exposed as a governed tool. Either of the first two patterns gets surfaced to the agent as something callable, most commonly today through an MCP server or a function-calling endpoint. The agent never touches a query engine directly. It calls a tool, and the tool handles the data access and the governance underneath it. Through this mechanism, the agent invokes a capability built to answer exactly that kind of question about gross margin, with identity, permissions, and freshness checked at the moment of the call rather than assumed from whatever session the agent happens to be running in.

Which mechanism handles a given question changes what the data layer underneath it has to provide. Semantic context, governance enforcement, and freshness guarantees all have to show up differently depending on whether the agent is writing SQL, selecting a metric, or calling a tool. The next three sections take each of those requirements in turn.

What the semantic layer must supply per mechanism

A semantic layer is not a reporting nicety bolted onto a warehouse for the sake of tidy dashboards. For an agent, it's the mechanism by which meaning that used to live in an analyst's head becomes something machine-readable, something a model can actually act on correctly. That meaning takes three concrete forms: metric definitions for things like revenue, active users, churn, pipeline, and margin, written down in governed, machine-readable form; annotations on tables and columns that give a name like acct_st meaning an agent can understand and call; and declared relationships between entities so an agent can join two tables correctly.

Take revenue as the running example. Without a semantic layer specifying which definition of revenue applies, one agent might sum gross bookings, another might net out returns, and a third might include deferred revenue that hasn't actually landed yet. All three would generate valid SQL. All three would return different numbers with equal confidence. Nexla draws the line between AI-ready and agent-ready data on exactly this point: AI-ready data is stored in a warehouse with its schema recorded in a catalog, while agent-ready data has tool descriptions and semantic metadata that an agent can discover and call at the moment it needs them.

The stakes of getting this wrong are asymmetric. Text-to-SQL agents running without semantic grounding tend to produce confident wrong numbers. Give the same model full semantic context, and accuracy reaches 98.2%, Ganz and Perigaud found. A refusal is recoverable: someone notices, asks again, fixes the gap. A confident wrong number often goes uncaught.

The semantic layer also has to behave differently for agents than it ever did for human dashboard users. A BI semantic layer gets queried at human speed, mostly for reading, mostly at design time, when someone is building a report. An AI semantic layer gets queried many times within a single agent run, at machine speed, and every one of those calls needs a definition of revenue that matches every other call, regardless of which of the three mechanisms is doing the asking.

For the compiled-metric mechanism, the semantic layer is the source of truth for which named metrics exist and what they compute. An agent selecting "gross margin" from a catalog gets the same SQL every time only because that semantic definition is the canonical reference behind the catalog entry. For the API-tool mechanism, that same definition has to travel with the tool declaration itself. A tool that claims to return "revenue" without specifying which revenue, gross, net, deferred, leaves a semantic gap sitting at the tool layer, not just somewhere back in the query logic.

Why agent governance can't sit at the model or session layer

Access control built around a human logging into a session fails for agents, because agents act on their own, chain multiple steps together, and in a growing number of cases write directly to live systems. Governance built for humans checks who's logged in once and trusts that session for the rest of the workday. An agent making thousands of calls across a single workflow needs every one of those calls checked on its own terms, against the actual identity behind the request and the actual content of what's being asked.

A fraud-detection agent makes the stakes concrete. It can see transaction amounts and timestamps without ever seeing a payment card number or a customer's name, because governance filters what comes back at the moment of the query, matched against the agent's declared identity and its declared purpose. That is attribute-based access control doing exactly the job session-based permissions can't: deciding, call by call, what a specific requester is allowed to see for a specific reason.

Read access and write access carry different risk profiles, so that difference has to be built into the governance model directly. An agent that can only read a report is a very different kind of risk than an agent that can trigger an action in an ERP system or update a record in a CRM. Governance that treats both kinds of agents the same, under one shared service account, isn't fine-grained enough to catch the moment a write action goes wrong.

Nexla frames the requirement this way: row-level security, role-based access, and consent flags all have to travel with the data through every transformation it passes through, into embeddings, into whatever retrieves it at the end. Governance that only gets applied before the data lands in the warehouse never reaches the agent making the call six steps downstream.

Audit logging has to scale to match. A log that records only that a service account queried a table is close to meaningless once an agent is making thousands of calls inside a single workflow. What regulatory review actually needs is identity, intent, and lineage captured together, on every single call, because any one of those calls might need to be reconstructed later.

None of this requires flipping a switch overnight. Not every agent workflow needs to run fully autonomous from the start. Organizations can design workflows that escalate to a human at defined decision points while the query-time enforcement infrastructure is built out underneath. That staged approach buys time without pretending the underlying problem has already been solved.

Freshness for decision-making agents versus dashboards

A data snapshot that's perfectly fine for a dashboard can already be wrong by the time an agent acts on it. Agent-ready data needs a declared time-to-live attached to it, so the agent knows whether what it's looking at is fresh enough for the decision in front of it, not just fresh enough to render on a screen.

Nexla lists this as one of five things agents need from data: agent-ready data carries a known time-to-live, and the pipeline tells the agent the moment an answer is too old to trust. An agent acting on stale data does worse than an agent that simply declines to answer, because the stale answer gets acted on with full confidence.

The stakes shift with the mechanism in play. A text-to-SQL agent queries whatever the underlying table currently holds. If that table is a batch snapshot refreshed once overnight, an agent querying it at 2pm to decide whether to approve a transaction is working from data that could be hours old, with no signal anywhere that this is the case. Compare that to a discount-approval agent deciding in real time whether a customer qualifies for a promotion: a nightly snapshot might say yes when the live account status says no.

Compiled named-metric queries and API-wrapped tools are built to carry that freshness contract directly. The tool declaration can specify a maximum age for the data it returns, and the tool can hand back a freshness signal alongside the answer itself, so the agent or whatever is orchestrating it can decide whether to act on the number, retry the call, or kick the decision up to a human. Thoughtworks states the underlying requirement without hedging: data has to be accurate, fresh, and validated before the agent ever sees it. The confidence a human analyst used to supply by noticing a suspicious date or a total that looked too round now has to be built into the data itself, before the agent ever touches it.

Freshness isn't a single number that applies uniformly across every use case. A dataset that's fresh enough for a weekly margin report is nowhere near fresh enough for a fraud-detection agent evaluating a transaction as it happens. The data layer has to support freshness declarations set per product, not one pipeline cadence applied across everything an organization runs.

How data products package semantics, governance, and freshness

A data product is the unit that makes all of this usable by an agent in practice. It bundles semantic context, governance enforced at query time, and a freshness contract into one discoverable, callable asset, so an agent never has to reconstruct any of those properties on its own from raw schema at the moment it needs an answer.

Data products, in this sense, are curated, documented, governed datasets that carry their own metadata, their own lineage, their own access controls, and the contextual information an agent needs to use them correctly. What separates a data product from raw data sitting in a warehouse is that every property an agent needs in order to reason correctly about it is already attached to the product itself, not something the agent has to infer or guess at.

Discoverability belongs on the same list as semantics, governance, and freshness, because a data product nobody can find isn't functioning as one. Without a registry of what exists, what each thing actually means, and what a given agent is allowed to do with it, every agent interaction turns into a one-off integration built from scratch. A data product sitting behind a discoverable catalog is something an agent can call. A data product buried in a swagger doc somewhere is effectively invisible to it.

The same data product serves all three consumption mechanisms at once. A text-to-SQL agent injects the product's semantic metadata as context before it writes a query. A compiled-query agent selects one of the named metrics the product exposes and gets deterministic SQL back. An API-tool agent calls the product directly as an MCP server or a function-calling endpoint, bypassing a query engine. One underlying asset, three different ways in, and that's precisely the point: build the semantics, the governance, and the freshness contract once, into the product itself, and every mechanism an agent might use to reach it inherits all three automatically.

More in Features