Why a Schema That Satisfies a Human Analyst Breaks an AI Agent
Incomplete schemas work for humans but fail silently when AI agents act on them.

A schema built for a human analyst breaks the moment an AI agent tries to act on it, because the schema was never actually complete. It only looked complete to the people who already knew what it left out. That gap between what a schema says and what it means is where confidently wrong answers come from. Tracing it shows what a data layer has to carry when the reader is a machine instead of a person.
Reading schemas: humans versus AI agents
A column named rev_adj means nothing on its own. An analyst reads it correctly anyway, because somewhere along the way, someone told them it stands for adjusted revenue net of returns. The same analyst knows the customers table quietly drops churned accounts, and that the Q4 number in the quarterly deck always comes with a seasonal asterisk that no column in the database records. That context lives nowhere in the schema itself. It lives in the analyst's head, built up through onboarding, hallway corrections, and a few painful meetings where someone explained why last quarter's number didn't match.
An AI agent gets none of that. It has no onboarding, and it carries no memory from one session to the next. It reads the schema exactly as written, and nothing more. That's not a capability gap in the model. The agent is working from a specification that was built to be filled in by a human with institutional context, and an agent has no institution to draw on. The schema was always incomplete. Human readers just never noticed, because they supplied the missing context from memory without realizing they were doing it.
When ambiguity humans navigate through convention becomes a query error for agents
The dangerous failure is a query that runs cleanly and returns a number that's wrong in a way nothing in the schema can flag. Take revenue. Finance counts only booked revenue after returns processing. Marketing counts gross transaction value. Asked to calculate revenue, an agent has no embedded definition to fall back on, so it calculates whatever pattern it finds in the raw transaction tables, sometimes including test data, sometimes excluding international sales, depending on how the question happened to be phrased. This kind of drift existed long before AI showed up. Humans managed it through institutional knowledge and regular reconciliation meetings, where two slightly different numbers got argued into agreement. An LLM asked the same business question doesn't have that luxury. It needs one deterministic answer, every time, regardless of who asked or how.
Gross margin shows the problem at its sharpest. In most organizations, gross margin involves cost allocation rules, returns processing, promotional discounts, and shipping adjustments, and that logic often varies by business unit and sometimes by product line. An agent asked to calculate gross margin by region cannot go read a governance policy document and apply the right formula. It has no mechanism for that. The logic has to already be built into the infrastructure it's querying, or the agent will improvise, and improvisation on a financial metric is exactly the failure mode that matters.
Silent filters are the quieter version of the same trap. A customers table that excludes churned accounts by convention will return a count that is structurally valid and analytically wrong for any question about total customer base. Nothing in the table tells the agent a filter was ever applied. The query runs fine. The number comes back fast. It's simply the wrong number, delivered with the same confidence a correct one would carry. That's the core of what gets called silent failure: a wrong answer with no explicit refusal, no error message, nothing for the agent or a human downstream to catch. Businesses tolerated this kind of ambiguity for years, with conflicting definitions and undocumented logic smoothed over in footnotes and meetings. That tolerance ran out the moment AI agents started consuming the same data directly. The technology didn't change the underlying ambiguity. It changed what that ambiguity costs.
Agent autonomy turns a single schema flaw into a systemic problem
An agent doesn't pause on a wrong answer and reconsider it. It takes that answer and uses it as the input to the next step, and the next, compounding the error through every subsequent action. A chatbot hands a draft to a human and waits. An agent acts. When it misreads the data, it executes on the misreading, and there's no review step built in to catch it before the action is taken.
The math behind this is unforgiving: an agent working through many sequential steps carries each prior error forward as a new premise for the next decision. A small inaccuracy introduced at step one gets amplified by the time the workflow finishes. DeepMind has framed this directly: a small error rate, compounded across thousands of planning steps, makes the odds of the final answer being correct close to random. Multi-agent systems make the exposure worse, not better. One agent's bad output becomes the next agent's bad input, and the error cascades through the whole chain before a human ever sees the result.
The cost of this isn't theoretical. In late 2025, an AI résumé-screening system used in ICE recruiting misread keywords, flagging people who'd simply mentioned "compliance officer" as experienced law enforcement. Hundreds of recruits were routed into the wrong training track before anyone caught the error and reversed it, a month later. The root cause wasn't a bad model. The data itself held meaning the system had no mechanism to interpret correctly.
The scale at which enterprises are discovering this gap in production
The dominant explanation for why AI pilots stall before reaching production is no longer model quality. It's the data layer underneath them, specifically context fragmentation. Specifically, it's context fragmentation: enterprise data that lacks the meaning, lineage, and governance an agent needs to act on it safely, regardless of how much data there is or how capable the model is. Analyst Tony Baer's "Data 2026 Outlook" argues that the next phase of AI adoption will be decided by semantics, meaning who controls the definitions, context, and relationships AI systems rely on to reason. Baer describes this as the rise of "semantic spheres of influence," a landscape where ambiguity that used to be tolerable no longer is, and meaning has to be made explicit.
Teradata's Arrested Automation report, which surveyed senior technology and data leaders, found that the leading barriers they cite are data lacking the metadata, context, and relationships agents need, and data fragmented across systems that can't be connected in real time. A BARC Spotlight report by analyst Florian Bigelmaier found that most enterprises admit their own data and analyses lack reliability and interpretability. That finding applies just as directly to any agent querying those same sources.
The distance between how fast enterprises want to deploy agents and how ready their data actually is remains wide. Adoption of task-specific agents is forecast to grow sharply over the next few years, yet a substantial share of agentic AI projects are expected to be canceled before they ever reach completion. The reasons cited are escalating costs, unclear business value, and risk controls that weren't adequate for the job.
What a schema built for an AI agent must carry that a human-analyst schema does not
Making a schema ready for an agent isn't a matter of cleaning it up further by the standards analysts already apply. It requires embedding an entirely different category of information, the kind human analysts have always supplied from memory instead. A schema built for people documents table provenance and gives columns reasonable names. A schema built for an agent has to carry the full business meaning behind those names: metric definitions, entity relationships, domain-scoped logic, and the reasoning behind every silent filter that a human would otherwise explain in a hallway conversation.
Most enterprise data stacks are missing three things an agent needs before it can be trusted with a query. Fresh grounding is one: retrieval has to reflect current reality rather than a stale export, since a snapshot that's perfectly fine for a weekly dashboard is already out of date for an agent making a decision in real time. Clean, semantically described schemas are the second: structured records the agent can query without guessing at field meanings or applying the wrong aggregation on its own. Permissioned access scoped to the actual end user is the third: the agent has to operate within that person's real entitlements, not inherit the broad access of a shared service account.
A BI semantic layer and an AI semantic layer stop being the same thing here, even though they sound alike. A BI layer's primary consumer is a human, looking at a dashboard or a report. An AI semantic layer's primary consumer is an autonomous or semi-autonomous agent, acting on what it retrieves. A BI layer generally covers metric and dimension definitions, enough for a person to build a chart correctly. An AI layer has to cover the full business meaning behind the data: the entities, the relationships between them, and the logic specific to each domain. And a BI layer gets queried at human speed, a handful of times a day by a handful of people, while an AI semantic layer gets queried many times within a single agent run, at machine speed. Consistency and determinism matter more than flexibility did for a dashboard.
Gross margin again makes the case concretely. Define it once in the semantic layer, with the expressions for time-based calculations, conditional logic, and nested aggregations built in, and both a Tableau dashboard and an AI agent return the same number. The definition lives in infrastructure at that point, not in an analyst's memory, which is the only way it survives contact with a system that has no memory of its own. None of this works as decoration. A governance layer is the mechanism that turns a raw schema into something an agent can actually trust, and without it, agents hallucinate calculations straight from raw data structure because business context was never there to constrain them. Sensitivity has to be evaluated where data gets combined, too, not just field by field. A schema that marks individual fields correctly but carries no logic for what those fields reveal once joined together isn't sufficient for agent-level access, no matter how clean each field looks on its own.
Enforcing meaning at query time separates governed agents from ungoverned ones
Embedding business meaning into a schema matters, but it isn't enough on its own. That meaning has to be enforced the moment an agent sends a query, not assumed to have survived the trip from a data dictionary into the model's reasoning. Plenty of organizations already did the documentation work: training sessions were held, data dictionaries were maintained, governance policies were written down in detail. Then AI arrived, and all of that documentation turned out to be insufficient, because it was never embedded in the infrastructure itself. A policy a human can read and a policy a machine can enforce are not the same artifact.
Permissions have to be evaluated at query time, scoped to the actual person the agent is acting for, not inherited from a service account's broad access. A finance agent should see finance systems and nothing more. A sales agent should see the CRM and nothing more. Built this way, the blast radius of any single agent going wrong stays bounded by design. Write actions need even tighter control than reads, since the two no longer carry the same risk. An agent that can act on its own output, with no human reviewing the step in between, needs stricter governance on anything that changes a system than on anything that merely reads one.
Audit logs have to capture identity, intent, and lineage together. At the query volume an agent generates, a log that records only what was queried, without recording who the query was effectively acting for and what business question it was meant to answer, stops being useful for debugging or for compliance. Access also needs to be revocable at the level of a single workload. Tying it to a shared service account means pulling access breaks every other workload riding on that same account, which makes revocation something nobody wants to do, which makes the whole system harder to govern in practice.
The industry has converged on this as the answer. At Informatica World 2026, Salesforce announced what it called the industry's first unified agent and context catalog, built with native Model Context Protocol support so governed data management capabilities can be exposed as reusable services any agent can call directly. At dbt Summit 2026, held September 15 to 18 in Las Vegas, the dbt Semantic Layer, MetricFlow, and a Model Context Protocol server were presented together as governed context agents can query through one standard interface, framed explicitly as a prerequisite for agentic AI rather than an optional optimization layered on afterward.
Research into semantic-layer failures has found they behave differently from raw-schema failures. Raw-schema failures tend to produce silent hallucination, confident wrong answers that propagate with no signal attached to them anywhere in the output. A governed semantic layer changes that failure mode. Enforcing meaning at the point of query makes an agent fail loudly, in a way a human can catch, instead of failing silently, in a way nobody notices until the bad number has already been acted on three steps downstream.
What an AI-ready data layer looks like in practice
Most enterprises are currently 0-for-3 on what an agent actually needs: no fresh grounding, no semantically described schemas, no permissioned access scoped to the real user. The fix for that is not a new warehouse, and it doesn't require migrating existing data or rebuilding pipelines that already work. It requires adding the layer the data was missing all along.
A semantic layer built for AI agents sits on top of the warehouses, SaaS tools, and operational systems already in place. It exposes a single interface where table descriptions, metric definitions, and entity relationships are available to every query that passes through it, whether that query comes from a dashboard or from an autonomous agent acting on a business process. The underlying infrastructure doesn't need to change. What changes is whether the meaning a human analyst used to carry in their head is now written somewhere a machine can read it, enforced at the moment it matters, every single time a query gets asked.

