What a Data Dictionary Entry Looks Like for an AI Agent
AI agents need data dictionaries rebuilt for automated readers, not human ones.

A human analyst who hits a confusing column name can walk over to a colleague's desk and ask what it means. An AI agent can't do that. It has to work from whatever is in front of it at the moment it runs the query, and nothing more. That difference is the whole story behind why data dictionaries are being rebuilt right now. The data dictionary built for human analysts was designed for a different kind of reader, one who can ask a colleague what a column means, spot an anomaly on sight, and lean on years of institutional memory to fill in whatever the schema leaves unsaid. An agent has no such memory to lean on. Whatever meaning a human would infer, recall, or look up on the fly has to be written directly into the entry itself, because at query time the entry is all the agent has. That turns the dictionary from a reference document someone consults when stuck into a working memory substitute the agent depends on for every single answer it gives.
Why the standard six-component dictionary falls short for agents
The standard data dictionary was never built with this reader in mind, and it shows. Most conventional entries cover six components: technical specifications like data types, field names, nullability, length, precision, allowed values, and defaults; validation rules and constraints such as primary keys, foreign keys, and check constraints; descriptive statistics and quality indicators like range, distribution, completeness, and cardinality; entity-relationship diagrams and lineage; ownership and governance metadata including the data owner, steward, subject matter expert, and approval status; and usage metadata like access frequency, popular queries, and downstream dependencies. Atlan's data dictionary reference treats these six as the recognized standard, and they map almost entirely to what a database administrator or analyst needs, not what an agent needs.
A concrete illustration: a field documented as "cust_id INTEGER(10), primary key, not null". That entry is syntactically complete. The standard format assumes the reader already carries business context in their head, already knows what the company's metrics mean, and can infer relationships the schema never spells out.
What an entry ready for automated use must contain
The first is semantic business context. A field needs a plain-language definition that encodes what it means to the business, beyond what data type it holds. "Active customer" has to be defined the way the business defines it, not reverse-engineered from a join condition. Entity tags and ontology links let an agent place a field in the right conceptual category on its own, rather than guessing from context clues.
The second is metric definitions with full logic. The entry has to specify not just what a field is, but how any metric derived from it gets calculated, including every join, filter, exclusion, and edge case involved. ISG Software Research analyst David Menninger points out that most semantic models today are limited to SQL definitions, and SQL isn't rich enough to capture all the logic that goes into a real business model. A revenue figure that looks straightforward in a dashboard often hides several judgment calls about what counts and what doesn't, and an agent that doesn't have those judgment calls spelled out will make its own, silently.
The third is relationship maps that go further than foreign keys. An agent needs to know not just which tables relate to each other, but what those relationships mean in business terms, which joins are actually valid for which kinds of questions, and which combinations of fields produce sensitive outputs that require extra controls. A map that stops at the foreign key level misses the risk that matters most, because sensitivity often emerges only when fields get combined, not when any single field is viewed in isolation.
The fourth is access scope expressed at the field level: who can read a given field, under what conditions, and whether that access gets evaluated at query time under the actual end user's identity. A service account's broad permissions aren't a safe stand-in for this, because an agent is often generating queries on behalf of many different people, each of whom should see something different.
Two more dimensions round out the anatomy. Field-level lineage traces specific fields, transformations, and downstream consumption events, beyond table-to-table provenance. Without it, a team can tell a dataset has a problem but can't locate where the problem started. And continuously measured quality scores for completeness, freshness, accuracy, and consistency need to be attached live to each entry rather than logged once during an annual audit, because bad data fed into an agent doesn't produce a visible error. It produces a confident wrong answer.
The semantic layer as the mechanism that makes these entries work at runtime
None of this matters if the system can't deliver it to the agent at the exact moment it's generating a query. A well-built dictionary entry sitting in a wiki somewhere does nothing for an agent that can't retrieve it in time. The entry and the semantic layer are complementary parts of the same system, not two separate solutions to two separate problems.
A semantic layer is a single, governed set of business definitions, covering metrics, dimensions, and the logic behind them, that sits between a data warehouse and everything reading from it. For years that meant BI dashboards. The shift now is that it includes AI agents too. Without that layer in place, an agent pointed straight at raw tables has to re-derive metric definitions and join paths on every single prompt, which produces inconsistent answers across queries and opens the door to data leakage wherever access rules aren't enforced inside the layer itself.
Snowflake VP of AI Baris Gultekin put the underlying issue this way: "Agents need shared meaning and context, not just shared data. If an AI agent doesn't understand the business definitions behind the data powering it, its reasoning breaks down quickly." The industry's build-out reflects that: dbt Labs renamed its flagship conference from Coalesce to dbt Summit (September 15–18, Las Vegas) and rebuilt the agenda around AI-ready data, with the dbt Semantic Layer paired with MetricFlow and MCP as the headline product news, making governed metric definitions the explicit interface agents use to read a business. There's also a standardization push underway: the Open Semantic Interchange, formed in September 2025 by a group including Salesforce and Snowflake, is attempting to set common ground for semantic modeling, though Menninger cautions that semantic modeling still has a way to go before it reaches the ubiquity MCP has achieved, and the consortium's work may or may not end up as an adopted standard.
What breaks when entries are missing these dimensions
Missing any of these dimensions produces predictable failures that compound over time.
Metric drift is the first one. Without a shared metric definition living in the entry, different agents or sub-agents each re-derive KPIs on their own. In a multi-agent system, small definitional differences stack up across chained queries into serious output drift, where two agents answering the same business question give completely different numbers, each with total confidence.
Entity fragmentation follows close behind. Without a single governed record for a customer, patient, or product, agents can't reliably tell whether two references point to the same entity. Aggregation or deduplication tasks, the kind agents get handed constantly, expose that failure immediately.
Stale answers cause a quieter kind of damage. A data snapshot that's perfectly fine lag for a weekly dashboard is already too old for an agent making an operational call in real time. Governance violations from field combination are the hardest to catch of all: sensitivity that isn't flagged at the individual field level can emerge only once fields are joined, and an agent that can't evaluate that at the point of combination produces an output that violates access policy without tripping any error.
The 2026 CIO Customer Data Readiness Report, commissioned with UserEvidence and drawing on senior IT leaders at large global enterprises, traced the root of these failures to a consistent set of causes: insufficient data quality or availability, inconsistent governance across tools and regions, and limited visibility into how data can actually be used across platforms. Promethium's agent readiness guide gives a sense of what "ready" looks like in concrete terms: a data catalog covering at least the majority of metadata for high-value datasets, no competing business glossaries with conflicting term definitions, null rates under 1% for required fields in agent-consumed data, and duplicate rates under a fraction of a percent for unique identifiers. Those numbers aren't a checklist to memorize so much as a picture of the distance most organizations still have to close.
How organizations are building agent-ready entries in practice
A handful of platforms are already shipping infrastructure built around this exact problem, each adding a governed context layer on top of what already exists rather than asking anyone to migrate off their current warehouse.
Informatica from Salesforce announced at Informatica World 2026 on May 20 that it had become the first enterprise data management platform to deliver fully headless data management, exposing every data management capability as a reusable, governed service that any AI agent can invoke instantly, with zero code and no architecture rebuild, through native MCP support. Its data governance and catalog capabilities, published as MCPs, now live inside the Agent Fabric Context Catalog, described as the industry's first single destination to discover, govern, and operate both enterprise data and agents.
Databricks took a similar approach from a different angle. With AI Gateway now generally available and MCP support in Databricks Marketplace in public preview, enterprises can govern every model and external MCP server through Unity Catalog, with all model activity logged, rate-limited, and auditable. The MCP Catalog integrates with Agent Bricks to centralize discovery, access control, and auditing, so every agent action stays governed and compliant no matter where it runs.
Snowflake unveiled sharing of Semantic Models, in private preview, at Snowflake Summit 2025, letting users bring AI-ready structured data into Cortex AI apps and agents, sourced from internal teams or from third-party providers including CARTO, CB Insights, Cotality powered by Bobsled, Deutsche Börse, IPinfo, and truestar.
Atlan approaches the same problem through automation at the lineage level. Its system maps column-level provenance automatically, and its Context Agents write descriptions, definitions, and quality scores at better than a 90% acceptance rate by Atlan's own reporting, producing what it calls an Enterprise Data Graph that functions as context infrastructure every AI agent reads and every human searches.
Different companies, different entry points, same underlying pattern: governance enforced at runtime, semantic context delivered where the agent actually works, and nothing torn out to make room for it.
Scaling entries into agent-consumable assets through a data products marketplace
A single, beautifully built entry still isn't worth much if nothing can find it. Distribution requires its own layer, separate from authoring. A data products marketplace is a centralized, self-service platform that lets an organization discover, access, and use trusted data products, and it's the mechanism that turns individual dictionary entries into something agents can pull from at scale, rather than something only a human analyst clicking through a catalog will ever find.
Alation's Data Products Marketplace certifies every data product against an organization's own standards for quality, ownership, and compliance before it gets published, so only trusted, approved products surface. Access requests go through the marketplace directly, whether the one asking is a person or an agent.
Entropy Data takes a contract-first approach to the same goal, offering a marketplace enforced by data contracts and built on open standards, natively supporting the Open Data Contract Standard and the Open Data Product Standard, designed to serve human analysts and AI agents from the same governed surface. The contract is what makes the arrangement enforceable rather than aspirational: it encodes the quality thresholds, ownership, freshness requirements, and access scope the entry defines, and it makes those commitments machine-readable, so an agent can check them at the moment of access instead of discovering a violation after the fact.
The marketplace completes the entry's lifecycle, and authoring a rich entry is necessary but not sufficient. An agent still needs a governed way to find that entry, confirm it's trustworthy, and pull it at the moment a query demands it, and that's precisely what the marketplace layer is built to do.
The practical objection: semantic layer work is chronically underfunded
Everything described above requires real, sustained investment, and the obstacle standing in the way is organizational and economic. Building a rich, governed entry for every field in a warehouse takes time from people who could be shipping features that show up on a roadmap. Semantic layer work produces no visible output of its own. Nobody opens a dashboard and sees "semantic layer" as a line item. An agent in production gives a confidently wrong answer because a metric definition was missing or a field's sensitivity was never flagged, and that is the moment the investment becomes visible.
That's a hard case to make to anyone who funds budgets on a quarterly cycle. Pre-competitive infrastructure, the kind that benefits every team that touches the data rather than any single product line, tends to lose out to work with a name attached to a launch date. The pattern holds across the platforms described above: none of them asked anyone to solve this alone from scratch. Each one built shared, governed infrastructure and made it reusable across every agent and every team that touches it, which is the only way this kind of work gets funded in practice, because a single team rarely has the incentive or the budget to build it by itself. The organizations making real progress are treating the semantic layer the way they'd treat any other piece of core infrastructure: funded centrally, maintained continuously, and judged not by what it produces on its own but by how many failures downstream it quietly prevents.


