AI-ready data
FeaturesLong read

Data Virtualization Tools Compared for AI Workloads

Most AI projects stall at the data layer, not the model.

Correspondent · · 10 min read
Features · September 30, 2026 · 10 min read · 2,152 words

Enterprise AI spending hit record levels in 2025, and most of it didn't work. Not "underperformed expectations," it delivered no measurable business value, full stop, per research compiled by Pertama Partners. MIT Project NANDA found 95% of generative AI pilots show no measurable P&L impact. Gartner expects more than 40% of agentic AI projects to get canceled by the end of 2027, the majority of the market failing to do the one thing it was funded to do. Those are the majority of the market failing to do the one thing it was funded to do. That's the majority of the market failing to do the one thing it was funded to do.

The instinct is to blame the model. Wrong instinct. Transcend's survey of senior IT and business leaders at enterprises with 5,000+ employees found 81% had at least one AI initiative delayed, scaled back, or abandoned in the past year. Of those, 93% hit permission and governance problems somewhere in the AI lifecycle. Engineering hours mostly go to data repair and governance workarounds, not building features. Teams hired for AI spend most of their time fixing plumbing.

The Cloudera and Harvard Business Review Analytic Services report puts a finer point on it: only a small fraction of enterprises say their data is completely ready for AI. Per Transcend's State of Customer Data survey, 81% had at least one AI initiative delayed, scaled back, or abandoned in the past 12 months. Only a minority of IT decision-makers rank governance a top priority, yet far more point to governance approval as where initiatives stall or reverse. Most of these organizations didn't define AI governance until systems were already live or nearly so. They built the car before deciding who gets the keys.

Put those numbers together and a pattern falls out. The bottleneck isn't the model's reasoning, its context window, or its training data. It sits one layer down, at the data layer itself: whether the right data reaches the right agent, at the right freshness, under the right permissions, with the right meaning attached. Data virtualization is the technology built to sit exactly at that layer. The question is whether it was built for this particular job, or a different one.

What data virtualization was originally built for

Data virtualization gives an organization one unified way to reach scattered data without copying it anywhere. A virtual layer sits atop source systems, sends queries to wherever data lives, and stitches results together on the fly. No replication, no waiting for an overnight batch job to land in a warehouse.

Three pieces do the work: connectors for each source, a federated query engine that splits requests into parallel sub-queries, and a semantic layer giving it consistent shape. A query fans out to Oracle, Snowflake, Salesforce, or elsewhere, and returns assembled, without the requester needing to know which system answered which part.

Implementations generally fall into three types. Pure virtual access stays fully live, with every query hitting the source in real time. Hybrid setups virtualize some data and cache or ingest other data, based on what performance actually demands. Materialized views take the data that gets hit constantly and store it in a pre-optimized format, trading a bit of freshness for speed.

None of this was designed in a vacuum. The original target customer was a centralized IT team supporting human analysts. Those analysts ran queries on a known schedule, tolerated some latency, and worked inside governed BI toolchains that IT had already blessed. A well-built DV platform can answer for last quarter's regional sales regardless of whether that number lives in a flat file, a Salesforce object, or an old Oracle table. The unified model handles the abstraction.

But it solved for access: a person asks, the system answers, a person interprets and decides what's next. AI agents don't work that way. They don't query on a schedule, don't tolerate ambiguity like a trained analyst, and increasingly act on data rather than just reading it. The next section spells out exactly how different.

What AI agents demand from a data layer that human analysts never did

An analyst who hits a vaguely named column like cust_stat_cd can ask someone on the data team what it means. An AI agent can't do that. It has no institutional memory, no colleague to tap, and no instinct to pause and double-check. A stale or partially wrong answer from an agent looks exactly as confident as a correct one. There's no hesitation in the output to tip anyone off.

That gap defines seven concrete requirements distinguishing an AI-ready data layer from a conventional DV platform.

Semantic context has to be built in, not inferred. Table descriptions, metric definitions, and entity relationships must be handed to an agent directly, since guessing from a schema name is how errors start. Every metric needs exactly one agreed definition. When analytics engineering, BI tools, and AI copilots each maintain their own metric logic, any change's cost multiplies across systems, and the same question can return divergent answers.

Lineage must run end to end from source to the agent's final answer, capturing identity and intent, not just a table-level note that data passed through. Freshness must match the use case, not a one-size-fits-all schedule: a snapshot fine for a dashboard may be stale for an agent making a live decision. Access must be revocable per workload, not tied to a shared credential that breaks other things downstream if pulled. And agents need structured, queryable answers from structured data. Forcing structured questions through retrieval-only systems, instead of queryable governed data, produces confidently wrong results.

Gartner's Data & Analytics conference named this failure mode: the ungoverned semantic layer as the new ungoverned data lake, the same access-without-governance mistake one layer higher. At dbt Summit 2026, the semantic layer was reframed from a BI nice-to-have into AI infrastructure: agents need governed meaning, definitions, tests, and lineage, not just open access to tables. Most apparent AI hallucinations are actually semantic failures. The model picks the wrong table, joins at the wrong grain, or aggregates something it shouldn't have, a plumbing failure wearing a reasoning failure's clothes. That's not a reasoning failure. That's a plumbing failure wearing a reasoning failure's clothes.

The semantic layer's place in this picture and its lagging governance

A semantic layer and a data virtualization platform do different jobs, and conflating them is how organizations end up with one but not the other. The platform gets an agent to the data. Per MIT Sloan, the semantic layer is between data and its users, explaining what data represents, how assets relate, and what rules govern its use. Access without meaning gets a confident wrong answer fast, while meaning without access gets no answer. Neither one substitutes for the other.

The uncomfortable part comes next. Gartner's conference data shows a growing share of organizations already implemented a semantic layer in 2025, with more planning to by 2027. Sounds like progress. Yet only 14% say their data is confidently secured and governed. Adoption is climbing a lot faster than governance is catching up. That gap is the whole problem in miniature: enterprises are wiring meaning into their systems faster than they're deciding who gets to see what, and faster than they're checking whether that meaning is even correct.

Gartner analyst Anthony Mullen said that by 2030, universal semantic layers will rank alongside data platforms and cybersecurity as non-negotiable infrastructure that gets automatic budget. That's a big claim, but the failure pattern already visible today backs it up. The most common failure mode is the same metric defined separately across analytics engineering, BI, and AI copilots. A governed semantic layer catches most such errors before a query runs. An ungoverned one scales these errors, repeating wrong logic across every agent that touches it.

So what actually counts as "AI-ready" for a semantic layer, specifically? A handful of things, and they're all checkable, not aspirational. One agreed definition per core metric, shared across every team and every agent that touches it. Tests and data contracts that run before anything downstream is allowed to read the data. A named owner attached to every definition, someone accountable when a number looks wrong. Freshness treated as a numeric service level, not a vague preference. And a cost model predictable as agent query volume climbs, not a surprise bill later.

How the leading data virtualization platforms handle AI workload requirements

Judged against those criteria, not against old-school benchmarks like connector count, the field splits pretty clearly by design center.

Denodo is the dominant pure-play in this market, having effectively defined enterprise data virtualization for roughly two decades. It's built for large enterprises with dedicated architecture teams favoring centralized, tightly governed logical data warehouses. Denodo's governance model is a logical data management approach that connects distributed data across hybrid, multi-cloud, on-premises, SaaS, and third-party environments without data movement. It requires real upfront architectural planning and semantic modeling, with rollout running weeks to months depending on source complexity.

Starburst runs on the Trino query engine, federating access across cloud platforms, data lakes, and traditional databases without moving data. It brings cost-based optimization, workload management, governance features, and integrations with Unity Catalog, Iceberg, and dbt. Its cloud-native design autoscales under high concurrency, which matters for agent-driven query spikes rather than scheduled human queries. It fits multi-cloud, lakehouse-first organizations best.

AtScale blends virtualization with semantic modeling aimed squarely at the "one metric, one definition" problem. Define a business metric once, and it's accessible across Excel, Tableau, Power BI, and Looker. It supports live query translation, pushdown optimization, and caching, so performance isn't traded for consistency, plus automatic lineage tracking and a self-adjusting query optimizer. Best suited to organizations that care most about one number meaning one thing everywhere it's read.

K2View organizes around business entities (a customer, an order, an account) instead of tables. AI automatically discovers sources and builds the entity-based model that becomes its semantic layer, letting agents reach data without understanding the underlying schema. Query execution runs at sub-second speed across distributed sources, with pull and push delivery via APIs, JDBC, streaming, and messaging. A strong fit for entity-centric, real-time, compliance-heavy environments.

TIBCO Data Virtualization centers on exposing governed, queryable views and reusable datasets across analytics, reporting, and applications, backed by centralized metadata controls and a self-service directory of already-governed datasets. It's built for large-scale enterprise integration across many source types, and that's where it earns its keep.

Snowflake took a different path in early 2026 with Semantic View Autopilot, moving from manual semantic modeling to AI-powered automated generation of semantic views. Because the layer is native to the warehouse, there's no separate infrastructure to stand up, and integration with Snowflake's query engine is tight. Semantic Views let teams define logical tables for entities like customers, orders, and products, then combine them into metrics and dimensions. The trade-off: tight native integration is an advantage on Snowflake, but multi-cloud organizations may find a standalone platform gives more flexibility.

Kyvos Insights applies AI-powered virtualization to build hypercubes supporting instant BI querying over very large datasets, suiting enterprises modernizing analytics and AI infrastructure together rather than sequentially.

Informatica strengthens the virtualized layer through governed data integration, metadata intelligence, and AI-powered automation, as one piece of a much broader data management platform rather than a virtualization specialist.

Delphix starts from compliance, automating data privacy handling for GDPR, CCPA, and HIPAA, and helping organizations move data from private to public cloud to speed both cloud migration and AI adoption.

Alongside these sits a newer category: platforms built ground-up as an AI-ready data layer, with semantic context and governance as core design, not an add-on. The design center here differs from legacy DV in a real way. Rather than requiring migration, they sit atop existing infrastructure, exposing one interface with semantic context and governance enforced at query time. This addresses the institutional memory gap: table descriptions, metric definitions, and entity relationships get embedded so an agent produces a correct answer, not just a valid-looking query. Permissions are enforced at query time under the requesting user's identity, and every query is logged with identity, intent, and lineage, not reconstructed later. For enterprises needing AI-ready access without ripping out an existing warehouse, this category fits that exact constraint. It is best fit for organizations already deeply invested in the Snowflake platform. It is best fit for high-volume analytics modernization alongside AI adoption.

The tools differ in shape, in what they modernize, and in how much infrastructure change they require. What they share is a common test: not whether they can move data, but whether an agent with no colleague to ask and no instinct to pause can trust what comes back. AI capabilities include zero-copy data access, unified semantics, centralized compliance, natural language search, and data marketplaces, designed to support agentic AI with accurate, governed data. It is best fit for organizations prioritizing metric consistency across BI and AI consumers.

Sources

  1. Exclusive report: The hidden force behind every stalled AI initiative | Transcend | The only real-time data governance and decision layer
  2. Only 7% of Enterprises Say Their Data Is Completely Ready for AI, According to New Report from Cloudera and Harvard Business Review Analytic Services
  3. 10 Things We Learned at the 2026 Semantic Layer Summit | AtScale
  4. Gartner D&A 2026: Where the Context Layer Became a Budget Line Item

More in Features