AI-ready data
FeaturesLong read

What Combined Data Fields Reveal That Individual Fields Don't

Combining data fields reveals relationships and insights that single fields cannot deliver alone.

Correspondent · · 10 min read
Features · October 3, 2026 · 10 min read · 2,299 words

A hotel guest's note that they prefer a high floor sits in a reservation system as a scrap of free text. Alone, it is trivia, the kind of detail a front desk clerk might or might not notice. Paired with reservation data, that same note becomes something else entirely: a priceable preference, an attribute the hotel can act on, honor, and even charge for. One field describes a fact. Two fields, joined, describe a decision.

That pattern repeats everywhere data gets collected. Cenfri's work on dataset merging in agricultural finance found the same gap: counting how many farmers had registered with a program told almost nothing on its own. Only once registration records were merged with mobile money transaction data, showing which of those registered farmers actually transacted, did the numbers say anything about real financial inclusion. The registry said who signed up. The transaction log said who used anything. Neither file, read by itself, could answer the question that mattered.

Transit planning in Kigali follows the same logic. E-ticketing ridership numbers describe how many people boarded a bus. GPS coordinates from motos and buses describe where vehicles were at a given moment. Neither dataset, taken alone, can tell a city planner where to redraw a bus route. Combined, ridership and location data reveal where demand and supply actually diverge, which is the only information a route optimization decision needs.

The common thread is structural: it is about how data is organized. A field is a descriptor of state: it tells you what something is, what value it holds, what category it belongs to. It does not tell you how that value relates to any other value sitting in a different table, system, or document. Structured data answers what happened, numbers on a sales dashboard showing revenue dropped. It takes unstructured data, like the customer reviews that explain the drop came from a specific product change, to answer why. Piling on more fields of the same kind does not close that gap, because the insight was never going to live inside a single column. It lives in the space between columns, at the point where two values get read together.

Emergent meaning from combination

Combination does not just add information, it creates it. The signal produced by two fields read together exceeds whatever each field could yield on its own, because the pairing exposes a relationship, a rate, a pattern, or a contradiction that no single field was ever built to encode. A number next to another number becomes a ratio. A timestamp next to a location becomes a route. A complaint next to a purchase becomes a cause.

Researchers studying multimodal data have a name for the excess value this produces: inter-modality synergy, the predictive gain achieved by integrating multiple sources that goes beyond what the best single source could deliver alone. The gain is a different category of insight, one that didn't exist in either source before they were joined.

Pharmaceutical research shows this at a mechanical level. Study protocols, PDFs, data tables, and diagnostic scans, considered individually, are reference documents sitting in a file system. Preserving how they relate to each other, across compounds, across trials, across patients, through a data fabric layer turns them into something a researcher or a model can reason over. Flatten those same documents into plain text to feed a language model, and the structure that made the reasoning possible disappears with it. The documents are the same. What survives or collapses is the relationship between them.

Industrial equipment monitoring works the same way. A vibration sensor reading, a calibration log, a thermal image, and a note about ambient humidity each describe one dimension of a machine's condition. None of them, alone, diagnoses a failure. Unified through a semantic layer that tracks lineage and gives real-time governed access across all four, they let an engineer trace a fault back to its root cause, something no single sensor feed could support on its own.

Sigma Computing's research on data blending describes the same mechanism in healthcare. A hospital blending electronic health records with patient satisfaction surveys and operational scheduling data can see how wait times affect satisfaction scores, and which treatments produce the best outcomes. That's a causal pattern. It doesn't exist in the health records alone, and it doesn't exist in the satisfaction surveys alone. It exists only in the blend. Relationship reveals causality, and causality is what makes prediction possible. That's the chain combination builds, every time.

Sensitivity and risk as properties of combinations

The same mechanism that produces insight produces exposure. A field that looks completely harmless on its own can turn identifying, discriminatory, or commercially sensitive the moment it sits next to a second field, and that risk has no home in either field by itself. It only exists at the join.

This turns a familiar governance assumption on its head. Classifying sensitivity field by field only works if risk lives in fields, and it doesn't. Because risk lives in combinations, evaluating sensitivity one column at a time will systematically miss the risk that actually matters, since that risk only materializes once fields are joined.

The Rwanda agricultural finance case makes the stakes concrete. Merging the Smart Nkunganire System, the Smart Kuhangara System, One Acre Fund registration records, and mobile money transaction data, researchers working under Cenfri's Rwanda Economy Digitalisation programme found that nearly two-thirds of transacting farmers were male, and more than a fifth were 55 or older. Neither the agricultural registry nor the transaction ledger could produce that demographic picture alone. The registry had no transaction history. The ledger had no demographic data. Only the merge produced an inference with real policy weight, and real privacy weight alongside it. The sensitivity was created at the moment of combination, not before it.

That reframes what a sound governance model has to look like. An organization that classifies fields individually, grants access by field, and audits by field has built a system for a world where data gets consumed one column at a time. Analysts don't work that way anymore, and AI agents never did. Both consume data in joins, pulling multiple fields together to answer a single question. A governance model built around columns has no mechanism for catching the risk that occurs only when those columns meet.

How data inconsistency across sources undermines combination

None of the insight described above is reachable if the underlying data can't actually be combined cleanly, and inconsistency across sources is what most often stands in the way. The obstacle occurs at three distinct levels, and each one breaks combination differently.

Structural inconsistency is the most mechanical of the three. One system records a date as DD-MM-YY, another as MM-DD-YY, and when those two fields get merged, the result is a silent error, a date that's wrong without ever throwing a flag. The merge completes. The merge completes, the record looks fine, and the date is wrong anyway.

Semantic inconsistency runs deeper. Asking an AI agent to pull "revenue" across a finance system tracking recognized revenue, a marketing system tracking pipeline revenue, and an operations system tracking booked revenue produces three different numbers, with no way for the agent to know they disagree. No layer told it what "revenue" means for this particular business, so it has no basis for flagging the mismatch. Researchers call this the semantic gap, exactly the kind of gap a human analyst would catch on instinct and a model won't catch.

Definitional inconsistency is the root both of the other two usually trace back to. When a model's revenue figure doesn't match finance's own number, the problem is usually in the organizational decisions governing the data, not the raw data underneath. Nobody ever decided which system counts as authoritative, which of three overlapping customer records should resolve to the canonical one, or which definition of a shared metric applies in which context. That's an organizational decision nobody made.

Sigma Computing's research draws a useful line here between data blending and data integration. Integration consolidates everything into one centralized warehouse through ETL pipelines built in advance. Blending has to handle mismatches in structure, format, and granularity on the fly. Join conditions demand matching value types, matching formatting, and matching values before they'll work. A mismatch as small as "London" against "LDN" is enough to silently break a blend, with no error message to say so.

The underlying cause sits above the data layer. The ERP, the CRM, the financial planning tool, and the marketing platform that make up a typical enterprise stack were never built around one shared data model, one set of consistent field definitions, or one synchronized update schedule. Each system was built to solve its own problem. When a model needs to pull a consistent signal across all four, there's no clean seam connecting them, because nobody designed one.

What a semantic layer must provide for trustworthy combined fields

Solving inconsistency at the level of individual pipelines treats a definitional problem as a plumbing problem. What the inconsistency above actually calls for is a layer that carries definitions, not just data, and that works differently depending on who, or what, is asking.

Human analysts and AI agents need different things from that layer, because they fail differently. An analyst who hits a vague column name can ask a colleague what it means, notice when a number looks implausible, and lean on institutional memory to resolve the ambiguity. An agent can't do any of that. It resolves ambiguity probabilistically. A stale answer or a partial one comes back looking exactly as confident as a correct one. There's no hesitation in the output to warn anyone something's off.

Traditional semantic layers were built for static dashboards, checked by humans who could catch an error before it spread. Agents query dynamically and at volume, so the governance around them has to enforce deterministic definitions automatically, on every single query, not only when a human happens to double-check the result.

That requires more than a BI-era semantic layer carried forward unchanged. A semantic layer built for agents needs table descriptions, metric definitions, and entity relationships laid out explicitly, giving an agent the context to generate an answer that's actually correct and a query that's syntactically valid. It needs certified definitions with lineage attached: which system is authoritative for a given metric, which of several overlapping records of the same customer resolve to one canonical entity, which definition of "active customer" or "revenue" applies in this particular context. It needs permissions enforced at query time, under the real identity of the person the agent is acting for, rather than inherited from a broad service account that quietly erases the distinction between what the agent can see and what the human it's working for is allowed to see. And it needs audit logs that capture identity, intent, and lineage together, because at the volume agents operate, a log that records only which table got touched gives nobody a way to reconstruct what actually happened when something goes wrong.

The gap between a schema-only setup and one carrying this kind of context isn't marginal. Research on the subject shows answers go from mostly wrong to reliably correct when this context is present. That's a system that works versus one that doesn't, a structural gap rather than a tuning problem to be improved at the margins.

Why AI initiatives stall without accounting for combination

Most enterprise AI initiatives that stall don't stall because the model underperforms. They stall because the data feeding it was never designed to be combined, governed, or understood by anything other than a human reading one column at a time.

This pattern occurs well beyond any single failed project. Enterprises that have already woven AI into core functions, but never resolved cross-system data access, are operating under what researchers describe as an AI readiness illusion. AI is running. The signal it's consuming is partial, inconsistent, or ungoverned at exactly the point where fields get combined, and nobody built the layer that would catch it.

The pressure to fix this has moved up the org chart. Data quality problems that a human analyst used to quietly catch and correct, a wrong join, a stale definition, a mismatched format, don't get caught anymore once an agent is the one acting on the data. The agent acts at volume, without flagging the error, and the consequence lands as a business decision rather than a data bug. That's why this is now a C-suite concern rather than an IT ticket.

The standards that made a field "clean" for a human analyst were built for column-by-column inspection: consistent formatting, no missing values, a reasonable range. Those standards say nothing about what happens when that field gets joined to three others at query time, which is exactly the situation an AI agent is asked to reason inside of constantly. Sensitivity, permissions, lineage, and semantic meaning all have to be evaluated at the point where fields come together, not before it and not after. That makes a governed semantic layer a prerequisite for AI to function correctly on enterprise data, not an enhancement layered on once the basics are working.

The organizational form this takes is the data product: data packaged together with its context, its governance, and its quality assurances as one consumable unit, built once and reused rather than assembled fresh every time a query needs to cross a system boundary. An agent that queries a certified data product inherits the combination-level guarantees already built into it. An agent that queries raw fields across four disconnected systems has to assemble those guarantees itself, on the fly, with no memory of how it was done last time. One of those approaches scales. The other recreates the same inconsistency problem at every single query, forever.

Sources

  1. What we can learn from merging datasets
  2. The Key to Better Insights: Structured and Unstructured Data Together - T-Gency
  3. Multimodal Data Insights
  4. What Is Data Inconsistency? Causes, Examples, and Fixes
  5. Integrating Disparate Data Sources: Challenges and Solutions
  6. Semantic Layer Architecture: Components, Design Patterns, and AI Integration
  7. Why Most Generative AI Projects Stall at the Data Layer
  8. Council Post: Avoiding The AI Failure Zone: Why Context And A Unified Data Layer Matter

More in Features