Too Much Data, Too Few Answers
The contemporary organisation rarely suffers from an absolute shortage of data. It suffers from a shortage of answers that can be produced with sufficient speed, semantic consistency, and evidential traceability. The distinction is not rhetorical. Data volume is a property of stored representations; an answer is an epistemic object generated under a question, a context, and a decision horizon. This essay argues that the central problem of a Data Platform is therefore not accumulation but computable intelligibility. It develops a technical model of the distance between raw observations and actionable answers, introduces the notions of semantic latency and decision friction, and shows why a platform should be evaluated by its ability to stabilise meanings, preserve provenance, and reduce the cost of asking the next question. The underlying thesis is simple: an organisation becomes data-driven not when it owns more data, but when it can explain how a relevant answer was produced and reproduce it under controlled conditions.
- — Semantics
- — Lineage
- — Platform Economics
- — Organisational Culture
1. The paradox of abundance
A hospital, laboratory, manufacturer, insurer, transport network, or public administration may generate millions of observations per day and still be unable to answer a modest operational question without days of preparation.
The paradox disappears once we distinguish storage from knowledge. A database can contain an abundance of inscriptions while the organisation remains epistemically poor. The mere presence of data does not imply that the data are identifiable, comparable, temporally aligned, semantically compatible, or authorised for the intended purpose.
Let the available data assets be represented by a set
where each source has its own schema, clock, identifiers, encoding conventions, error distribution, and institutional owner. Let a business or clinical question be . Producing an answer is not a direct lookup but a composite computation:
where selects relevant evidence, normalises representations, reconciles identities and meanings, and derives the answer. In weak data environments, these functions are repeatedly improvised by analysts. In a mature Data Platform, they are progressively made explicit, versioned, tested, observable, and reusable.
The operational complaint “we have the data, but cannot answer the question” therefore points to a missing computational institution. The organisation owns records but lacks a dependable machinery for transforming records into warranted claims.
2. An answer is not a datum
The philosophy of data is useful here because it resists a common category error. A datum is not a miniature fact waiting passively inside a table. It is a difference recorded according to a system of distinctions. Gregory Bateson’s famous formulation of information as a difference that makes a difference is technically productive: a value becomes informative only relative to a model in which the difference is significant.
Consider a numeric observation . Without unit, measurement method, reference interval, subject, time, and purpose, the value has almost no operational meaning. A table may preserve the lexical token while losing the conditions that allow it to function as evidence. Thus the information content relevant to a decision is not reducible to Shannon entropy. Shannon’s model deliberately abstracts from semantics; a Data Platform cannot.
We can express the effective usefulness of an observation for a question as
where is semantic completeness, temporal relevance, provenance confidence, and governance admissibility. Each term may be normalised to . If any critical term approaches zero, the observation may be technically present but practically unusable. This multiplicative form is deliberate: missing identity, unknown provenance, or prohibited use cannot always be compensated by high volume.
The platform’s task is therefore not merely to preserve values but to preserve the conditions under which values can be interpreted and legitimately recombined.
3. Semantic latency and decision friction
Traditional performance engineering measures ingestion latency, query time, throughput, or storage cost. These metrics are necessary but insufficient. An organisation can ingest a feed in seconds and still need a week to agree on what a field means. We need a broader metric: semantic latency.
For a question , let
where “accepted answer” means an answer whose definitions, evidence, and ownership are sufficiently agreed for action. Semantic latency includes discovery, access negotiation, identifier reconciliation, data cleaning, definition disputes, validation, and explanation. It is often much larger than computational latency.
Decision friction can be modelled as a weighted sum:
where is discovery effort, integration effort, semantic mediation, validation effort, and regulatory or access resolution. A Data Platform creates value when it reduces the expected friction across a distribution of future questions, not merely when it accelerates one predetermined report.
This is why one-off pipelines are deceptively successful. They minimise friction for a single known query while externalising maintenance and future ambiguity. Their local optimisation increases global entropy. A platform must instead reduce the marginal cost of the next legitimate question.
4. From repositories to epistemic machinery
A repository stores representations. An epistemic machinery organises how claims are produced from representations. This requires at least five capabilities.
First, identity must be explicit. Patient, customer, device, order, episode, shipment, or contract identifiers cannot remain informal assumptions. Second, temporal semantics must be represented: event time, processing time, validity time, and publication time are different variables. Third, transformations must be inspectable and reproducible. Fourth, quality rules must be attached to the data contract rather than applied as ad hoc filters.
Fifth, governance must determine not only who can read a table but which use of which attribute is legitimate in which context.
These capabilities transform the platform into what cybernetics would recognise as a regulated system. The platform observes sources, compares states with contracts, isolates deviations, and adjusts processing while preserving evidence of the adjustment. The system does not “know” in a human sense, but it materialises the institutional conditions for reliable knowing.
A useful quality function for an answer can be written as
where accuracy, freshness, completeness, explainability, and reproducibility are weighted according to the use case. No universal weighting exists. Emergency operations may privilege freshness; regulated reporting may privilege reproducibility and provenance; research may privilege completeness and methodological transparency. Architecture is partly the art of making these weightings explicit.
- 01. Too Much Data, Too Few Answers
- 05. The question as an architectural object
- 06. The politics of definitions
- 07. What success should mean
- 02. A Data Platform Is Not Another System to Use
- 03. The Data Platform as an Organisational Control Room
- 06. Multi-loop governance
- 07. Visibility, power, and ethical limits
- 08. The architecture of a credible control room
- 04. The Hidden Cost of Fragmented Data
- 06. Risk, audit, and the cost of explanation
- 07. The philosophical structure of fragmentation
- 08. Stable contracts as real options
- 05. Why Putting All Data in One Place Is Not Enough
- 06. Every Source Speaks a Dialect
- 06. Preserving the original expression
- 07. Late binding as translational ethics
- 08. Dialects evolve
- 09. The politics of the interlanguage
- 07. From Collecting Data to Producing Answers
- 08. The Business Value of a Data Platform: Efficiency, Quality, and Innovation
- 10. The counterfactual business case
- 09. Separating the Fact from the Local Format
- 10. Data Is Not the Table That Contains It
- 11. Dismantling the Source, Reconstructing Information
- 06. Idempotence and replay
- 07. Quarantine as a third state
- 08. Recomposition is not a return to the source
- 09. The philological analogy
- 12. What Is a Data Object?