Executive perspective
Executive perspective
Meaning depends on context. A field such as status = closed may describe an administrative action, not a successful customer outcome. Definitions, timestamps and workflow stage determine what a record can legitimately demonstrate.
Labels can embody scarce knowledge. Specialist decisions and adjudications may be harder to reproduce than the source documents, but their quality depends on consistent rules and auditable origins.
Outcomes add evaluative power. Linking decisions to subsequent results can support testing and analysis, provided future information is not exposed to models making earlier decisions.
More metadata is not automatically better. Its incremental benefit must justify reconstruction expense, privacy risk, contractual constraints and the requirements of a specific buyer task.
1. From stored facts to interpretable evidence
An enterprise rarely collects operational data with future AI licensing in mind. Its systems capture the information needed to complete transactions, manage staff, meet reporting obligations and resolve problems. The resulting records are convenient for day-to-day operations, but they are not necessarily intelligible outside the organisation. A field called value might mean transaction amount, risk grade, priority, elapsed hours or a category number; closed might mean paid, rejected, transferred, written off or simply no longer assigned to an employee.
The original record therefore describes only part of the potential information asset. Its context includes field definitions, units, original source, event time, business process, decision-maker role, policy version, subsequent events and the distinction between observed facts and human opinions. It is this surrounding structure that allows another analyst to interpret the record rather than guess. The original chapter's chain, record → definition → provenance → decision → outcome, is an effective shorthand, but not every dataset needs every component. What matters is which context is essential for a defined use.
Consider a warranty-claims system. A file containing 500,000 photographs of failed components could be useful for visual classification. If the same images can be linked to model number, manufacturing batch, operating conditions, technician assessment, replacement decision and recurrence, they could also support fault diagnosis or the evaluation of repair recommendations. These are different tasks with different evidential needs. The richer dataset is not intrinsically better for every buyer: the extra information may be irrelevant to image recognition, difficult to verify or restricted by customer agreements.
Dataset-documentation research has long identified this problem. Datasheets for Datasets recommends recording motivation, collection, composition, processing, recommended uses and limitations; Google's Data Cards work similarly emphasises documentation of decisions and assumptions that shaped a dataset [1,2]. These approaches do not assign an economic price to metadata. They explain why a buyer can make a better-informed assessment of suitability, limitations and risk when that metadata exists.
Exhibit 1. The context ladder: how the same record changes in usefulness
| Evidence available | Illustrative claim record | What it can support | What remains unknown |
|---|---|---|---|
| Raw fields | status = closed; value = 4 | Basic counting, if fields are interpreted correctly | Meaning, unit, time, outcome |
| Defined record | value = severity grade 4/5; closed means workflow ended | Consistent categorisation | Who decided, why, under which rule |
| Provenance and decision | Claim ID, timestamps, source, adjudicator role, policy version, action | Audit, contextual search, decision-pattern analysis | Whether action achieved the intended result |
| Verified outcome | Follow-up inspection, independent test, appeal or recurrence window | Defined outcome evaluation and longitudinal analysis | Counterfactual result under a different decision |
Interpretation: Each layer broadens possible tasks. It does not establish better model performance, legal entitlement or a higher licence fee. In particular, administrative closure should never be silently relabelled as successful resolution.
2. Definitions, provenance and relationships make records usable
Context has at least three distinct forms. Semantic context states what a field means, including units, allowable values and changes in definition. Provenance describes how a record was created, transformed and attributed. Relational context connects records across events, actors and systems. These functions overlap, but should be tested separately. A field can have a precise definition yet an unknown source; a case can have excellent provenance but no trustworthy link to its outcome.
Industry standards offer useful methods without requiring every business to implement a complex knowledge graph. The World Wide Web Consortium's PROV model distinguishes entities, activities and agents and describes relationships such as was derived from and was generated by 3. W3C's DCAT version 3 supports interoperable descriptions of datasets and versions 4. A smaller operating company may implement the essential principles through a documented schema, data dictionary, transformation log and stable internal identifiers. Adopting a formal ontology is an engineering choice, not a prerequisite for demonstrating value.
Versioning is particularly important. A risk classification of “high” might have meant a probability threshold above 15% in one year and above 8% in another. Without the effective date and scoring-policy version, a model may learn inconsistent patterns from ostensibly identical labels. Historical data can remain useful when such differences are documented; when the definitions have been lost, joining longer histories may introduce noise rather than information. This is a specific application of the comparability problem examined in Chapter 4.
Relationships create both opportunity and risk. A warranty claim may reside in the service system, its parts replacement in inventory software, and the follow-up inspection in a quality-management platform. Stable case keys and timestamps can connect the three. Fuzzy matches based on customer names or descriptive text can generate false associations, especially where names are common, records were merged or different incidents occurred close together. A data export should distinguish confirmed joins from probabilistic ones and preserve join-confidence information where relevant. A buyer needs the denominator as well: the share of source records that cannot be linked should not disappear from the documentation.
Provenance is also about transformation. If original complaint narratives are summarised, identifiers are replaced, severity categories are standardised, or dates are generalised for privacy, those operations should be logged. Otherwise a buyer may mistake a processed description for a contemporaneous observation. In its generative-AI risk-management profile, NIST recommends recording data origins, content lineage and transformations for documentation and evaluation 5. This is governance guidance, not proof that documented data will command a particular price.
3. Expert labels are potentially valuable, but not unquestionable truth
Operational businesses routinely generate judgments that are difficult to infer from a raw document: whether a manufacturing defect is critical, whether a payment exception is legitimate, whether a complaint merits escalation, whether a code change is acceptable, or whether a claim falls within a contractual exclusion. These labels may represent expensive specialist attention. Their potential value is not simply that a human entered them, but that they encode a task-specific distinction which a model or evaluator would otherwise need to obtain elsewhere.
The quality of such labels varies. A customer-service priority category may be chosen to manage queues rather than to represent intrinsic severity. A credit-risk rating may combine quantitative analysis, current policy and committee judgment. Two experts may legitimately disagree where criteria are ambiguous. An adjudication can also be overturned. The dataset should therefore distinguish the initial label, reviewed label, final decision and, where observable, subsequent outcome. These cannot safely be collapsed into a single supposed ground truth.
Label quality depends on a documented rubric, who performed the classification, when it occurred, what evidence was available and whether independent review confirmed it. Where disagreement is material, recording multiple assessments and adjudication procedures may be more informative than suppressing uncertainty. An audit sample can estimate disagreement, error and missingness in the proposed product before large-scale extraction begins. The sample design should represent routine cases as well as difficult exceptions; otherwise a pleasingly high agreement rate on easy cases may conceal weak performance exactly where the data is most distinctive.
The value of expert labels is also use-dependent. A buyer developing a classification model may need a consistent target and sufficient examples by category. An evaluation provider may require withheld cases and reproducible scoring criteria. A retrieval system may only need reliable tags that help staff locate precedents. MLCommons' DataPerf initiative, which benchmarks data-centric machine-learning methods, provides research evidence for evaluating data with respect to specific tasks rather than treating dataset size as a universal quality measure 6. But none of these publications establishes a general market rate for a labelled claim, audit or case review.
Exhibit 2. Annotation evidence required for different buyer tasks
| Proposed AI use | Most useful contextual information | Main test | Material failure mode |
|---|---|---|---|
| Classification training | Label definitions, class balance, reviewer policy, sampling history | Performance against independently checked labels | Labels proxy operational convenience rather than task truth |
| Evaluation benchmark | Locked cases, agreed scoring rubric, adjudication evidence | Repeatable scoring without train/test contamination | Outcomes leaked into decision-time inputs |
| Retrieval or specialist assistance | Source, document section, timestamp, version, expert tags | Relevance and correctness of cited precedents | Obsolete or unauthorised content retrieved |
| Decision analysis | Initial assessment, alternatives, later verified outcome | Appropriately controlled comparisons | Historical selection effects mistaken for causation |
Interpretation: The same labelled archive can support several products, but each demands a different evidence standard. Additional labels should be commissioned only after the intended test is clear.
4. Outcomes improve evaluation, provided time and causality are respected
A recorded recommendation becomes more informative when there is evidence of what followed. A service case may have been reopened, a component may have failed again, or an invoice exception may have been reversed after audit. Such information enables a different question from “what did an employee decide?”: namely, “was the objective met under the specified observation window?” This is potentially useful for evaluating AI assistance and ranking candidate actions, even when a complete workflow reconstruction is unavailable.
A crucial design rule is to separate features available at decision time from later labels used for assessment. If a model is meant to predict which complaints will reopen within 30 days, the eventual reopen flag cannot be included as an input. Nor should a post-resolution summary written with hindsight be presented as though it was available to the original handler. Data leakage can make a benchmark look excellent while measuring nothing useful about future decisions. Evaluation specifications should record the prediction cut-off, observation period, exclusions, split method and what constitutes a verified result.
Outcomes must also be interpreted cautiously. A claim paid after appeal does not automatically prove the initial assessment was negligent; a support issue that never reopens may mean the customer disengaged. Where the outcome is missing, it may reflect non-response rather than failure or success. Periods with shorter follow-up naturally contain more unresolved outcomes. The buyer should understand these censoring and selection mechanisms, especially in domains where final results emerge only after months or years.
Nor can outcome-linked records, by themselves, establish which historical action caused success. Employees may route difficult cases to specialists, so the specialists' observed resolution rate reflects case complexity as well as skill. A model trained to imitate historical decisions can reproduce those patterns without knowing what would have happened under alternatives. Causal claims require an appropriate experimental or quasi-experimental design. A labelled archive can nevertheless support descriptive analysis, supervised training and carefully specified evaluation where the limitations are explicit.
Documented outcomes can increase the defensibility of a dataset, but the link itself must be demonstrated. A buyer should be able to inspect a sample and trace each label to a defined event, external confirmation or approved adjudication procedure. If verification requires reading unrestricted customer messages, new extraction controls or a narrower product may be necessary. More context is useful only if it survives both the evidential and permissions tests.
5. Quantifying the preparation economics of context
The commercial question is not whether context is philosophically valuable. It is whether contextual enrichment produces enough incremental usefulness to justify the cost of creating a reliable, legally usable product. A seller might have millions of records but only a small proportion with trustworthy cross-system links and observed outcomes. Even when those records exist, a buyer could prefer a lower-cost alternative, collect its own task examples or decide that enriched metadata adds little to the performance of its particular system.
An initial cost model should therefore distinguish source volume from candidate yield. The relevant unit might be a complete claim episode, a validated diagnosis or an independently adjudicated decision, not a row. Sellers should calculate extraction and joining costs, expert verification, de-identification, rights review, documentation, secure transfer and ongoing support. They should separately test whether a prospective buyer can use the resulting evidence, preferably through a permitted, bounded evaluation. This is more defensible than applying a generic currency amount to each field or label.
Exhibit 3. Illustrative economics of enriching a claims archive (hypothetical only)
Assume a company has 100,000 archived claims. The percentages are sequential assumptions, not industry statistics or observed rates. The seller considers a rights-cleared pilot and allocates £48,000 to extraction, cross-system joining, specialist QA, privacy/security controls and documentation.
| Case | Defined and linkable | Verified outcome among survivors | Permitted, complete and unique among survivors | Candidate episodes | Cost per candidate episode |
|---|---|---|---|---|---|
| Constrained | 50% | 50% | 50% | 12,500 | £3.84 |
| Base | 75% | 70% | 60% | 31,500 | £1.52 |
| Strong | 90% | 85% | 80% | 61,200 | £0.78 |
Formula: Candidates = 100,000 × linkage × outcome verification × permitted/completeness yield. Cost per candidate = £48,000 ÷ candidates; values rounded to two decimals.
Interpretation: Additional context can improve the quality of the usable denominator while raising preparation expenditure. These figures describe seller costs, not transaction prices, expected revenue or buyer demand. In each scenario, product viability would still depend on legal feasibility, a buyer's task, demonstrated incremental performance, negotiation and further costs.
A useful distinction is between the cost of recovering context once and the cost of maintaining it. The first may include one-off work to reconstruct past labels, align legacy schemas and repair unmatched identifiers. The second involves continuing data-dictionary governance, quality sampling, new policy versions and secure buyer support. A one-off archive may tolerate expensive historical reconstruction if it produces uniquely useful evidence; a recurring data feed needs stable automated controls and a sustainable refresh budget. Contracts should specify what happens when source definitions change, rather than promising an immutable product drawn from evolving operating systems.
A disciplined pilot compares a base extract with a context-enriched version on the same agreed task. For instance, compare search precision or adjudication accuracy using independent test cases, recording staff effort and privacy constraints. A measured improvement might justify enriching more records; a negligible change might justify stopping. For some uses, basic definitions and provenance will suffice. For others, outcome linkage is the critical asset. Context should be treated as a targeted investment, not a mandate to build a complete representation of the entire business.
6. Governance and practical implications for management
Context is not automatically safe to share. Cross-system linking can reveal identities, employment histories, commercially sensitive relationships or previously separated information. Replacing direct identifiers with pseudonyms can preserve episode-level relationships, but pseudonymised personal data remains subject to UK data-protection law. The Information Commissioner's Office explains that data minimisation requires personal data to be adequate, relevant and limited to what is necessary; it also stresses that pseudonymisation is not the same as anonymisation [7,8]. UK purpose-limitation guidance requires organisations to consider whether a proposed new use is compatible with the original purpose and to establish an appropriate lawful basis 9.
The practical balance is not between retaining everything and deleting every identifier. It is between a buyer's legitimate evidential needs and the narrowest authorised disclosure. Techniques may include randomised case tokens, separately controlled matching keys, suppression of rare combinations, limited field access, aggregation, derived task descriptions and secure evaluation without a broad raw-data export. Removing links that are unnecessary for the stated buyer task may improve both privacy and efficiency. Retaining links without an authorised purpose may undermine the transaction entirely.
Contractual and intellectual-property diligence is distinct from data protection. A business may control the database infrastructure but lack permission to sublicense certain customer submissions, supplier content or employee-created material under particular agreements. The UK Intellectual Property Office distinguishes database rights from copyright in a database's selection or arrangement 10. This is a reason to inspect provenance and the rights chain, not an assurance that any database can be licensed freely. Buyer-specific transfer restrictions, confidentiality obligations and permitted model uses should be addressed before a pilot leaves the organisation.
For executives, the most useful starting point is a context inventory, not a wholesale data export. Pick one economically meaningful process, identify the task a prospective AI developer might want to train or test, and map the records required to make a small sample interpretable. Appoint data owners to document definitions and policy versions; check the presence of stable joins; distinguish expert opinion from verified outcomes; quantify missingness; and obtain legal and security screening. Only then prepare a sample and test it against a concrete use case.
Exhibit 4. Management diligence checklist for a context-rich data product
| Question | Evidence to request | Decision consequence |
|---|---|---|
| What exact buyer task is contemplated? | Task statement, test metric and minimum required fields | Reject enrichment unrelated to task |
| Can records be interpreted consistently? | Dictionary, schema and policy version history | Limit scope or correct ambiguity |
| Can key relationships be proved? | Join audit, unmatched rate, lineage sample | Exclude unverified associations |
| Who supplied labels and outcomes? | Rubrics, reviewer roles, audit samples, follow-up windows | Distinguish opinion, decision and verified result |
| Is the proposed use permitted and secure? | Contract review, lawful-basis analysis, minimisation and access design | Stop, redesign or narrow the pilot |
| Does enrichment change measured usefulness? | Side-by-side buyer evaluation and full preparation budget | Scale only where incremental benefit is demonstrated |
Interpretation: A credible proposition is not “we possess rich metadata”. It is “for this permitted purpose, we can supply an interpretable, tested body of evidence with known limitations at a defensible cost”.
Practical conclusion. In a data inventory, count relationships and definitions alongside records. The commercially meaningful product may be a modest collection of carefully documented, outcome-linked examples rather than a massive archive. Conversely, a rich context layer may have greater value for internal audit, quality assurance or retrieval than for third-party model training. Management should preserve that option: failing the external licensing test does not make the underlying governance work worthless.
Sources and further reading
- Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J.W., Wallach, H., Daumé III, H. and Crawford, K. (2021), ‘Datasheets for Datasets’, Communications of the ACM, 64(12), pp. 86–92. Microsoft Research publication record. https://www.microsoft.com/en-us/research/?p=563964
- Google Research / Google for Developers, Data Cards Playbook: Transparent documentation for responsible AI. https://developers.google.com/learn/pathways/data-cards-playbook
- W3C (2013), PROV-O: The PROV Ontology, W3C Recommendation. https://www.w3.org/TR/prov-o/
- W3C (2024), Data Catalog Vocabulary (DCAT) – Version 3, W3C Recommendation, 22 August. https://www.w3.org/TR/vocab-dcat-3/
- NIST (2024), Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, especially MAP 2.1 and MAP 2.2. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- Mazumder, M. et al. (2022), DataPerf: Benchmarks for Data-Centric AI Development, arXiv:2207.10062 (research preprint; see MLCommons project). https://arxiv.org/abs/2207.10062
- Information Commissioner's Office, Principle (c): Data minimisation (guidance; under review following 2025 legislative changes as stated on the ICO page). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-protection-principles/a-guide-to-the-data-protection-principles/data-minimisation/
- Information Commissioner's Office, Pseudonymisation (guidance; review notice on ICO page). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
- Information Commissioner's Office (updated 23 March 2026), Principle (b): Purpose limitation. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-protection-principles/a-guide-to-the-data-protection-principles/purpose-limitation
- UK Intellectual Property Office, Sui generis database rights. https://www.gov.uk/guidance/sui-generis-database-rights