Data Monetisation 101 / Section 1 / Chapter 3

Section 01 · Understand the asset · Chapter 03

What Makes One Operational Dataset More Valuable Than Another?

Data value is not measured by record count alone. This chapter explains how task relevance, independent examples, distinctive context, data rights and delivery costs determine whether an operational archive can improve an AI system and support a viable commercial transaction.

15 min read8 referencesSite last updated 9 October 2026

Executive perspective

Executive perspective

Buyer relevance comes first. Two businesses may hold equally large archives but offer very different value to an AI developer. What matters is whether the records improve a particular task against credible alternatives, not the number of rows they contain.

Independent evidence matters more than repetition. Case coverage, verified outcomes and interpretable context can be more useful than years of duplicated routine activity.

Differentiation needs a test. Proprietary content or uncommon expertise may improve a dataset's position, but only where the information is difficult to substitute and the seller can demonstrate a permitted use.

Commercial viability is a separate hurdle. Rights review, extraction, quality assurance and secure delivery can consume the benefit available to a buyer. A valuable-looking archive is therefore a candidate for structured diligence, not an asset with an automatic licensing price.

1. Value begins with the problem, not the database

Companies commonly begin a data monetisation discussion with inventory: ten years of service tickets, millions of transactions, thousands of engineering reports or a large repository of customer correspondence. Inventory identifies where evidence may be found. It does not establish economic usefulness. A buyer needs to specify a capability it wants to develop or measure and how additional data could help.

For an AI developer, a collection of fault reports might serve several materially different purposes. It could supply examples for supervised classification, realistic scenarios for evaluating an existing agent, evidence for retrieval-based assistance, or sequences of actions for studying how technicians investigate failures. The relevant fields, quality tests and rights assessment will differ for each use. A record that is excellent for searching product manuals may be poor evidence of whether an autonomous troubleshooting agent actually resolved an incident.

The buyer's counterfactual matters just as much. If public manuals, synthetic examples, existing customer data or less expensive licensed sources deliver comparable task performance, the incremental contribution of a proposed archive may be modest. Conversely, a smaller private collection of verified outcomes from difficult industrial incidents could fill a gap that alternative sources do not cover. The distinction is useful information not otherwise available at acceptable cost.

Research on data valuation reinforces this conditional approach. Ghorbani and Zou's Data Shapley framework measures how individual training examples contribute to a specified predictive task, rather than assigning every record an independent market price 1. Its methods are not an off-the-shelf way to quote a licensing fee for an enterprise archive, but the underlying principle is transferable: test the marginal improvement associated with the data. A practical pilot compares a baseline system with the same system using the candidate material, under an agreed evaluation method. This does not remove commercial uncertainty, but it turns an abstract sales claim into testable evidence.

Exhibit 1. Buyer task determines what constitutes useful data

Potential buyer taskMost informative operational evidenceA meaningful testMain limitation to investigate
Diagnose equipment failuresSymptoms, component state, actions, verified repair, recurrenceAccuracy on held-out failure typesMissing or unreliable outcome labels
Evaluate customer-service agentsRequests, constraints, actions taken, escalation and resolutionSuccess on realistic cases with known outcomesCustomer confidentiality and inconsistent case definitions
Improve invoice exception handlingException reason, policy, resolution path, subsequent reconciliationFewer incorrect decisions at the same control standardPolicy changes and access to financial records
Support document retrievalVersioned technical documents, identifiers and authoritative answersAnswer accuracy against current documentationObsolete versions and unresolved document rights

Interpretation: A seller should prepare a task-specific evidence pack, not a general claim that all of its stored information will train an AI model. The best technical form depends on the use case.

2. Depth, coverage and the diminishing return of repetition

More observations can help an AI system distinguish a genuine pattern from an isolated accident, but only if the observations bring additional relevant information. A support platform might contain two million messages but only 120,000 distinct cases. Counting messages as if each were an independent outcome inflates the apparent evidence base. An executive should ask what one observation represents: a customer, a case, an action, an event, an interaction or a complete resolved workflow.

Coverage is as important as volume. Does the archive span different products, jurisdictions, shifts, operating conditions and failure severities? Are rare but consequential cases represented? Are unsuccessful interventions recorded, or only routine successes? A model trained on repeated straightforward examples might perform well on the common case while remaining unreliable precisely where skilled human judgement is most valuable. For the buyer, the distribution of cases must match the intended deployment environment.

Historical breadth adds another dimension. Longer records may include unusual economic periods, regulatory changes, supplier disruptions or slow-moving outcomes. They can also mix incompatible policies and obsolete systems. Time therefore provides potential diversity, not a simple multiplier of worth. Chapter 4 evaluates the additional information contributed by older cohorts; at this stage, management should simply identify whether history changes coverage or the distribution of usable cases.

Repetition can have a cost. In language-modelling research, Lee and colleagues found that removing duplicated training text reduced memorisation and improved aspects of model training and evaluation 2. Their findings do not prove an identical effect for every enterprise dataset, but they demonstrate why a large file can exaggerate its informational content. A useful audit should therefore estimate unique cases, duplicate rates, coverage by category and the fraction linked to credible outcomes. Likewise, training and evaluation records should be separated carefully to prevent leakage from producing flattering but misleading results.

A seller should resist the temptation to delete every imperfect or failed case. Errors and corrections can be highly informative when their causes and outcomes are documented. The issue is not whether the records look tidy; it is whether they can be interpreted accurately. Real-world exceptions may be more educational than sanitised summaries, provided that they remain lawful and safe to use.

3. Scarcity is valuable only when a buyer cannot replace the signal

Uniqueness is often invoked as though a proprietary database automatically commands a premium. More precisely, differentiation exists when a buyer cannot obtain equivalent task-relevant information through a reasonable substitute. A specialised maintenance archive, for example, may record unusual failures that public documents explain only theoretically. It might also contain practical interventions, unsuccessful attempts and verified recovery, making the evidence unusually rich. That is a stronger proposition than confidentiality alone.

Scarcity has several dimensions. Some collections represent hard-to-observe events, such as major industrial breakdowns or infrequent credit restructurings. Others record specialised work carried out in a small number of organisations, or connect otherwise separate systems through stable identifiers. Records in under-represented languages or operational environments may be differentiated for particular applications. Yet diversity or rarity can also work against use: an exotic category with only a handful of unreliable labels may be difficult to train on or evaluate convincingly.

The competitive question is what happens if the buyer does not obtain this particular dataset. Could it achieve most of the intended improvement with domain experts, simulations, public records, instrumented pilot deployments or data from another provider? Is the proposed archive relevant to several buyers or only a narrow technical requirement? The answers affect both bargaining power and the prospects of repeat transactions.

Exclusivity is a separate commercial decision. A buyer may pay more for carefully defined exclusive rights where exclusivity creates demonstrable advantage, but it restricts the seller's future options. The benefits and restrictions should be modelled rather than assumed. Similarly, an agreement permitting evaluation but prohibiting model training has a different commercial scope from a broader licence. Dataset scarcity, contractual exclusivity and permissible use are related, but they are not interchangeable.

This suggests a disciplined negotiating posture: articulate the gap the data fills, demonstrate it on a representative sample, disclose material limitations and compare the cost of substitutes. Without that evidence, statements about being “unique” are marketing descriptions rather than economic conclusions.

4. Context, quality and documentation convert records into evidence

A buyer needs to understand what happened, how the record was collected and what can reasonably be inferred. The most important metadata may be mundane: stable case identifiers, event timestamps, system versions, role definitions, units, escalation codes, resolution status and links between actions and outcomes. If these fields are absent or unreliable, sophisticated content can become difficult to use. An isolated sentence such as “replaced valve; fixed” has much less explanatory value than a linked case showing equipment type, symptoms, diagnostic tests, intervention and a subsequent period without recurrence.

The original distinction between structured, unstructured and hybrid data is useful, provided it does not become a ranking. Structured tables enable joins, filtering and quantitative outcomes. Unstructured messages, images and documents can preserve linguistic detail, judgement and context. Hybrid collections may combine both, but require careful identity matching and often substantial preparation. No format is intrinsically best: utility is determined by the buyer's objective and the integrity of the relationships.

Data documentation is itself an economic asset. Gebru and colleagues' Datasheets for Datasets and Google's Data Cards advocate recording composition, origins, collection methods, appropriate uses and limitations [3, 4]. These approaches help a buyer assess suitability and can reduce repeated diligence questions. They do not certify that a dataset is lawful or accurate; rather, they create a more transparent basis for inspection.

Quality should be measured along explicit dimensions: completeness of decision-critical fields; validity of dates, values and identifiers; consistency across systems; accuracy or auditability of outcome labels; provenance; timeliness; and representation of the population to which the buyer will apply the data. Research interviewing 53 practitioners in high-stakes AI described how neglected data issues can cause compounding downstream problems, underlining the operational importance of upstream evidence quality 5.

Exhibit 2. Two equal-sized archives can have different practical utility (hypothetical)

CharacteristicArchive A: 500,000 recordsArchive B: 500,000 records
Primary unitMessages from unresolved and resolved tickets mixed togetherDistinct cases with consistent IDs
Outcome linkageResolution available for 25% of casesVerified outcome available for 85%
ContextProduct version missing; free-text categoryProduct, version, intervention and severity recorded
RepetitionMultiple updates to the same routine incidentsMix of routine, exceptional and failed interventions
Rights statusSome customer contracts unreviewedRelevant contracts reviewed; permitted scope defined
Most plausible first useLanguage analysis after filteringCase-based evaluation or training pilot, subject to tests

Interpretation: The figures are constructed solely to illustrate diligence questions. Archive B is not automatically more valuable; a buyer studying language variation might favour material in A. Neither archive has an evidenced sale price.

5. Rights, provenance and deliverability are commercial constraints

Data utility does not entitle a seller to use or share information for any purpose it chooses. Operational archives can contain personal data, employee communications, customer secrets, suppliers' intellectual property, licensed third-party material and documents subject to contractual restrictions. Control of the technical database is not the same as owning every right relevant to an AI use. The UK Intellectual Property Office distinguishes copyright in certain selections or arrangements from database rights relating to qualifying investment in obtaining, verifying or presenting contents 6. These protections do not, by themselves, resolve rights in underlying materials.

Where personal data is involved, the purpose, legal basis, fairness, transparency, minimisation, security and relevant cross-border rules need careful assessment. The UK Information Commissioner's Office treats giving another organisation access to personal data as a form of sharing and provides a structured code of practice 7. The ICO notes that parts of the code are being reviewed following legislative change, so any transaction should rely on up-to-date, jurisdiction-specific advice rather than a generic assertion that business data can be sold. Anonymisation should be tested for re-identification risk; masking names alone does not necessarily remove identifiability.

Provenance also affects technical reliability. A buyer should know which systems produced the records, what transformations took place, who annotated outcomes, and whether the extract is reproducible. NIST's Generative AI Profile recommends review and documentation of data accuracy, representativeness, relevance and suitability across the AI lifecycle 8. A clear evidence trail makes it easier to identify both risks and plausible uses.

Deliverability has a direct cost. Old systems may require vendor assistance or bespoke extraction; data from acquisitions may use conflicting identifiers; redaction may eliminate essential context; secure hosting, access controls and deletion processes require effort. A buyer may accept controlled remote evaluation or a restricted analytical environment rather than receiving a full raw copy. Such alternatives can preserve some utility while reducing exposure, but need to be designed and priced. An archive with enormous theoretical promise may have no viable near-term transaction if the legal and engineering work exceeds the expected benefit.

6. From technical usefulness to a defensible commercial case

A commercial assessment should separate model utility from monetary value. Technical utility asks whether the data improves a named training, evaluation or retrieval task relative to a credible baseline. The buyer's economic benefit asks whether that improvement reduces errors, saves time, increases throughput or lowers other costs in actual operations. The amount a buyer might pay is narrower still: it depends on substitutes, integration, uncertainty, negotiating leverage and the share of benefit attributable to the dataset.

A useful decision sequence is therefore: define the use case; benchmark incremental performance; model how that improvement affects operations; discount uncertainty and deployment costs; assess what other suppliers could offer; and then test whether proposed licence terms leave an acceptable return for both parties. This is a valuation process, not a market-pricing formula. Research methods such as Data Shapley help assess model-level contributions but do not establish a market-clearing licensing price 1.

Exhibit 3. Sensitivity of annual operational benefit (illustrative only)

Assume a buyer processes 120,000 service cases annually and the fully loaded labour cost is £30 per hour. An improved AI assistant saves an average of 2, 5 or 8 minutes on 5%, 10% or 20% of cases. Assume saved time is genuinely released into productive work; do not count cases without realised time savings.

Formula: annual gross benefit = cases × share affected × minutes saved ÷ 60 × loaded hourly cost.

Share of cases affected2 minutes saved5 minutes saved8 minutes saved
5%£6,000£15,000£24,000
10%£12,000£30,000£48,000
20%£24,000£60,000£96,000

Interpretation: These numbers are not observed AI performance improvements, expected licensing fees or estimates of an actual buyer's willingness to pay. They demonstrate how modest operational assumptions change the economic opportunity. The entire benefit cannot automatically be assigned to the dataset, because software, implementation and other information may also contribute.

Consider the middle case: £30,000 annual gross benefit. For illustration, suppose, strictly as a negotiating scenario, that the buyer assigns 30% of that benefit to this particular data contribution: £9,000 a year. This 30% is an arbitrary scenario assumption, not an industry benchmark. On the seller's side, suppose preparing the records and legal documentation costs £18,000 once, with £5,000 a year of secure delivery and compliance costs. At £9,000 annual receipts, the seller's first-year contribution before general overheads is minus £14,000; the cumulative three-year contribution is minus £6,000 (3 × £9,000 − £18,000 − 3 × £5,000). Under these assumptions, a one-buyer deal is unattractive unless the price, scope, repeat-use opportunity or cost structure changes. A technically useful dataset is not necessarily a monetisable dataset.

This example also warns against generic price-per-record calculations. The more credible negotiation starts with use-specific benefit, alternatives and constraints. If the material can be licensed lawfully to several independent buyers, its development cost may be spread across contracts; if rights require exclusivity, the opportunity cost rises. Future receipts remain uncertain until buyer demand and permitted uses are substantiated.

Exhibit 4. Executive diligence matrix: questions that determine the next step

DimensionTestable questionEvidence to requestCommon reason to pause
RelevanceWhich buyer task could improve?Written use case and baseline metricNo defined application
Depth and coverageHow many independent, informative cases?Case counts, distributions, exceptionsMostly duplicates or unrepresentative cases
DifferentiationWhat do substitutes lack?Comparable alternatives and sample analysisEquivalent cheaper information exists
Context and qualityCan actions be linked to outcomes?Schema, joins, completeness and label auditBroken identifiers or opaque labels
Rights and provenanceWhich uses may legally be licensed?Contract and data-flow reviewMaterial restrictions unresolved
DeliverabilityCan a safe, useful extract be produced?Extraction plan, controls and cost estimatePreparation costs overwhelm benefit

Interpretation: This is a portfolio-level feasibility screen, not a numerical valuation model. A single critical failure can outweigh strengths elsewhere. Chapter 5 develops the more specific acceptance tests needed for reconstructing and testing agent workflows.

Practical implications for operating companies

The sensible initial exercise is a narrow, evidence-led feasibility review, not an immediate export of the full archive. Select one operational domain and document the intended record unit, typical workflow, key outcome fields, years covered and principal systems involved. Produce a small representative sample through approved security and privacy controls, then assess its completeness, case diversity and possible use cases. The sample should preserve enough structure for testing while withholding material that is not authorised for the review.

Ask a potential buyer to identify the task, baseline, success measure and acceptable evidence of improvement. In parallel, have legal, security and data owners establish whether the proposed use and transfer are permitted, whether controls can meet the risk, and what preparation and ongoing servicing would cost. Treat unresolved rights as a gating issue rather than a minor deduction in an assumed value score.

Finally, compare the buyer's realistic incremental benefit with the cost of producing and licensing the data. The commercial objective is not to prove that the company owns a large archive. It is to prove that a defined, permitted, deliverable slice of that archive adds value which a buyer cannot obtain more efficiently elsewhere. Sometimes the conclusion will be that the data is better used internally, or that no viable transaction exists. Reaching that conclusion early is itself a useful management outcome.

Sources and further reading

  1. Ghorbani, A. & Zou, J. (2019), ‘Data Shapley: Equitable Valuation of Data for Machine Learning’, Proceedings of the 36th International Conference on Machine Learning, PMLR 97, pp. 2242–2251. https://proceedings.mlr.press/v97/ghorbani19c.html
  2. Lee, K. et al. (2022), ‘Deduplicating Training Data Makes Language Models Better’, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pp. 8424–8445. https://aclanthology.org/2022.acl-long.577/
  3. Gebru, T. et al. (2021), ‘Datasheets for Datasets’, Communications of the ACM, 64(12), pp. 86–92. https://www.microsoft.com/en-us/research/publication/datasheets-for-datasets/
  4. Pushkarna, M., Zaldivar, A. & Kjartansson, O. (2022), ‘Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI’, ACM FAccT. https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/
  5. Sambasivan, N. et al. (2021), ‘“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI’, ACM CHI. https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/
  6. UK Intellectual Property Office (2020), ‘Sui generis database rights’, GOV.UK. https://www.gov.uk/guidance/sui-generis-database-rights
  7. Information Commissioner's Office (ICO), ‘Data sharing: a code of practice’, GOV.UK/ICO guidance hub. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/ (check current updates and applicability before a transaction).
  8. National Institute of Standards and Technology (NIST) (2024), Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf