Data Monetisation 101 / Section 1 / Chapter 1

Section 01 · Understand the asset · Chapter 01

Why AI Companies Are Looking Beyond the Public Internet for Training Data

The public web offers scale and breadth, but often lacks verified decisions, operational context and clear permissions. This chapter explains when enterprise records offer distinct evidence, how buyers might use it and why commercial value must be demonstrated rather than assumed.

13 min read8 referencesSite last updated 9 October 2026

Executive perspective

Executive perspective

Public data remains essential. Web pages, books, code and openly available research expose AI systems to varied language and knowledge. Their abundance does not, however, guarantee evidence of what happens when a professional makes a difficult decision or whether that decision works.

Enterprise data can supply missing context. A linked record of inputs, actions, exceptions, corrections and independently checked outcomes may help a developer train or evaluate a specific capability. Its usefulness depends on task fit, not simply on being private.

Permission and preparation are part of the product. Internal records can be duplicated, biased, incomprehensible or restricted by contracts and data protection obligations. Some archives are uneconomic to prepare or unsuitable to share at all.

The commercial question is measurable. Management should determine which eligible subset can improve a named buyer task, what proving that benefit would cost and whether the proposed use is permissible before attaching a monetary value.

1. What the public internet teaches, and what it cannot prove

The public internet is one of the great sources of breadth in modern AI development. It contains explanations, dialogue, software, photographs, technical writing, history and expressions of ordinary human activity. Broad collections support the learning of general patterns and concepts. The proposition that public-web data has ceased to matter is therefore false. The more precise question is whether adding more similar material improves the particular behaviour a developer wants to achieve.

A webpage describing how to resolve an insurance claim is not necessarily an example of a successfully resolved claim. A forum discussion may contain several plausible diagnoses without identifying the one confirmed by the subsequent repair. A public code snippet may describe a procedure without the permissions, state changes or failure conditions encountered inside an operating company. Public sources often make knowledge visible while leaving the execution and results of work unobserved. That distinction becomes consequential when an AI system must select actions, handle exceptions or be evaluated against real outcomes.

The limitations are not confined to missing outcomes. Source datasets can repeat the same passages, contain obsolete information or include material for which permissions are disputed. The statistical frequency of an assertion does not establish its correctness. Research by Lee and colleagues found substantial duplication in language-model training datasets and showed that deduplication could improve aspects of memorisation and evaluation 1. In vision-language modelling, Goyal and colleagues demonstrated that data-filtering choices depend on the available training computation and the diminishing utility of repeated examples 2. Neither result establishes that private business data is generally superior. Both show that record count alone is an inadequate measure of useful learning signal.

The difference also varies across tasks. A general-purpose system may need diverse public examples that one company's archive could never supply. A specialist document assistant might primarily need retrieval of authoritative current policies. A support agent being evaluated on escalation decisions might instead benefit from authentic case histories and verified resolution outcomes. A data supplier should therefore avoid pitching a generic “AI training dataset” before identifying what a buyer intends to learn, retrieve, test or automate.

Exhibit 1. Four useful distinctions between public and operational sources

DimensionBroad public or web-derived materialOperational company recordsWhat the buyer must test
CoverageBroad topics, styles and participantsRepeated evidence from particular workflowsDoes the target task need diversity or procedural depth?
OutcomesOften narrated, asserted or absentMay link a decision to subsequent eventsIs the outcome genuine, mature and independently checked?
DocumentationCollection and rights can vary by sourceInternal fields and process versions may be reconstructableCan definitions, changes and exclusions be explained?
RestrictionsAccessibility does not resolve permissionConfidentiality, contracts and personal data may constrain reuseCan the proposed downstream use be authorised?

Interpretation: These are tendencies, not universal properties. Public benchmark datasets can be exceptionally well annotated; private records can be unlabelled and unusable. The right comparison is between actual candidate datasets for a specified purpose.

2. The missing asset may be the decision, not the document

Consider an equipment-support operation. The customer describes a fault, a technician consults a manual, checks a diagnostic reading, chooses a corrective action and records a repair. If the archive preserves only the customer's first message, an AI developer has a description of the problem but little evidence of the resolution. If it also preserves product version, tests performed, sequence of actions, escalation and the absence or return of the fault, a much richer task becomes observable.

Such a sequence can serve different purposes. A training example may teach a model to recognise an appropriate next action. An evaluation case can test whether a model would escalate when conditions demand it, but only if the expected response or outcome is credible and evaluation cases have not contaminated training. A retrieval system may need the latest authorised procedure rather than past decisions. A workflow agent may require complete state and tool-action histories, not just a final answer. Calling all four products “training data” can conceal materially different quality and contractual requirements.

The relationship among fields matters as much as their content. A ticket marked closed may have been administratively completed without solving the customer's problem. A repair marked successful could have been reopened a week later. A complaint recorded after a policy change may not be comparable with an older case under different rules. Without timestamps, field definitions and a way to distinguish observed facts from retrospective annotations, the apparent labels may be misleading. The technical task is therefore to reconstruct a defensible chain of evidence, rather than merely export documents.

Exhibit 2. A decision-linked record as an AI evidence product

Request and operating conditions
    ↓  event date · product version · applicable policy
Information available at that decision
    ↓  exclude knowledge discovered only later
Human or system action
    ↓  who acted · sequence · tools · exceptions
Subsequent observed outcome
    ↓  independent check · elapsed time · corrections
Rights, quality and provenance controls
    ↓  eligible records · documented exclusions · version
Buyer evaluation against a defined baseline
    ↓  measured task result, cost and limitations
Process or decision structure reproduced from the approved manuscript.

Interpretation: The missing elements are not interchangeable. A correct-looking answer without the original problem may be impossible to evaluate; a completed case without follow-up may not prove successful resolution. Conversely, a document lacking outcome labels may still be useful for authorised retrieval.

A business with ten years of this material does not necessarily have ten years of comparable evidence. Systems change, companies acquire new subsidiaries, employees alter classification practices and older records lose context. A well-designed data description records collection procedures, intended uses, limitations, population characteristics and known omissions. Gebru and colleagues' Datasheets for Datasets and Google's Data Cards research offer frameworks for such documentation 34. These studies concern transparency and dataset practice, not licensing prices or proof that documentation alone improves model performance.

Chapter 5 examines complete human workflows, Chapter 13 examines contextual relationships, and Chapter 18 compares different archives by evidentiary yield. Here the lesson is introductory: an executive assessing an archive should ask what relevant behaviour is actually observable within the records before asking how many there are.

3. What documented buyer interest does, and does not, establish

Published AI-company materials show that private data partnerships are a real approach to sourcing some training and domain-specific material. OpenAI described public and private dataset partnerships in November 2023, specifically identifying material not easily accessible online and expressing interest in information that represents human intention 5. Its published data approach also distinguishes selected public data, proprietary partnership data and human feedback 6. Those are examples of publicly described sourcing approaches, not evidence that any named company is presently seeking a particular seller's records, or a quoted market price.

This distinction matters commercially. “AI companies need proprietary data” is too broad to support an investment decision. Different counterparties have different products. A general-model developer may value specialised language or modalities; a sector-specific software company may need a narrow set of reliably labelled decisions; an evaluation team may need genuinely independent hard cases; an intermediary may require repeatable rights-clearance and delivery processes. The same archive might satisfy one proposition and fail another. A seller needs a buyer-task hypothesis and a method for testing it against alternatives.

Commercial supply also differs from simply making a dataset available. A buyer must be able to understand scope, permitted purpose, geographic and sector restrictions, quality expectations, format, retention and the treatment of corrections or deletions. Depending on the use, a company may deliver a static licensed archive, a time-limited evaluation sample or a controlled refresh. The supplier should not assume that a broad perpetual licence, or onward use in any future model, is commercially or legally acceptable. Chapters 7, 8 and 16 consider product cadence, licensing structures and exclusivity in detail.

For sensitive records, governance may be the limiting factor even when the contents would be technically interesting. The UK Information Commissioner's Office explains that data-sharing agreements can document purposes, roles, security and responsibilities, but do not themselves guarantee legal compliance 7. Its guidance is under review following legal changes and must be checked for the proposed jurisdiction and facts. Confidentiality and third-party intellectual-property constraints may exist even where a dataset contains no personal data. Pseudonymisation reduces certain risks but does not itself make personal data anonymous or remove applicable obligations.

The commercial implication is straightforward: technical attractiveness, legal permissibility and contractual saleability are separate tests. Strong technical evidence cannot cure absent rights. Conversely, having permission does not create a buyer or prove the usefulness of the records. A credible early discussion is often more effective when it acknowledges limitations, permitted uses and missing evidence instead of describing the entire archive as “AI-ready”.

4. From records owned to usable evidence: a worked screening example

An internal record count is at most the beginning of an inventory. Suppose a company has 100,000 support cases and wants to test whether its records could help evaluate a specialist triage system. Management may initially present that figure as its potential data supply. A staged screening exercise could show a much smaller sample capable of supporting the proposed evaluation.

Exhibit 3. Illustrative usable-case waterfall, not a market benchmark

StageHypothetical retention assumptionRemaining cases
Source-system inventoryStarting total100,000
Within preliminary policy and permission scope80% of total80,000
Have relevant, verifiable outcomes60% of previous stage48,000
Meet deduplication and documentation threshold75% of previous stage36,000
Fit the selected buyer task and sampling criteria50% of previous stage18,000

Method: Each fraction is applied sequentially to the remaining population: 100,000 × 0.80 × 0.60 × 0.75 × 0.50 = 18,000 cases. The assumptions are fictional and illustrative. Preliminary policy eligibility is not a completed legal clearance. Different buyers may accept different subsets; sampling rules must also protect against bias.

If documenting, screening and preparing the pilot costs a hypothetical £45,000, the direct preparation cost is £45,000 ÷ 18,000 = £2.50 per candidate case. If the last-stage task-fit rate halves to 25%, the usable yield becomes 9,000, increasing that unit preparation cost to £5.00 if total work remains unchanged. Neither number is a licensing fee or a measure of buyer willingness to pay. It is a way to expose the effect of data quality and task fit on the provider's cost base.

Even these cases are only candidates. An evaluation pilot must specify what the system is expected to do, the task distribution, a credible baseline, a hold-out method, scoring rules and failure severity. Cases that look persuasive because they include rich explanations may not represent the everyday population. Records created by earlier automated systems may also carry inherited errors. In regulated settings, validation may require expert adjudication and stronger controls. NIST's voluntary Generative AI Profile offers a broader framework for identifying and managing generative-AI risks across the development and use lifecycle 8; it is not an enterprise-data pricing method.

An executive evaluating a proposal should ask both what the buyer measures and what the seller must spend to generate reliable evidence. Large apparent value may disappear if rights review, manual labelling, de-identification, secure extraction and recurring delivery consume more resources than a commercially credible agreement would support. The opposite can also happen: a modest, well-governed corpus might be relatively inexpensive to prepare and unusually relevant to a narrow task. These outcomes require testing, not extrapolation from the size of the database.

5. An executive test before investing in monetisation

The first sensible deliverable is a short feasibility dossier, not a raw-data transfer. Start with two or three workflows in which the organisation has repeatedly made decisions or assessed outcomes. Identify the systems involved, the time range, record grain, operational meaning of key fields and any known migrations. Specify what an external system might learn, retrieve or evaluate, and where that evidence would improve on publicly accessible material or alternatives already available to a buyer.

Next, ask an operational owner to explain how correctness is established. Is success inferred from an administrative status, independently checked by a specialist or measured after an appropriate period? How are failed attempts, exceptions and disagreements recorded? Does the evidence show what was known at the time? If a model is being evaluated, avoid using future information to explain past decisions. Record known weaknesses rather than silently filtering away difficult cases.

The legal, privacy and security review should occur before any external production sample is shared. It should address the contractual rights chain, confidentiality, employee and customer information, third-party inputs, permitted recipients and uses, minimisation, retention, incident response and potential withdrawal or deletion obligations. A small synthetic illustration or schema description may be suitable for an initial conversation when real records cannot safely be supplied. Synthetic examples should be labelled as such and not presented as evidence of actual dataset quality.

Finally, estimate the cost of turning the candidate subset into something a buyer can test. Allow for data engineering, technical documentation, sample selection, legal review, expert adjudication, secure access and later requests for correction. The decision should not be reduced to a speculative fee multiplied by the total number of records. Management should identify an affordable diagnostic test and pre-agreed criteria for stopping if the dataset proves unusable.

Exhibit 4. A first-stage management decision screen

GateEvidence sufficient to proceed to a controlled pilotReason to pause
Task fitDefined capability, evaluation metric and plausible counterpartyNo identifiable incremental use beyond public alternatives
EvidenceLinked inputs, decisions, context and credible outcomes where requiredLabels cannot be explained or independently checked
Provenance and qualityDocumented systems, definitions, time range and samplingMajor unknown gaps, duplicates or irreproducible extraction
PermissionsProposed sample and downstream use reviewed by appropriate ownersUnclear customer, employee, vendor or third-party restrictions
Unit economicsPreparation scope and decision thresholds are costedWorkload, continuing obligations or exposure cannot be bounded

Interpretation: Passing this preliminary screen does not establish legal clearance, buyer demand or commercial price. It authorises a more disciplined question: can a small, controlled sample demonstrate distinct usefulness at a cost and risk the organisation can accept?

Practical conclusion

The public internet remains valuable because it offers extraordinary breadth; company records matter when they reveal something the public record cannot reliably show. The distinguishing feature is rarely secrecy by itself. It is a defensible combination of task relevance, decision context, outcome evidence, documentation and permitted use. Some operational archives will satisfy none of those requirements. Others may justify a small pilot and, eventually, a licence or partnership. Chapter 1 sets the discipline for the entire course: test the specific evidence and its cost before making claims about value.

Sources and further reading

  1. Lee, K. et al. (2022), ‘Deduplicating Training Data Makes Language Models Better’, Proceedings of ACL 2022, pp. 8424–8445. https://aclanthology.org/2022.acl-long.577/
  2. Goyal, S. et al. (2024), ‘Scaling Laws for Data Filtering: Data Curation Cannot Be Compute Agnostic’, Proceedings of CVPR 2024, pp. 22702–22711. https://openaccess.thecvf.com/content/CVPR2024/html/Goyal_Scaling_Laws_for_Data_Filtering--_Data_Curation_cannot_be_Compute_CVPR_2024_paper.html
  3. Gebru, T. et al. (2021), ‘Datasheets for Datasets’, Communications of the ACM, 64(12), pp. 86–92. https://doi.org/10.1145/3458723
  4. Pushkarna, M., Zaldivar, A. and Kjartansson, O. (2022), ‘Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI’, ACM FAccT 2022. https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/
  5. OpenAI (9 November 2023), ‘OpenAI Data Partnerships’. A dated description of an approach, not evidence of current purchasing demand. https://openai.com/index/data-partnerships/
  6. OpenAI, ‘Our approach to data and AI’. Public description of selected public material, data partnerships and human feedback. https://openai.com/index/approach-to-data-and-ai/
  7. Information Commissioner's Office, ‘Data sharing agreements’, Data Sharing Code of Practice (guidance under review following the Data (Use and Access) Act). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/data-sharing-agreements/
  8. National Institute of Standards and Technology (2024), AI Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. https://doi.org/10.6028/NIST.AI.600-1