Executive perspective
Executive perspective
The buyer’s task determines the type of evidence required. Training examples, independent evaluation cases, current retrieval documents and ordered agent trajectories are distinct assets. A single collection may contain all four, but that must be demonstrated.
Performance is comparative, not intrinsic. A data supplier can measure coverage, error rates and preparation yield; it cannot promise model improvement without a defined buyer baseline and a controlled evaluation.
Temporal and information leakage can invalidate attractive results. Labels attached after an outcome must not accidentally appear as inputs to earlier decisions, while training and test examples require genuine separation.
The economic unit is a qualified case. Preparation effort, legally eligible records, validation costs and usable evidence matter more than the number of archived messages.
1. Four distinct routes from enterprise data to AI capability
Exhibit 1. Four routes from enterprise records to AI capability
| Use | How the data is used | What tends to make it useful | Common failure |
|---|---|---|---|
| Training or adaptation | Examples are used to teach or tune a model to perform a task or exhibit a desired behaviour. | Clear inputs, reliable target outputs, relevant edge cases and consistent examples. | Contradictory labels or examples encode poor practice. |
| Evaluation | Held-back cases test whether a model or system performs as intended. | Trusted outcomes, realistic difficulty, coverage of important conditions and independence from training. | Leakage or flawed cases make performance look better—or worse—than it is. |
| Retrieval | A system searches approved information at answer time and supplies relevant material as context. | Current, authoritative, well-organised content with usable access controls and metadata. | Stale or poorly chunked documents produce irrelevant or misleading context. |
| Agent workflows | Sequences of observations, tool calls, decisions and outcomes are used to test or improve multi-step systems. | Correct event order, recoverable state, valid tool results and verified outcomes. | Logs show actions but not whether the task actually succeeded. |
These are not interchangeable product categories. A buyer evaluating a model for a specialised task may want demonstrations and corrections. A buyer building a knowledge assistant may instead need current, permissioned documents and reliable retrieval metadata. A team assessing an agent may need complete traces, including tool responses and failures.
A single archive may support more than one use, but the preparation should keep the resulting assets and permissions distinct. In particular, evaluation cases should be held apart from training examples when the goal is to measure generalisation. If a model has already seen the evaluation answers or close variants, its score may not reflect performance on unseen work. OpenAI’s evaluation guidance discusses contamination and evaluation validity.
The technical route also changes what the provider should promise. Fine-tuning alters a model through examples; a retrieval system indexes documents and supplies them at inference time; an evaluator holds examples apart to measure behaviour. Lewis and colleagues’ work on retrieval-augmented generation illustrates this architectural distinction, not a guarantee of accuracy for a particular enterprise corpus. Organisations often combine these approaches, but the contract must identify each authorised use and the documentation should make resulting subsets distinguishable.
For training, the decisive unit may be an input paired with a verified target response, including corrections or a choice among alternatives. For retrieval, the same response may be irrelevant unless the authoritative source document, effective date, access permissions and refresh procedure are present. For evaluation, the source must not be incorporated into the same model’s training or a derivative corpus in a way that invalidates independence. Differences in task design can make the same 50,000 archived interactions either useful evidence or little more than raw material.
2. Work backwards from the capability and the test
Before assembling examples, define the capability in operational terms. “Improve customer support” is too broad to guide data selection. A more testable goal might be:
- classify incoming issues into the correct category;
- propose an approved next action;
- identify cases that require escalation;
- produce a concise resolution summary;
- select the right tool and arguments for a workflow.
For each goal, specify what the system receives, what good performance looks like, how errors are scored and which errors matter most. That makes it possible to select relevant records instead of exporting an entire archive and hoping it helps.
The task specification
Exhibit 2. Task specification and baseline design
| Design question | Example for support triage |
|---|---|
| Input | Customer message, product type, recent events and permitted account context. |
| Expected output | Issue category, urgency, next action and escalation decision. |
| Ground truth | A verified resolution code and reviewer-confirmed escalation outcome. |
| Important conditions | Ambiguous descriptions, multiple issues, missing information, unusual products and policy exceptions. |
| Failure severity | Incorrectly closing a safety-related issue is more serious than selecting a less precise subcategory. |
| Test method | Compare model outputs with independently reviewed cases and report performance by relevant condition. |
This approach exposes gaps early. If the archive contains messages and final resolution codes but no reliable record of the next action, it may support classification but not demonstrate the full decision process. If an “escalated” flag is present without a definition or consistent use, it may be a weak target until audited.
Data quality research reinforces why volume alone is insufficient. Work on the FineWeb corpus documents filtering and deduplication choices as important parts of dataset curation; its results concern particular language-model experiments and should not be treated as a guarantee that any one curation method will improve every domain or model. FineWeb paper.
Build a baseline before counting data as an improvement
A buyer should begin with a fixed task specification and a baseline system. For a triage model, that might mean measuring current category accuracy and the proportion of critical faults incorrectly routed to ordinary service. For retrieval, the baseline is whether the system finds the right current procedure and cites the governing version. Only after fixing the baseline should additional data be tested. Otherwise a simultaneous model change, prompt change and dataset change leaves the cause of any apparent gain ambiguous.
Training and evaluation populations should be split by the unit at which leakage can occur. If multiple notes describe the same case, splitting randomly by message allows near-duplicates across sets. If a customer or asset appears repeatedly, dependence can extend across cases. Where future conditions are being predicted, a time-based holdout may be more realistic than random assignment. Preserve an independent final test set and record which model versions accessed which examples. OpenAI’s published fine-tuning guidance similarly distinguishes training and test files and recommends early evaluation design; it does not suggest that every fine-tuning method improves every task.
3. Worked example: a field-service archive
Imagine a field-service company holding work orders, technician notes, parts records, sensor readings, dispatch changes and return-to-service outcomes.
A useful record might have fields such as:
work_order_id asset_type equipment_version fault_code reported_at inspection_at technician_note action_taken parts_replaced procedure_version escalated closed_at return_to_service_verified repeat_failure_within_30_days
The fields are only meaningful with definitions and provenance. For example:
- Does
closed_atmean a technician submitted the form, a supervisor approved it or the asset returned to service? - Was
technician_notewritten before the repair, after it, or edited later? - Does
escalated = truereflect a safety threshold, customer request or workflow rule? - Does
procedure_versionidentify the instructions used at the time, or the current instructions? - Is a repeat failure linked to the same asset and fault, and how is the 30-day window calculated?
A buyer could derive several separate assets from the archive:
Exhibit 3. Different products derived from a field-service archive
| Candidate asset | Example package | Essential quality evidence |
|---|---|---|
| Adaptation examples | Fault description → verified diagnosis and approved next step. | Review status, conflicting diagnoses, policy version and error categories. |
| Evaluation set | Held-back cases with expected classification, escalation and rationale. | Independence from training, case coverage, scoring rules and reviewer agreement. |
| Retrieval collection | Current manuals, approved procedures and equipment-specific bulletins. | Effective dates, superseded versions, access rights and source authority. |
| Agent trajectories | Alert → dispatch → inspection → parts request → repair → verified return to service. | Timestamps, event order, tool outcomes, retries and evidence that the workflow completed. |
The archive should not be treated as a ready-made package for every use. It may need customer identifiers removed or restricted, free text screened for sensitive information, old procedure versions separated, and outcomes checked against source systems. Removing context too aggressively can also damage usefulness: an equipment version or timestamp may be necessary to interpret the diagnosis correctly.
The practical aim is not “clean everything.” It is to preserve the context needed for the stated task while excluding or protecting information that is unnecessary or not authorised for that use.
Treat task-specific signal as a curated product
Suppose the hypothetical field-service archive contains 50,000 closed work orders. A rights review finds that 80% can be considered for the intended use; of those, 70% have sufficiently linked inspection and repair information; of those, 60% have a verified post-repair outcome. The sequential yield is 50,000 × 0.80 × 0.70 × 0.60 = 16,800 candidate cases. These rates are illustrative, not observed performance. A candidate case still needs label and representativeness checks; the other 33,200 records have not necessarily become worthless, as some may support retrieval or language analysis instead.
Assume £12,000 of one-off extraction and validation work, plus £1.40 of review cost for each of 16,800 candidates. Direct preparation costs would be £35,520, or about £2.11 per candidate, before contractual clearance, secure delivery, governance and the buyer’s own experiments. Increasing the verification rate from 60% to 75%, with other rates unchanged, would yield 21,000 candidates, but the buyer would still need to test whether those additional examples materially improve the target task. This separates measurable preparation economics from unverified licensing-market value.
4. What buyers can measure—and what those measures miss
A buyer should evaluate the relationship between the dataset and a capability, not just count rows, tokens or files. Useful diligence usually starts with a data profile and then tests representative samples.
Measurement checklist
Exhibit 4. Buyer evaluation and measurement checklist
| Dimension | Questions and possible measures |
|---|---|
| Coverage | Which products, workflow stages, operating conditions and difficult cases are represented? What important segments are absent? |
| Label reliability | Are target outputs reviewed? Are definitions stable? How often do reviewers disagree? |
| Outcome validity | Does the recorded outcome demonstrate success, or merely administrative closure? |
| Duplicates and leakage | Are records repeated across splits? Do training records contain answers or fields unavailable at real use time? |
| Temporal integrity | Were timestamps preserved? Could later edits or future information leak into the model’s input? |
| Representativeness | Does the sample reflect current operations, or a historic process, region or product version? |
| Safety and escalation | Are rare but high-impact cases present and correctly labelled? Is the cost of an error considered? |
| Retrieval quality | Are documents current and authoritative? Can the correct source be found and attributed? |
| Agent validity | Are tool calls, arguments, responses, retries and final outcomes captured in sequence? |
| Uncertainty | Which fields are missing, inferred, weakly labelled or subject to changing definitions? |
Metrics must match the task. A high overall accuracy can conceal poor performance on a rare but consequential class. A retrieval system may return relevant documents that are stale or unauthorised. An agent may make the correct tool call but fail to complete the task. Report slices, errors and uncertainty—not only one aggregate score.
Evaluation cases also need to be valid. A benchmark can be difficult yet still measure the wrong thing if prompts are underspecified, expected answers are inconsistent or scoring rewards superficial matches. Audits of coding benchmarks have documented such data-quality problems, illustrating that evaluation material itself requires quality assurance. OpenAI’s SWE-bench Pro audit.
Interpret the loss function, not only the average score
Consider 200 held-back hypothetical support cases, of which 40 genuinely require urgent escalation. A baseline recalls 30 of those 40; a candidate trained with new data recalls 35. The additional five detected critical cases matter even if overall category accuracy changes little. Yet this improvement could be offset by more false escalations, inconsistent reviewer labels or mistakes on a different vulnerable group. The buyer should therefore report critical-case recall, unnecessary escalation rate, calibration where relevant, uncertainty intervals and results by policy version, location or equipment type.
Reliable evidence also needs control conditions: the same evaluation cases, comparable model capabilities and disclosed prompt or workflow changes. A field-service example does not by itself prove transfer to a different manufacturer or country. Published work on data cascades found that neglected data-quality issues can compound downstream; it is a reason to audit the data path, not evidence that every additional record improves a model. Where records contain only administrative closure, training to imitate them may reproduce poor practices. Independent assessment of final customer or asset outcomes can be more valuable than polished labels alone.
5. Package, protect and commercialise the capability evidence
Preparation is part of product design. A credible delivery might include a task-specific subset, schema, data dictionary, version information, provenance notes, quality statistics, known limitations and a correction process. Where permitted, separate training, evaluation and retrieval collections can prevent accidental mixing and make the intended use clearer.
A practical preparation sequence:
Define capability → select records → validate labels and outcomes → preserve provenance → separate use-specific subsets → document limitations → test a sample → agree permissions and delivery controls
The preparation process should also record how the data changed: filters applied, fields removed, transformations made, versions included and quality checks performed. Otherwise, a recipient may be unable to reproduce a result or understand why two deliveries differ.
Avoid overclaiming
A provider can describe the dataset and its checks, but should not promise “model improvement” based only on volume or perceived uniqueness. Demonstrating improvement requires a defined model, baseline, task, evaluation method and comparable conditions. Results on one configuration may not transfer to another.
Similarly, distinguish technical capability from permission. A collection may be technically suitable for evaluation while still containing personal, confidential or third-party material that cannot be shared for the proposed use. For personal data, parties need to establish their actual roles and responsibilities; a label such as “processor” in a contract does not by itself decide the legal relationship. ICO data-sharing guidance.
For certain general-purpose AI model providers, Article 53 of the EU AI Act includes obligations concerning training-content summaries and copyright policies. These obligations apply to the relevant model providers; they do not themselves grant a dataset supplier permission to provide data or establish that a particular use is lawful. EU AI Act, Article 53.
Executive takeaway
Enterprise data is most useful when it connects a real operational capability to evidence: examples that teach it, held-back cases that test it, authoritative information that supplies it, or complete trajectories that reveal how work unfolds.
Start with the task and the evaluation method. Then identify the smallest defensible package that can support that purpose, document what it does and does not show, and confirm that the intended use is authorised.
Commercial implications of test design
For a provider, successful diligence is not simply receiving a good benchmark score from a prospective buyer. The team should agree in advance whether the task is classification, document retrieval, risk-sensitive escalation, evaluation or tool-using agency; which subset is being tested; what the buyer may retain; and what acceptance standard constitutes useful incremental evidence. A buyer might find strong technical performance but decline a licence because alternatives are cheaper, permissions are narrow or the data cannot be refreshed. Those are commercial conclusions, not contradictions of the technical experiment.
A credible first package therefore contains the task specification, extraction filters, excluded-record categories, a versioned schema, a rights and privacy summary, independent test protocol and known shortcomings. This has a cost, but it gives the supplier something the counterparty can evaluate without relying on an unsubstantiated claim about the volume of its archive.
Sources and further reading
- OpenAI, Fine-tuning best practices, training versus test sets. https://developers.openai.com/api/docs/guides/fine-tuning-best-practices
- Lewis, P. et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. https://nlp.cs.ucl.ac.uk/publications/2020-05-retrieval-augmented-generation-for-knowledge-intensive-nlp-tasks/
- Sambasivan, N. et al. (2021), Data Cascades in High-Stakes AI, ACM CHI. https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/
- Gebru, T. et al. (2021), Datasheets for Datasets, Communications of the ACM. https://doi.org/10.1145/3458723
- Google Cloud, Retrieval-augmented generation product/use-case description. https://cloud.google.com/use-cases/retrieval-augmented-generation
- European Union, AI Act, Article 53; model-provider duties, not an authorisation to share third-party data. https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- OpenAI, Trustworthy third-party evaluations; testing independence and contamination issues. https://openai.com/index/trustworthy-third-party-evaluations-foundations/
- Penedo, G. et al. (2024), The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, preprint. https://arxiv.org/abs/2406.17557