Executive perspective
Executive perspective
Size is a capacity measure, not a valuation method. A million customer conversations may be attractive for broad linguistic coverage, while a smaller collection of expert decisions could be more useful for a narrowly specified specialist task. Neither claim is universal.
Signal density depends on the intended use. A missing outcome is a major weakness for evaluating task completion, but may be much less important when studying customer language or retrieving common support requests.
An unusually informative archive can still be biased. Rare cases, failed attempts and verified outcomes matter, but so do ordinary cases, changing operating rules and groups that are under-represented in the records.
Commercial value emerges only after testing usable yield. Deduplication, validation, documentation, legal clearance and buyer alternatives may matter more to a transaction than the nominal row count.
1. The autopsy: two assets with different kinds of information
Imagine two businesses considering whether their historical operational records might support an AI developer. Company A has one million customer-service conversations. They span multiple channels, include natural customer language and cover repeated problems, but only 12% link to confirmed resolution. Password resets and other routine enquiries recur frequently. Company B holds 100,000 expert decisions for a specialist product line. Each case links its inputs, decision, recorded rationale, action and stated verified outcome, although terminology changed during a system migration.
These are hypothetical archives, not examples of actual licensing transactions or established industry averages. Their purpose is to expose a frequent analytical mistake: treating a record as a standard unit of AI usefulness. A conversational turn, a completed case and a professionally adjudicated decision are fundamentally different observations. Even the denominator requires care. One support case may comprise twenty messages; one expert decision may be amended several times. Counting messages instead of distinct cases can inflate apparent volume before any technical analysis begins.
Dataset A could be useful to an assistant intended to recognise everyday customer requests, handle varied wording or retrieve relevant precedents. It offers examples of how people ask for help, what terminology confuses them and when they seek human support. Dataset B may be more suitable for assessing whether an AI system correctly identifies exceptions, requests evidence or follows specialist escalation rules. Because the second collection is narrow, however, it is a poor substitute for language diversity across products and sectors.
Academic research makes the distinction between amount and selection concrete. The DataComp benchmark holds training procedures and compute arrangements constant while comparing strategies for curating image–text training data; different data choices materially affect downstream performance 1. Google's DataPerf initiative similarly treats training-data selection, cleaning and acquisition as measurable design problems rather than assuming more examples are always better 2. These results concern specified benchmark environments, not enterprise-data prices. Their relevance is methodological: assess what the additional records enable, not how impressive the inventory sounds.
Exhibit 1. Dataset A versus Dataset B: a buyer-task matrix
| Proposed AI use | Dataset A: 1m conversations | Dataset B: 100k expert decisions | Critical validation question |
|---|---|---|---|
| Customer-language coverage and intent recognition | Potentially strong breadth and ordinary phrasing | Specialist vocabulary; limited breadth | Does the sample represent actual users, languages and channels? |
| Support retrieval and human assistance | Potentially useful even without final outcomes | Specialist precedents may help in one domain | Are passages current, searchable and safe to expose? |
| Specialist decision support | Often weak if evidence and decision rules are absent | Promising if rationale and policy context are reliable | What was known at the time of the decision? |
| Task-completion evaluation | Only the resolved, independently checkable subset may qualify | Potentially strong for the covered product line | Can success be verified independently without leakage? |
| General-purpose model improvement | May provide conversational variety but substantial redundancy | Limited breadth; could add rare expertise | What measurable improvement occurs versus a buyer's baseline? |
Interpretation: “Better” is conditional on the buyer's task and target population. The table describes possible applications, not confirmed buyer demand or a ranking that applies to every transaction.
2. Signal density means evidence per relevant, usable case
The term signal density is useful only when defined precisely. Here, it means the proportion and richness of task-relevant evidence within a case that survives validity and usability checks. For a support assistant, the relevant signal might include natural language, intent, sentiment, escalation or retrieval cues. For an evaluator of specialist decisions, it could include contemporaneously available inputs, the applicable policy, the permitted action and a trustworthy later result. There is no scientifically established universal “signal-density score” that converts these ingredients into an asset valuation.
A useful audit therefore separates at least three things. Completeness asks whether essential fields exist. Reliability asks whether those fields mean what their labels suggest. Distinctiveness asks how much new information the next case contributes. An archive can be complete but repetitive, varied but poorly labelled, or well curated but irrelevant to the intended task. The order of these tests matters: high-quality labels for the wrong task do not rescue the dataset.
Repeated observations can be valuable when they represent genuine variation. Multiple password enquiries might expose differences in customer vocabulary, device type, accessibility needs or escalation risk. But repeated copies of the same script provide less new coverage. In an influential study of language-model training data, Lee and colleagues found substantial text duplication; removing it reduced memorisation and improved training efficiency in their experiments 3. That finding supports deliberate duplicate analysis, not an instruction to delete all recurrent business events. For some operational uses, frequency itself is the signal: an agent expected to handle routine demand must encounter the routine as well as exceptions.
Label quality is just as important. Northcutt, Athalye and Mueller documented errors in widely used machine-learning test sets and showed that incorrect evaluation labels can affect model comparisons 4. In a business archive, “resolved” might mean the ticket was administratively closed, the customer stopped replying or the fault passed an independent verification check. Those states should not be treated as interchangeable. An expert's title is not proof that every recorded conclusion is correct. Versioned rules, reviewer disagreement, overrides and reversals should be documented rather than silently smoothed away.
These observations suggest a practical distinction between raw count, distinct cases, task-qualified cases and licensable task-qualified cases. Moving through those stages requires sampling and evidence, not a single multiplication by an arbitrary quality score. It also requires enough explanation for a prospective buyer to reproduce the assessment on a controlled subset.
3. Rare cases, failures and representativeness pull in different directions
Dataset B's apparent advantage is concentrated professional judgement. Well-described exception handling, rejected options and later reversals may expose difficult operational decisions that a conversational archive rarely records. Such cases can help test whether a model knows when not to act, how to request missing information and when escalation is necessary. However, a rejected action is not automatically a negative training label: it may have been rejected because of a temporary policy, missing evidence or a risk tolerance specific to that institution.
Rarity requires similar restraint. Uncommon events may be expensive to reproduce, and a small number of well-documented adverse outcomes could be essential for testing a high-consequence process. Yet rare events can also be statistical outliers, recording errors or products of obsolete conditions. More rare cases do not automatically create a better training distribution. Research on long-tailed classification explains how naive training can favour frequent categories, motivating methods that account for imbalanced labels 5. The proper response is to design sampling and evaluation around the deployment objective, not to assume all rare cases deserve equal weight.
A buyer should therefore ask two different questions: does the dataset cover the situations the system will face? and does it adequately stress-test the situations in which failure matters most? These call for different sampling strategies. A representative training sample might preserve the mix of routine and complex work, while a separate stress-test set deliberately over-samples difficult, unusual or harmful failures. Results should be reported by case segment; a single average accuracy figure can conceal systematic weaknesses.
The WILDS benchmark established that models trained under conventional conditions can perform substantially worse when tested on realistic distribution shifts, including differences across sites, time or geography 6. Dataset B's migration-era vocabulary change is therefore not just a cosmetic cleanup problem. It might mark a new product version, altered decision rules or a shift in the customer population. Removing the old vocabulary without documenting the regime would conceal real historical differences; combining regimes without labels could also create misleading examples.
For executive review, the strongest dataset is not necessarily the one with the highest fraction of difficult decisions. It is the one whose coverage can be described, whose limitations can be measured and whose intended training or evaluation role is clear. A dataset containing only successful resolutions may omit the very situations in which an assistant must refuse, seek evidence or hand control to a professional.
Exhibit 2. A sampling design to expose the hidden weaknesses
| Sample stratum | What to inspect | Why it matters |
|---|---|---|
| Routine high-volume cases | Repetition, language variation, distinct customers and outcomes | Separates useful everyday coverage from duplicates |
| Difficult or escalated cases | Evidence available at decision time, rejected paths, hand-offs | Tests whether hard cases are legible rather than just labelled |
| Cases without outcomes | Reason for missingness, observation window, case status | Avoids treating “not observed” as failure or success |
| Pre- and post-migration records | Schema, vocabulary, policy and system versions | Identifies genuine operating change versus coding artefacts |
| Rights-sensitive segments | Customer contracts, employee information, free-text risk | Estimates excluded or restricted material before full extraction |
Interpretation: Draw random samples within each stratum, record population proportions and separately report deliberately over-sampled cases. Do not present stratified sample percentages as population-wide findings without appropriate weighting.
4. A reproducible yield calculation, not a made-up market price
The original worked example provides two known premises: one million conversations with 12% confirmed-resolution linkage, and 100,000 specialist decisions with inputs, actions, rationale and outcomes linked. A company should not turn those observations directly into a transaction price. Instead, it can run a staged audit that converts raw inventory into a defensible count of candidate cases for a named application.
The following assumptions are entirely illustrative, not reported deal economics or statistical estimates. Dataset A is assessed specifically for outcome-linked support evaluation, so its missing-resolution rate matters greatly. Dataset B is assessed for specialist decision evaluation, not broad conversational fluency. Their candidate counts are thus not interchangeable units of value. Each percentage applies sequentially to cases remaining after the previous filter.
Exhibit 3. Hypothetical candidate-yield and preparation-cost sensitivity
| Gate or result | Dataset A: outcome-linked support evaluation | Dataset B: specialist decision evaluation |
|---|---|---|
| Starting cases | 1,000,000 | 100,000 |
| Confirmed outcome / independently rechecked case | 12% of A; source premise | 90% pass follow-up reliability review; assumption |
| Sufficient contemporaneous context | 75%; assumption | 90%; assumption |
| Rights and privacy clearance for proposed pilot | 70%; assumption | 80%; assumption |
| Distinct and in-scope after deduplication / regime checks | 60%; assumption | 75%; assumption |
| Candidate task-qualified cases | 37,800 | 48,600 |
| Assumed total preparation budget | £120,000 | £95,000 |
| Preparation cost per candidate | £3.17 | £1.95 |
Calculation: A = 1,000,000 × 0.12 × 0.75 × 0.70 × 0.60 = 37,800. B = 100,000 × 0.90 × 0.90 × 0.80 × 0.75 = 48,600. Cost per candidate = assumed preparation budget ÷ candidate cases. The 90% review pass rate for B does not contradict the original description of linked outcomes; it represents a hypothetical independent reliability check of those links.
Sensitivity: If A's confirmed-resolution linkage were 25% rather than 12%, with all other illustrative gates unchanged, its candidate yield would be 78,750 and the assumed cost per candidate would fall to about £1.52. At 40%, the figures would become 126,000 and £0.95. These are arithmetic scenarios, not forecasts; better linkage may also require extra spending, which this simplified example holds constant.
The example is intentionally capable of producing a smaller high-confidence subset from the larger archive. It does not demonstrate that Dataset B is worth more money. Dataset A's remaining conversations may still have significant value for language, routing or search use cases without confirmed resolutions. Similarly, Dataset B may become far less attractive if the buyer does not serve its particular product line. Neither budget represents a quoted market price, nor do the per-case costs measure willingness to pay.
Commercial testing requires an additional step: document baseline model performance, compare it with a carefully controlled addition of the proposed data, and assess improvement against the buyer's meaningful metrics. Those might include resolution accuracy, escalation appropriateness, error severity, coverage by customer segment or the cost of manual reviews. For evaluation datasets, the value could lie in revealing failures reliably rather than improving the model directly. A buyer may compare the proposal with synthetic tasks, self-collected data or licensed alternatives, and there may be no economic case after that comparison.
5. Combining the collections can improve coverage, or corrupt evaluation
If the two archives concern related services, their complementarity deserves testing. Conversation data could supply the customer's language and the sequence of help sought; expert records could contribute adjudicated labels, exceptional reasoning or objective outcomes. Yet joining them is not an innocent file operation. It needs stable case identifiers, accurate time ordering, a definition of authoritative records, and a map of rules and system versions. Similar text should not be treated as proof that two records describe the same incident.
A key technical risk is information leakage. If a model is evaluated on cases it effectively encountered during training, performance can look stronger than it will be on unseen work. Leakage can also occur when a task's answer, a later decision or an outcome field is included among inputs purportedly available at the decision point. Scikit-learn's guidance explains why training and test information should be separated before preprocessing, and why test data should not influence the fitting process 7. For operational episodes, splitting by customer, organisation, incident or later time period may be more appropriate than randomly splitting messages from the same case.
The separation should extend to benchmark construction. Suppose Dataset B contains an expert-approved outcome for a case that also appears in Dataset A as a conversation. Putting the conversation in model training and the expert decision in evaluation could create an artificial advantage unless the intended test explicitly permits prior exposure. Conversely, an independently held-back set of difficult cases can be useful precisely because it tests whether the system generalises beyond frequently repeated examples.
Documentation matters here. Google's Data Cards framework calls for explicit information about dataset origins, collection, annotation, intended use and limitations 8. For these archives, a concise dossier should also state the unit of observation, rates of missing outcomes, selection methods, regime changes, de-identification steps and evaluation-set construction. A package without such documentation obliges every buyer to redo basic diligence and may be unattractive even when its content is excellent.
There is a legal boundary as well. Customer conversations and expert decisions may contain personal information, confidential correspondence, commercially sensitive judgments, licensed documents or third-party intellectual property. The UK Information Commissioner's Office emphasises purpose limitation, fairness and data minimisation when personal data is used in AI systems 9. Its guidance also makes clear that pseudonymisation is not equivalent to anonymisation; pseudonymised material remains within data-protection law 10. Permissions for the proposed buyer and purpose must be checked, not inferred from lawful internal possession. A buyer-facing proof of value can sometimes use restricted access, sampled extracts or derived measures rather than transferring raw archives.
6. What executives should actually ask before selecting the better dataset
An executive decision should begin with a proposed task statement, not a dataset catalogue. “Improve customer support” is too broad. “Evaluate whether an agent escalates unresolved billing disputes under the current policy” is testable. It specifies the relevant population, inputs, permitted actions, outcome definition and failure costs. The dataset can then be assessed for that purpose, with exclusions stated as clearly as its strengths.
A practical sequence is to identify one use case; select a controlled, authorised sample across ordinary, difficult and different historical periods; measure completeness and independence of outcomes; review repeat cases and missingness; check policy versions and rights; and conduct a small buyer-relevant technical test. Keep enough evidence to explain every exclusion. Where the dataset is intended for evaluation, reserve cases that are not used in developing or selecting the model. For higher-consequence applications, incorporate domain reviewers and report errors separately for important case groups. NIST's AI Risk Management Framework stresses representative evaluation, documented metrics and assessments relevant to deployment conditions 11.
Exhibit 4. Executive decision gates before discussing licence economics
| Gate | Evidence that would justify progressing | Reason to narrow or stop |
|---|---|---|
| Use-case fit | A named capability and measurable baseline | Buyer cannot specify the task or test |
| Distinctive information | Demonstrated coverage of needed language, decisions or exceptions | Most cases are redundant with buyer alternatives |
| Outcome and regime integrity | Independently checked labels, time/version information | “Closed” codes stand in for unverified success; regime unclear |
| Permitted use and delivery | Rights map, minimisation plan, secure sample process | Customer/third-party restrictions or uncontrolled disclosure |
| Incremental economics | Pilot performance or evaluation gain exceeds preparation and control costs | No measurable utility or excessive reconstruction cost |
Interpretation: Failure at a gate need not condemn the entire archive. It may indicate that only a narrow segment, a different application or internal use is appropriate.
The autopsy's conclusion is more discriminating than “quality beats quantity”. Sometimes quantity is the quality, especially where realistic language coverage and routine-frequency patterns matter. Sometimes a smaller archive's linked decisions and outcomes supply the evidence the buyer cannot easily recreate. And sometimes both fail: one is too repetitive, the other too narrow, poorly cleared or expensive to reconstruct. The defensible commercial proposition is a task-specific, auditable body of evidence, with its limitations acknowledged before anyone attempts to put a price on it.
Sources and further reading
- Gadre, S. Y. et al. (2023), ‘DataComp: In Search of the Next Generation of Multimodal Datasets’, NeurIPS 2023, Datasets and Benchmarks. https://papers.neurips.cc/paper_files/paper/2023/hash/56332d41d55ad7ad8024aac625881be7-Abstract-Datasets_and_Benchmarks.html
- Mattson, P. and Paritosh, P. (2023), ‘Data-centric ML benchmarking: Announcing DataPerf’s 2023 challenges’, Google Research. https://research.google/blog/data-centric-ml-benchmarking-announcing-dataperfs-2023-challenges/
- Lee, K. et al. (2022), ‘Deduplicating Training Data Makes Language Models Better’, Proceedings of ACL 2022. https://aclanthology.org/2022.acl-long.577/
- Northcutt, C. G., Athalye, A. and Mueller, J. (2021), ‘Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks’, NeurIPS 2021, Datasets and Benchmarks. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/f2217062e9a397a1dca429e7d70bc6ca-Abstract-round1.html
- Menon, A. K. et al. (2021), ‘Long-tail learning via logit adjustment’, International Conference on Learning Representations. https://research.google/pubs/long-tail-learning-via-logit-adjustment/
- Koh, P. W. et al. (2021), ‘WILDS: A Benchmark of in-the-Wild Distribution Shifts’, Proceedings of ICML 2021. https://proceedings.mlr.press/v139/koh21a.html
- scikit-learn (documentation accessed October 2026), ‘Common pitfalls and recommended practices: Data leakage’. https://scikit-learn.org/stable/common_pitfalls.html
- Pushkarna, M., Zaldivar, A. and Kjartansson, O. (2022), ‘Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI’, ACM FAccT. https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/
- Information Commissioner’s Office (ICO), ‘How do we ensure fairness in AI?’, Guidance on AI and data protection. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-do-we-ensure-fairness-in-ai/
- Information Commissioner’s Office (ICO), ‘Pseudonymisation’, Anonymisation and data protection guidance. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
- National Institute of Standards and Technology (NIST), AI Risk Management Framework Playbook: Measure. https://airc.nist.gov/airmf-resources/playbook/measure/