Executive perspective
Executive perspective
A diligence pack should make claims independently testable. Record counts, outcome quality, coverage and rights status require denominators, methods and known exclusions. The intended AI use matters: information suitable for retrieval may not be adequate for outcome-based evaluation or agent training.
Disclosure should be progressive. Metadata and task hypotheses can establish initial fit; detailed samples require specific permissions, security controls and documented selection. Preparation is an investment decision, not a housekeeping task. A sensible provider prices extraction, cleaning, legal review and technical support before committing to deliver material whose incremental buyer value has not been demonstrated.
1. Describe the asset at a level a buyer can evaluate
Start with the underlying business activity—not the storage format or a large record count.
Explain what one record represents; who or what creates it; how records relate; what is observed versus inferred; what period and population it covers; and what a proposed buyer might do with it. A useful description might say:
“A case-level archive of industrial maintenance work orders, linked technician notes and sensor observations, covering three equipment families across four service regions from 2018–2025. Outcomes include recorded return-to-service and repeat-failure events. Older records have less complete outcome coding, and two months of telemetry are missing.”
That tells a buyer more than “millions of proprietary records.” A data inventory should also identify source systems, formats, approximate volumes, refresh pattern and known exclusions.
Exhibit 1. Core questions in an evidence-led buyer brief
| Buyer question | Useful evidence |
|---|---|
| What is it? | Plain-language description of the workflow and unit of record. |
| How much and how long? | Approximate record counts by year, date range, growth or refresh rate and relevant entity counts. |
| What is connected? | Schema, stable identifiers, timestamps, linked systems and data dictionary. |
| Where did it come from? | Source-system map, collection process, provenance notes and transformation history. |
| What is included or excluded? | Field list, modalities, versions, exclusions and known third-party material. |
| What use might it support? | A task hypothesis and evidence that would test it—not an unsupported promise of model improvement. |
OpenAI’s public data-partnership description provides one example of a buyer explaining interest in non-public archives and domain-specific material, as well as the need to clean or digitise some sources. It is an example of one organisation’s stated approach—not a universal list of buyer requirements. OpenAI Data Partnerships.
2. Turn each commercial claim into testable evidence
A buyer's first diligence question is often not “How many records?” but “What would an eligible example actually contain?” The inventory should distinguish raw records, reconstructed cases, cases with usable outcome labels and cases the provider may lawfully share. These populations need not be equal. A million ticket messages may represent only a fraction as many distinct cases, and fewer still may have a trusted resolution. Giving a denominator for each stage prevents the appearance of quantity from replacing usable evidence.
The seller should be able to reproduce counts from the source. That means documenting an extract's selection date, time window, deletion rules, de-duplication method, join logic and known unobserved groups. Counts should reconcile to the systems of record rather than an informal spreadsheet. Where a quality finding is based on a sample, the sampling approach matters. A hand-picked set of the cleanest cases can demonstrate format and context, but cannot establish archive-wide completeness or failure rates. Representative findings require an appropriate sampling frame and a transparent treatment of cases the extraction could not access.
Buyers should also distinguish administrative and substantive labels. A work order marked “complete” may mean the technician submitted a form, not that the equipment remained operational for 30 days. A claim marked “approved” may encode a human decision rather than the final financial outcome. Relabelling such fields as verified success introduces a measurement error precisely where the dataset appears most valuable. A defensible pack retains the original state label and, where available, a separately defined verified outcome, including the observation period and subsequent corrections.
A credible diligence pack follows a simple discipline:
CLAIM → EVIDENCE → LIMITATION → DECISION
Exhibit 2. Claim–evidence–limitation diligence matrix
| Claim | Evidence a buyer can inspect | Limitation to disclose | Decision supported |
|---|---|---|---|
| “The archive covers eight years.” | Counts by year, source and record type. | Two months of a source system are absent. | Whether the period is sufficient for the proposed task. |
| “Outcomes are verified.” | Outcome definitions, audit method and sample review. | Older records use administrative closure rather than verified success. | Whether to use all years or restrict the pilot to later records. |
| “Cases can be linked across systems.” | Join keys, matching method and linkage-quality checks. | Some legacy identifiers were reused. | Whether linked data are reliable enough for workflow-level use. |
| “The data are ready to deliver.” | Sample export, schema, security and delivery plan. | Free-text notes need additional review and redaction. | Whether to proceed with a sample or scope preparation work first. |
| “The archive is distinctive.” | Evidence of specialist context, rare workflow, longitudinal history or verified outcomes. | The uniqueness claim has not been tested against alternatives. | Whether the buyer sees fit worth exploring—not a guaranteed price. |
Good diligence is not the art of hiding imperfections. It is the practice of showing what is known, how it was checked and where confidence stops. Government guidance on commercialising knowledge assets likewise treats rights, obligations, access models and market fit as part of due diligence, not merely promotional material. UK Knowledge Asset Commercialisation Guide.
3. Demonstrate quality, context and provenance
For each quality indicator, the method should be described with enough precision for an independent reviewer to replicate it. Consider a customer-support archive that reports 92% label completeness. Does this mean the status field was filled in on 92% of messages, or that 92% of unique cases have a final expert-confirmed outcome? The first statistic may be perfectly correct yet irrelevant to a buyer testing successful resolution. A quality dashboard should specify the unit measured, denominator, segmentation and validation rule for every metric, preferably by source system and date.
The buyer may also be testing whether records preserve information that would have been available at the time of decision. An extract that links an initial diagnosis with a fault discovered months later can be valuable for retrospective evaluation, but using that later discovery as an input in a simulated historical decision creates information leakage. Record both event time and label-observation time. Similarly, a code that was valid under a 2020 policy may no longer have the same meaning after a 2023 process redesign. Version information and provenance explain differences that would otherwise be mistaken for model errors or improvements.
Dataset documentation provides established ways to describe composition, collection, recommended uses and limitations. Gebru and colleagues' Datasheets for Datasets proposes structured documentation questions, while W3C PROV provides a formal vocabulary for entities, activities and agents involved in data provenance [3, 4]. Neither method substitutes for actual field-level verification. The practical opportunity is to organise the seller's evidence so an AI developer can assess claims without repeatedly reconstructing how the business worked.
Quality evidence should relate to the proposed use. A buyer considering classification may care about label definitions and coverage of difficult classes; one considering retrieval may care about freshness, document authority and versioning.
Exhibit 3. Evidence for quality, provenance and outcome claims
| Dimension | Example evidence |
|---|---|
| Completeness | Missing-field rates by field, source and time period. |
| Consistency | Label definitions, allowed values, reviewer guidance and history of definition changes. |
| Duplication | Exact and near-duplicate rates, with the method used to detect them. |
| Linkage | Percentage of records matched across systems, plus known false-match and unmatched cases. |
| Outcome quality | Whether outcomes are verified, inferred, self-reported or administrative. |
| Time integrity | Timestamp meanings, time zones, event-order checks and later-edit handling. |
| Coverage | Distribution by operating condition, equipment version, geography, language or case type. |
| Provenance | Original source, extraction date, transformation steps and data version. |
Worked example: explain a workflow, not just a count
Exhibit 4. An auditable candidate dataset description
| Snapshot field | Example description |
|---|---|
| Asset | Case-level industrial maintenance archive, with linked technician notes and selected sensor observations. |
| Grain | One row per work order; related events may be linked by asset and work-order identifiers. |
| Coverage | January 2018–December 2025; three equipment families and four service regions. |
| Outcomes | Recorded return-to-service, repeat failure within 30 days and parts replacement. |
| Context | Equipment model, operating condition, technician role, timestamps and procedure version. |
| Known limits | Two months of missing telemetry, workflow redesign in 2021 and incomplete outcome coding in older records. |
| Restrictions and exclusions | Direct identifiers removed from the sample; confidentiality review pending for free text; third-party manuals excluded unless separately cleared. |
This is a starting inventory, not proof that every field is accurate or that the data may be licensed for any use. Each claim needs supporting records, a method and an owner who can answer questions.
For provenance, preserve distinctions such as system-recorded versus analyst-interpreted codes, contemporaneous versus later-edited notes, procedure version at event time versus current procedure, measured versus imputed values, and missing versus explicit “no”. NIST guidance identifies provenance and metadata as important to assessing data quality and reliability. NIST Research Data Framework.
Worked evidence-yield calculation: from messages to qualifying cases
Suppose a company reports 240,000 support messages and asks whether an AI buyer could use them to evaluate durable case resolution. For illustration only, assume records can be grouped into 80,000 distinct cases using stable identifiers. A reconciliation then finds that 60%, or 48,000 cases, have an observed resolution state; 75% of these, or 36,000, have the required request-and-action context. If a legal and privacy screen determines that 80% of that remaining population is eligible for the proposed restricted pilot, the candidate pool is 28,800 cases, before any holdout design or independent outcome audit. The yield is 28,800 ÷ 80,000 = 36% of distinct cases, or 12% of original message count, but these denominators answer different questions.
The percentages and counts are invented, not industry benchmarks. They also assume the checks can be applied sequentially to the same population. If outcome observation is weaker in older years, the final pool may be concentrated in later cohorts; if repeated problems are disproportionately common, the cases may not offer adequate diversity. A sensitivity test could lower the context-qualified rate from 75% to 50%, reducing eligible cases to 19,200 (48,000 × 50% × 80%). Neither number proves an improvement in model performance. They describe an evidence pool whose sampling, labels and permitted use still require validation. This is why buyers need the numerator, denominator, period and method for every conversion.
4. Scope rights, sample access and delivery honestly
A useful rights review is record- and use-specific. A customer contract may permit internal service operations but say little about third-party AI training; supplier manuals may be independently protected; free-text notes may include personal information or privileged communications. The review should trace collection purpose, contracting entity, contributors, restrictions and proposed recipient use. Records should be categorised as cleared for a defined purpose, requiring further review, or excluded. A blanket “all proprietary data” designation is neither a rights analysis nor a data-protection assessment.
Sampling should minimise irreversible disclosure. A synthetic schema or empty example can establish technical fit without revealing confidential data, although a synthetic example is not evidence of real distributions. A de-identified sample can test format and some semantics, but the ability to link records or recognise unusual people or incidents may remain. Pseudonymisation does not by itself render personal information anonymous under data-protection law. The parties should select the smallest lawful and useful disclosure that answers the next diligence question, with rights, retention and access controls set before exchange.
There is also an operational cost asymmetry. The seller may spend weeks extracting, redacting and reconciling a sample while the buyer is only evaluating whether its use case merits a pilot. Management should gate spending: first validate a buyer's task and ingestion capability using metadata; then fund a small controlled sample; only then commit to expensive transformation or recurring feeds. The same sequential discipline protects both sides from the illusion that every attractive archive is immediately marketable.
Technical quality does not establish permission. A buyer may ask about customer terms, employee or contractor contributions, confidentiality, personal data, regulated information, third-party material and prior licences.
A concise rights summary should state what has been reviewed, what remains unresolved and what is explicitly excluded. Avoid categorical claims such as “we own everything” or “the data are fully anonymised” unless the evidence and legal analysis support them.
For personal-data sharing, the parties need to establish their real roles, purposes and responsibilities. The ICO recommends documenting the organisations involved, precise purposes, data items, access, security, retention and individual-rights procedures in suitable arrangements. ICO data-sharing agreements.
A staged sample protocol
- Agree the question. Specify the task and what evidence the sample should reveal.
- Choose a balanced slice. Include common cases, difficult cases, relevant time periods and operating conditions—not only polished successes.
- Minimise. Remove fields and free text that are not needed for the test; apply appropriate review and access controls.
- Keep it traceable. Record sample criteria and source versions so the sample can be reconciled with the inventory.
- Define handling. Agree who can access it, permitted use, retention and deletion or return.
- Review results. Note what the sample demonstrated, what it did not, and what further work would be required.
The sample must not be presented as representative unless its selection method supports that claim. Nor should a sample be sent before rights, privacy and security questions are appropriately addressed.
Delivery readiness
Practical delivery evidence should identify where records reside, who may authorise extraction, how exports can be reproduced, which preparation operations are required and whether the sample can be reconciled to the master inventory. A seller should also specify transfer method, access permissions, logging, retention, repeat delivery capability and any excluded systems. A technically impressive archive that cannot be exported without interrupting critical operations may not be ready for a paid pilot. These issues are different from questions about whether the buyer's model can use the content once received.
5. What “special” means—and when to pause
The most informative diligence answer is sometimes a bounded limitation. If pre-2021 cases have no reliable outcome coding, a pilot can focus on 2022–2025 while the older period remains available only for linguistic or contextual analysis. If the source has records from five subsidiaries but only two possess a defensible rights chain, the eligible asset should be described as those two. Narrowing scope may increase usefulness by reducing uncertainty even as the nominal record count declines.
A pilot must answer a falsifiable question rather than demonstrate that a vendor can display the data. For example, measure whether a model trained or evaluated with the additional fault cases identifies escalation-worthy incidents more reliably on held-out later-period cases than the existing baseline. Define the unit of evaluation, sample exclusions and an agreed acceptance metric before inspecting results. The buyer's development data should not overlap with the held-out evaluation population in ways that invalidate the comparison. The seller need not promise improved model performance; it needs to supply reliable evidence with which that proposition can be tested.
A practical executive review should price the preparation path. Allocate staff time to legal and privacy review, technical extraction, linkage validation, redaction, documentation, secure delivery and follow-up support. Estimate costs under a clean-sample case and a difficult-data case, then identify which tasks the buyer could fund. Reject a transaction where the cost to establish lawful and reproducible evidence exceeds plausible benefit, even if the raw dataset appears rare. Disciplined refusal is an outcome of diligence, not evidence that the archive lacks operational value.
A dataset is not commercially distinctive simply because it is large or internal. Its potential relevance may come from hard-to-reproduce characteristics such as a specialised workflow, verified outcomes linked to inputs and decisions, rare operating conditions, expert interventions, contextual metadata, a documented rights path or a provider’s ability to maintain and explain the data.
The claim must connect to a buyer task. “We have eight years of data” does not show that the archive teaches, tests or supplies a capability. “We can identify equipment type, fault, action, procedure version and verified repeat-failure outcome for a defined subset” is more testable.
Diligence red flags
- unexplained record-count changes;
- stable-looking fields whose definitions changed over time;
- closure codes described as verified outcomes;
- joins that cannot be reproduced or audited;
- unknown third-party material or unclear contributor rights;
- a sample that cannot be reconciled with the documented inventory;
- “fully anonymised” claims without a documented assessment;
- quality metrics with no method, denominator or time period;
- a delivery promise that ignores extraction, security or review work;
- a claim of uniqueness based only on volume.
Executive takeaway
Prepare for diligence as if a buyer will test every important claim. A specific, candid pack can support a useful pilot even when the data have gaps. A vague claim of scale cannot compensate for missing context, unclear rights or unverified outcomes.
Buyer-ready minimum pack: asset description; inventory and schema; data dictionary; coverage and quality profile; provenance and transformation notes; rights and restriction summary; representative sample protocol; delivery and security plan; known limitations; and a named contact for technical and governance questions.
Sources and further reading
- OpenAI (2023). Data Partnerships, description of the company's published interest in public and private datasets; a historical indication, not confirmation of current procurement demand. https://openai.com/index/data-partnerships/
- OpenAI. Data Partnerships enquiry form, describing information sought from prospective data partners, including size, rights and preparation needs. https://openai.com/form/data-partnerships/
- Gebru, T. et al. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92. https://doi.org/10.1145/3458723
- W3C (2013). PROV-O: The PROV Ontology, W3C Recommendation. https://www.w3.org/TR/prov-o/
- UK Government. The Knowledge Asset Commercialisation Guide. https://www.gov.uk/government/publications/knowledge-asset-commercialisation-guide/the-knowledge-asset-commercialisation-guide
- UK Information Commissioner's Office. Data sharing agreements (guidance under review following legislative changes). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/data-sharing-agreements/
- National Institute of Standards and Technology. Research Data Framework, SP 1500-18r2. https://nvlpubs.nist.gov/nistpubs/SpecialPublications/1500-18/NIST.SP.1500-18r2.html
- National Institute of Standards and Technology. AI RMF Playbook: Measure (task-specific assessment guidance). https://airc.nist.gov/airmf-resources/playbook/measure/