Executive perspective
Executive perspective
The journey from identifying a potentially valuable dataset to establishing an AI data partnership involves several distinct commercial, technical and legal decisions. Each should reduce uncertainty before a company commits significant resources or discloses sensitive information.
Four principles govern a disciplined process. Buyer relevance must be established before extensive preparation; a controlled sample must demonstrate something measurable; rights and confidentiality must be examined before access is granted; and commercial terms must reflect the full cost and risk of delivery.
A successful outcome may be a recurring licensing arrangement, a restricted evaluation project or a decision not to proceed. The objective is not simply to complete a transaction. It is to discover whether a defined, permissible and commercially worthwhile use of the company's data exists.
1. Discovery: identifying the asset and its potential use
A typical partnership discussion begins with a deceptively simple question: what data does the company possess?
For many operating businesses, the answer is less straightforward than management expects. Data may be spread across CRM systems, ERP platforms, operational databases, document repositories, communication tools and historical backups. Individual departments often understand their own records but lack visibility over the organisation's wider information assets.
The first stage is therefore an inventory, not an extraction exercise.
An initial inventory should establish which systems contain relevant records, how many years of history exist, what business activities are represented and whether the records can be linked to meaningful outcomes.
For example, a field-services company may hold ten years of service orders, technician notes, equipment histories, spare-parts consumption and customer complaints. These records could provide evidence of how technicians diagnose equipment failures and resolve operational problems.
But their existence does not establish demand.
A buyer developing general conversational capabilities might find little value in them. A developer building a specialist maintenance assistant could potentially find the same records interesting, particularly if they preserve diagnostic decisions and verified repair outcomes.
This distinction should shape the first commercial discussion.
Instead of approaching prospective partners with a generic proposition such as "we have ten years of operational data", the company should articulate a preliminary use-case hypothesis:
Our records may contain linked examples of equipment faults, diagnostic decisions, interventions and observed outcomes that could support the development or evaluation of specialist maintenance AI.
The hypothesis should explain what an AI system might learn, which evidence supports that possibility and what remains unproven.
OpenAI's published Data Partnerships programme, for example, describes an interest in selected non-public material, including data expressing human intention, while also identifying restrictions concerning sensitive and third-party information 1. This provides evidence of a particular developer's interests, not a universal procurement standard or guaranteed market for enterprise records.
Discovery should also establish whether the proposed partner is a suitable counterparty. A foundation-model developer, specialist AI vendor and systems integrator may have different requirements and technical capabilities. They should not be treated as interchangeable buyers.
The OECD's work on data sharing emphasises that the benefits and acceptable conditions of reuse depend on context, stakeholder interests and the nature of the data 2. Consequently, a partnership assessment must consider the specific buyer and use rather than assume that data has a single intrinsic licensing value.
2. Initial assessment: building evidence before sharing data
Once a plausible use has been identified, the next step is to determine whether the records can support it.
This is primarily a feasibility assessment, rather than a pricing exercise.
The company should investigate three interconnected questions.
First, is the evidence technically useful? Operational records need sufficient structure, context and consistency for the intended application. An archive of disconnected customer messages may have limited value for evaluating problem resolution, while linked cases with expert decisions and observed outcomes may provide stronger evidence.
Second, is the proposed use permissible? The company may possess the records but remain subject to customer confidentiality, intellectual-property restrictions, employment obligations, data-protection requirements or contractual limits on onward transfer.
Third, does the possible economic benefit justify further work? Data preparation, legal review, expert annotation, security and delivery can impose costs before any commercial agreement exists.
These questions should be answered progressively rather than through a single expensive investigation.
Exhibit 1: Five decision gates in an AI data partnership
| Gate | Principal question | Evidence required | Accountable functions | Decision |
|---|---|---|---|---|
| 1. Buyer relevance | Can the records address a defined AI task? | Use-case hypothesis and preliminary inventory | Commercial, operations | Continue or decline |
| 2. Rights feasibility | Is there a defensible path to the proposed use? | Initial rights, confidentiality and privacy screening | Legal, privacy | Proceed, narrow or stop |
| 3. Dataset quality | Can relevant signals be demonstrated reliably? | Profile, schema, quality measures and approved sample | Data, operations | Prepare a pilot or defer |
| 4. Technical and economic viability | Can improvement be measured at acceptable cost? | Pilot results, cost model and risk assessment | Technical, finance | Negotiate, revise or stop |
| 5. Operational delivery | Can the agreement be fulfilled securely and consistently? | Contract, security controls and delivery plan | Legal, security, operations | Launch or defer |
This is an OpDataLab analytical framework, not a mandatory regulatory process or a representation of a particular AI buyer's internal workflow.
The gates need not occur in an entirely linear sequence. Rights screening and quality assessment may proceed in parallel, while early buyer feedback can refine the use-case hypothesis.
However, their order reflects an important economic principle: avoid incurring substantial irreversible costs before establishing basic relevance and permission.
NIST's AI Risk Management Framework supports this emphasis on documenting intended uses, third-party data risks, relevant controls and evaluation methods 3.
A company should expect to make several decisions not to proceed. An early rejection can be economically rational if it prevents expensive preparation for an unsuitable buyer or an impermissible application.
3. Controlled samples and buyer diligence
A serious buyer may eventually require evidence beyond a verbal description of the dataset.
This does not mean the company should immediately send raw customer records or export its entire archive.
The first exchange can often be limited to descriptive material: source-system information, approximate record counts, field definitions, historical coverage, missingness statistics, language distribution and known restrictions.
These materials allow a prospective partner to assess relevance without obtaining the underlying records.
If the proposition remains promising, the parties may agree to a controlled sample. The sample's purpose should be specified before any access is granted, with appropriate confidentiality, rights and security arrangements in place.
A sample used to inspect data structure is different from one authorised for model development. Access for a limited feasibility assessment should not be presumed to authorise unrestricted training, retention or onward transfer.
Sample design requires particular care. A convenient extract containing only clean, recent records may not represent a ten-year archive. Similarly, a sample containing only difficult cases might exaggerate the density of specialist expertise.
The company should disclose how samples were selected, what they exclude and whether any transformations have changed their meaning.
Exhibit 2: Illustrative data-room documentation package
| Document | What it establishes | Frequent weakness |
|---|---|---|
| Dataset inventory | Systems, history, record counts, entities and owners | Incomplete source coverage |
| Data dictionary | Field definitions, types, allowed values and changes over time | Undefined categories and inconsistent terminology |
| Provenance record | How records were collected, transformed and linked | Missing history of source-system changes |
| Quality profile | Completeness, duplication, outcome coverage and known errors | Aggregate metrics conceal weak subsets |
| Rights and restrictions register | Contractual, confidentiality, privacy and third-party constraints | Assumptions of unrestricted ownership |
| Sample methodology | Selection criteria, exclusions and representativeness | Attractive sample differs from production archive |
| Security and delivery outline | Access controls, transfer options and retention constraints | No approved mechanism for controlled access |
A useful package should permit an informed decision without creating an unnecessarily broad disclosure.
Documentation also establishes accountability. If a buyer later discovers that a timestamp represented case closure rather than a verified resolution, the discrepancy should be traceable to a documented definition, not an unexplained interpretation.
A well-prepared company should distinguish between data it can describe, data it can expose for restricted evaluation and data it can lawfully license for a specified ongoing use.
These categories may be very different.
An archive may look technically attractive while containing extensive customer material whose onward licensing was never contemplated. In that case, the feasible product might involve aggregated statistics, independently reviewed derived data or controlled access rather than unrestricted raw-record delivery.
4. Rights diligence and the design of a meaningful pilot
Rights and privacy diligence should begin before any disclosure requiring legal permission. More detailed review can then proceed alongside technical evaluation, with access expanding only when the relevant safeguards are established.
The central legal question is not simply who owns the database. It is which rights and permissions permit the specific proposed use.
A company may own the system infrastructure while customer contracts restrict secondary processing. Supplier manuals may be protected by third-party copyright. Employee records may involve separate privacy considerations. Historical customer communications may contain confidential information whose transfer would exceed reasonable contractual expectations.
The diligence process should examine the actual rights chain, contractual arrangements, jurisdictions, relevant personal data and intended uses.
For UK personal-data sharing, the Information Commissioner's Office recommends a risk-based assessment and explains that data-sharing agreements help document purposes, responsibilities and controls. It also recommends considering a Data Protection Impact Assessment, with one required where processing is likely to result in high risk to individuals 4.
These considerations are not interchangeable with general commercial confidentiality provisions. A non-disclosure agreement does not itself establish the lawful basis for processing personal data or permission to license third-party intellectual property.
Cross-border transfers may introduce additional requirements. Where applicable, EU or UK data-protection rules may require an appropriate international transfer mechanism, such as relevant standard contractual clauses, together with any necessary supporting safeguards 5.
Requirements differ by jurisdiction and the circumstances of the transaction.
Designing the pilot
Once the relevant disclosure is authorised, a pilot can test whether the data actually improves a specific AI task.
Consider the field-services example. A specialist AI developer claims that access to historical technician cases might improve an assistant's fault-triage performance.
A controlled pilot could compare a baseline system against a version using approved dataset material, while reserving independent cases for evaluation.
The parties should agree the questions in advance.
What constitutes a correct diagnosis? When should the assistant escalate to a human? How will outdated procedures be identified? What errors create safety or commercial risk? How will improvement be assessed across equipment types and unusual failure modes?
A metric such as classification accuracy may be insufficient if incorrect recommendations can create material operational risk.
The pilot should identify relevant performance measures, a defensible baseline, permitted processing, sample size, evaluation independence, access restrictions and deletion or retention arrangements.
It should also establish who bears the cost of annotation, testing and specialist review.
The outcome should be a documented decision, not merely an encouraging demonstration. A technically successful pilot could still fail economically if scaling the dataset requires extensive manual review.
Likewise, a pilot that produces no meaningful performance improvement may be valuable if it demonstrates that an external licence is not worth pursuing.
A restricted evaluation agreement should specify whether the buyer may use the material only to assess performance or may also train or modify models. These rights cannot safely be inferred from the word pilot.
5. Commercial structure: turning technical potential into an economic agreement
If the pilot indicates a viable use, the commercial discussion becomes more specific.
The buyer is no longer acquiring an undefined collection of records. It is considering access to a documented data product for a defined purpose.
Possible structures include a one-time licence to a historical archive, time-limited evaluation access, a recurring data feed or an arrangement involving preparation and ongoing specialist services.
Each creates different economic obligations.
An historical extract may require substantial preparation but limited ongoing updates. A recurring feed can generate repeated delivery, maintenance, monitoring and quality-control work. An exclusive licence may constrain future uses and should therefore be evaluated against the options being surrendered.
The seller's cost base matters, but it should not be confused with market value.
A company's expenditure on preparing data does not prove that an AI buyer will pay enough to recover it. Conversely, a buyer might obtain substantial value from a carefully documented niche dataset whose preparation cost was relatively modest.
The commercial assessment should therefore consider both expected buyer benefit and the supplier's cost, contractual exposure and alternative uses.
Exhibit 3: Illustrative economics of a first-year licensing arrangement
Assume a company prepares a restricted dataset for an AI partner.
| Preparation component | Hypothetical assumption | Cost |
|---|---|---|
| Engineering and extraction | 100 hours × £90 | £9,000 |
| Domain-expert review | 120 hours × £60 | £7,200 |
| Legal and privacy diligence | 50 hours × £130 | £6,500 |
| Security and validation | 45 hours × £100 | £4,500 |
| Secure pilot infrastructure | Fixed allowance | £3,500 |
| Direct preparation cost | £30,700 | |
| Contingency | 20% | £6,140 |
| Total preparation requirement | £36,840 | |
| First-year support and monitoring | Illustrative allowance | £12,000 |
| First-year cost requirement | £48,840 |
All figures are hypothetical planning assumptions, not observed licensing prices or industry benchmarks. Costs are assumed to be fully attributable to the arrangement and exclude taxes, financing, opportunity costs and unpriced liabilities.
Three hypothetical annual licence fees produce different first-year contributions:
| Illustrative annual fee | Less first-year cost requirement | Illustrative contribution |
|---|---|---|
| £35,000 | £48,840 | −£13,840 |
| £50,000 | £48,840 | £1,160 |
| £70,000 | £48,840 | £21,160 |
The calculation is straightforward, but its implications are significant.
At £50,000, the arrangement produces little first-year contribution despite appearing to generate meaningful new revenue. A modest increase in expert-review effort could eliminate the surplus.
Preparation costs might be spread over future agreements if the dataset can legally and technically be reused. But renewal revenue should not be assumed, and rights granted to the first buyer may limit future opportunities.
Similarly, a proposed exclusive agreement should be assessed against the economic value of foregone future licensing opportunities, even where those opportunities are uncertain.
A sound negotiation must therefore address the allocation of costs and risk, not just the licence fee.
Exhibit 4: How deal structure changes commercial exposure
| Contractual dimension | Restricted evaluation | Non-exclusive production licence | Exclusive or highly restricted licence |
|---|---|---|---|
| Permitted purpose | Defined pilot or benchmark | Specified ongoing use | Specified use with restrictions on other licences |
| Duration | Short and controlled | Fixed term or renewal | Negotiated term and scope |
| Data delivery | Limited sample or controlled access | Approved archive or updates | Agreed licensed package |
| Revenue structure | Potentially paid pilot or evaluation fee | Fixed fee, recurring fee or negotiated variable component | Negotiated consideration reflecting exclusivity |
| Retention and derivatives | Narrowly specified | Detailed rules for source and derived materials | Particularly important to future optionality |
| Seller exposure | Bounded if controls are effective | Ongoing delivery and compliance obligations | Additional opportunity cost and negotiating constraints |
These are illustrative contractual dimensions, not standard market terms.
Particular attention should be given to model-training rights, derivative datasets, retention, onward transfer, permitted affiliates, subcontractors, warranties, liability limits, security incidents, audit rights and termination.
Deletion obligations need careful drafting where models have already been trained on licensed material. Deleting source files is not necessarily equivalent to reversing their contribution to model parameters.
In relevant EU contexts, general-purpose AI model providers also face scoped documentation, copyright-policy and training-content-summary obligations under Article 53 of the AI Act. These are obligations of covered model providers, not a universal checklist automatically imposed on every enterprise dataset supplier 6.
The seller should understand the buyer's downstream compliance needs without assuming that every regulatory requirement applies to the transaction.
6. Secure delivery, ongoing governance and the decision to proceed
Signing a licence does not eliminate operational responsibilities.
The data product must still be prepared, checked and delivered through the agreed mechanism.
Preparation may involve deduplication, field filtering, transformation, pseudonymisation, annotation, validation and conversion into a consistent format. Each transformation should be documented so the buyer understands what has changed.
A production arrangement may also require repeatable checks for new records, stable schema versions, incident procedures and agreed methods of handling corrections.
NIST SP 800-47 Revision 1 provides guidance on protecting information exchanges before, during and after transfer, including arrangements for managing responsibilities and security controls 7.
For a recurring feed, change management becomes particularly important.
If a company replaces its CRM, updates a ticket taxonomy or changes how outcomes are recorded, the dataset may become less comparable over time. An apparently routine operational change could undermine the use for which the licence was negotiated.
Both parties should therefore understand who monitors quality, reports deviations, resolves incidents and approves changes to the delivered material.
This requires clear internal ownership.
The commercial lead coordinates the partnership proposition and negotiation. Operations explains the underlying processes. Engineering prepares and validates the records. Legal and privacy teams assess permissions. Security oversees the approved delivery mechanism. Finance evaluates cost, pricing and financial exposure.
One accountable project owner should maintain the decision record and ensure that technical and commercial commitments remain aligned.
When the correct decision is to stop
A partnership process should include explicit circumstances in which the company pauses, narrows or declines the opportunity.
These include an inability to verify rights, unsuitable buyer use, unmanageable confidentiality risk, unreliable outcome labels, excessive preparation cost or security requirements the provider cannot reasonably meet.
An additional concern arises when a prospective partner seeks broad, perpetual reuse rights before establishing the value of the dataset.
Such a request does not automatically make a transaction unacceptable, but it changes the risk profile and should not be accepted merely as a preliminary evaluation convenience.
Management should be particularly cautious where the dataset embodies proprietary operational expertise that the company continues to use competitively.
The value of retaining exclusive internal access may exceed the benefit of licensing that knowledge externally.
This is ultimately a capital-allocation decision. Technical feasibility is necessary but insufficient. The project must offer an acceptable return after considering preparation, ongoing obligations, legal risk, strategic opportunity cost and realistic alternatives.
Practical implications
A company approaching an AI data partnership should be able to demonstrate a coherent chain of evidence: a defined buyer task, relevant operational records, documented quality, defensible permission, measurable technical utility and commercially acceptable delivery obligations.
The process should move from a low-cost inventory and initial assessment to increasingly specific tests and commitments. The company should not undertake a substantial transformation project merely because a buyer has expressed curiosity, nor should it disclose sensitive material before the intended use and necessary permissions are established.
For management, the most useful output is often a controlled decision: proceed to a restricted pilot, revise the proposition, explore internal use or stop.
The objective is not to move from an archive to a licence as quickly as possible. It is to reduce uncertainty at each stage until the value, cost and risk of the proposed transaction can be judged responsibly.
Sources and further reading
- OpenAI (2023). OpenAI Data Partnerships. Describes a specific developer programme and its stated interests, not general industry procurement standards. https://openai.com/index/data-partnerships/
- OECD (2019). Enhancing Access to and Sharing of Data: Reconciling Risks and Benefits for Data Re-use across Societies. OECD Publishing. https://doi.org/10.1787/276aaca8-en
- NIST (2023). Artificial Intelligence Risk Management Framework 1.0, Core. https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
- UK Information Commissioner's Office. Data Sharing: A Code of Practice, including guidance on deciding to share data and data-sharing agreements. The ICO notes that parts of this guidance are under review following the Data (Use and Access) Act. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/
- European Commission. Standard Contractual Clauses for International Transfers. https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/standard-contractual-clauses-scc_en
- European Union (2024, as amended). Artificial Intelligence Act, Article 53, obligations for providers of general-purpose AI models. https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- NIST (2021). SP 800-47 Revision 1: Managing the Security of Information Exchanges. https://csrc.nist.gov/pubs/sp/800/47/r1/final