Executive perspective
Executive perspective
An enterprise data partnership depends not only on the usefulness of the underlying records, but also on whether they can be extracted, prepared and delivered reliably.
Four considerations are central. The intended AI application determines what must be transferred; data preparation must preserve useful context while respecting rights and confidentiality; secure delivery requires both technical and organisational controls; and successful transfer does not guarantee that the resulting dataset is usable.
For a one-time evaluation, a company may need only a restricted, documented sample. A recurring licensing arrangement may require an operational pipeline capable of handling new records, schema changes, corrections and security obligations.
The delivery process is therefore part of the commercial asset. A dataset that cannot be reproduced, validated or maintained may be expensive for an AI partner to use, regardless of its apparent information value.
1. Starting with purpose, scope and source-system inventory
The practical starting point is an authorised delivery specification, not a database export. It should define the unit of delivery (a service case, document, event or linked workflow), relevant periods and source systems, fields necessary for the agreed AI task, excluded categories, recipient permissions, retention, and the evidence required for acceptance. A retrieval application may need current controlled documents; an evaluation of troubleshooting may instead require past diagnostic actions and observed outcomes. The distinction determines which records can be selected and how they must be connected.
A feasibility sample, a held-out evaluation set and a dataset licensed for training need not contain the same records or carry identical permissions. The provider and recipient should therefore agree a versioned input contract recording permissible use, minimum schema, target coverage, required context, delivery method, expected output and acceptance procedure. The document is operational: an engineer should be able to convert it into reproducible extraction instructions, while an accountable information owner can identify what must never be included.
A well-scoped pilot avoids an expensive full export, interference with business systems and exposure of irrelevant material. Chapter 9 considers the commercial diligence gates; the issue here is how the agreed scope becomes a controlled technical package.
Locating the records
Enterprise operational data is rarely held in one convenient repository.
Consider a company providing commercial equipment maintenance. Its historical information may be distributed across four systems:
- A service-management platform records customer complaints, engineer assignments and ticket statuses.
- An ERP system records equipment models, replacement components and service contracts.
- A document repository contains technical notes, inspection forms and diagnostic photographs.
- An asset-management platform records maintenance history, installed firmware and subsequent breakdowns.
An apparently simple request for ten years of equipment-failure data may therefore require the company to connect several different forms of information.
The integration problem is not just technical. The systems may use different customer identifiers, equipment references and timestamp conventions. A service case might refer to a product family rather than the exact equipment model. Historical ERP migrations may have replaced identifiers or changed business rules.
Without resolving these issues, the receiving partner may mistakenly interpret separate events as independent failures or treat administrative ticket closure as a successful repair.
A useful inventory therefore records the origin, ownership, structure and dependencies of each source.
Exhibit 1: Illustrative source-system inventory
| Source system | Relevant evidence | Technical dependency | Principal risk |
|---|---|---|---|
| Service-management platform | Fault reports, technician actions, escalation history | Stable case identifiers and event timestamps | Incomplete or ambiguous outcomes |
| ERP | Equipment specifications, parts and contractual information | Equipment and customer mapping | Restricted commercial information |
| Document repository | Technical procedures, inspections and attachments | Document versions and case references | Confidential and third-party content |
| Asset-management system | Repairs, recurring faults and subsequent equipment performance | Asset identifiers and longitudinal linkage | Historical gaps and inconsistent records |
Illustrative framework for a maintenance-data partnership. The appropriate source systems depend on the business and intended AI use.
The inventory should also record which team owns each system, who can authorise extraction, the date range available and whether the information can be accessed without disrupting normal operations.
This last point matters. Running large queries against a production database may interfere with operational workloads. Historical archives may be stored in formats that require specialist tools or restoration work.
An initial assessment should therefore estimate extraction effort before committing to delivery.
NIST's AI Risk Management Framework emphasises documenting intended uses, relevant information sources and risk-management responsibilities. Although voluntary, it provides a useful foundation for structuring this early assessment 1.
2. Extraction and transformation: creating a usable dataset without losing meaning
Once the scope and sources are understood, the company can design an extraction process.
Possible approaches include database queries, application programming interfaces, scheduled exports, cloud-storage transfers or outputs generated by an internal data engineering team.
The choice depends on scale, system capabilities, security requirements and whether the arrangement involves a one-time archive or recurring delivery.
For a pilot, a controlled export may be sufficient. For an ongoing partnership, manually repeating multiple exports may become expensive and unreliable.
Preserving relational context
Suppose the maintenance company's service database contains a record showing that a technician replaced a component on 12 March 2024.
That action cannot be interpreted reliably without knowing which equipment was involved, the reported symptoms, the applicable product version, the diagnostic tests performed and whether the intervention resolved the fault.
The extraction must preserve the relationships required for the intended AI task.
It may also need to distinguish event time from recording time. A technician could complete work on Monday but submit the report on Wednesday. Treating the submission timestamp as the repair date can distort calculations of response time or operating performance.
The same problem arises when data is updated retrospectively. A record extracted today may contain a final classification that was unavailable when the original decision was made.
For AI training or evaluation, this distinction can be material. A dataset may inadvertently expose information from the future relative to the decision the model is supposed to reproduce.
A sound extraction should therefore document relevant timestamp meanings, entity relationships, historical versions and known limitations.
Transforming the records
Raw operational systems are designed to support business processes, not necessarily external AI development.
Records may contain inconsistent categories, duplicate events, invalid dates, redundant attachments or unsupported file types.
Preparation can involve:
- Normalising formats and categorisations.
- Reconciling identifiers across systems.
- Deduplicating repeated events.
- Filtering irrelevant records.
- Preserving or reconstructing necessary context.
- Removing or restricting sensitive fields.
- Producing agreed output formats.
Each operation can alter the evidence.
For example, aggressive deduplication might remove legitimately repeated equipment failures. Merging two ticket categories may conceal important differences in fault severity. Removing timestamps can make it impossible to reconstruct a workflow sequence.
Transformation decisions should therefore be traceable and reversible where feasible.
Rights and privacy controls as extraction rules
Rights and privacy decisions must be converted into field- and record-level processing instructions. A rights schedule can identify excluded customers, third-party attachments, restricted date ranges, masked attributes and permissible recipient uses. The transformation log should record which rule removed or changed an item and how exceptions were reviewed, without exporting the sensitive evidence used to make that decision. Chapters 14 and 15 explain the underlying privacy and legal tests; delivery engineers must implement only the resulting approved scope.
Pseudonymisation may reduce identification risks, but it does not automatically render personal data anonymous under ICO guidance 2. Similarly, hosting information on a UK server does not eliminate transfer requirements where a separate overseas organisation can access UK personal data 3. The access design must match the actual recipient and permitted processing, not just the location of the files.
Where permitted, controlled computation or a restricted evaluation environment may avoid handing over a complete extract. Decisions to mask, exclude or retain context must be tested against the agreed AI task; otherwise the delivered package can be compliant with the extraction rules yet unusable.
3. Documentation: making the information intelligible to an external partner
A dataset can be technically complete and still be difficult to use.
Consider a file containing hundreds of thousands of service records with columns labelled status_code, action_type and resolution_flag.
Without definitions, the recipient may not know whether a resolution flag indicates technical success, ticket closure, payment approval or customer acknowledgement.
An AI developer could build a model using an incorrect interpretation of the labels, producing apparently credible results that do not reflect the company's actual operations.
Documentation reduces this risk.
Gebru and colleagues' Datasheets for Datasets proposes a structured approach to describing a dataset's motivation, composition, collection process, recommended uses and limitations 4.
For an enterprise data partnership, this principle can be implemented through a data dictionary, provenance record, transformation history and a broader dataset description.
Exhibit 2: Minimum delivery-documentation package
| Document | Purpose | Example contents |
|---|---|---|
| Dataset specification | Defines intended use and delivery scope | Included systems, periods, data categories, permitted purposes |
| Data dictionary | Explains fields and values | Field types, definitions, valid codes, null meanings |
| Provenance register | Tracks origin and transformations | Source systems, extraction dates, processing versions |
| Quality report | Identifies known limitations | Missingness, duplicates, rejected records, coverage gaps |
| Rights and restrictions schedule | Specifies permitted use | Sensitive fields, contractual exclusions, recipient limitations |
| Delivery manifest | Identifies the actual package | Files, versions, record counts, integrity checks |
| Change log | Explains changes over time | Schema revisions, corrections and exclusions |
A recipient should be able to identify which dataset version was delivered, how it was prepared and what limitations apply.
For recurring arrangements, versioning becomes especially important.
A company might modify its operating system, introduce a new ticket category or replace a supplier platform. If the change is not documented, the buyer could interpret shifts in data format or category frequency as changes in the underlying business.
Documentation should also explain what the dataset does not contain.
An archive may exclude certain regions, periods, customer categories or products. Historical outcome fields may be unavailable before a particular year. Certain records may have been removed because their rights could not be established.
These exclusions are part of the dataset's definition, not minor technical footnotes.
The ability to communicate them clearly is commercially useful because it allows a prospective partner to decide whether the records are suitable for its application.
4. Secure transfer and validation: why delivery is more than moving bytes
After extraction, preparation and authorisation, the dataset must be made available through an agreed technical mechanism.
Possible options include encrypted object storage, managed file transfer, controlled data rooms, approved APIs or direct cloud-to-cloud delivery.
No single method is universally preferable.
A small evaluation dataset containing limited information may be suitable for a controlled data room. A large archive may be more efficiently transferred through managed cloud storage. A recurring arrangement may require a monitored API or scheduled data pipeline.
The choice should reflect the classification of the information, the recipient's capabilities, the volume involved and the contractual arrangements.
NIST SP 800-47 Revision 1 provides guidance on protecting information before, during and after exchange, including managing responsibilities and risk across organisational boundaries 5.
Exhibit 3: Comparing common delivery mechanisms
| Mechanism | Potential advantage | Principal limitation | Typical application |
|---|---|---|---|
| Controlled data room | Restricts access and supports review | May be unsuitable for large-scale processing | Preliminary diligence |
| Managed encrypted file transfer | Established operational transfer mechanism | Requires credential and transfer management | Periodic file delivery |
| Cloud object storage | Scalable delivery with granular access controls | Misconfiguration can create exposure | Large historical archives |
| API-based access | Supports selective retrieval and ongoing updates | Greater engineering and reliability requirements | Recurring data services |
| Controlled computation environment | Can restrict raw-data movement | Additional infrastructure and governance complexity | Sensitive evaluation or analysis |
These are illustrative choices, not standard requirements imposed by AI data buyers.
Security should be designed around the actual risk.
Controls may include encryption in transit and at rest, restricted identities, least-privilege access, temporary credentials, appropriate audit logs and agreed retention rules.
A transfer should not require the recipient to obtain unrestricted access to the provider's production systems.
Validating technical integrity
Successful transfer requires evidence that the intended information arrived without corruption or omission.
Technical validation may compare file counts, record counts, file sizes and cryptographic checksums.
Cloud-storage platforms such as Amazon S3 support integrity verification using several checksum algorithms, including SHA-256 6.
However, a matching checksum proves something narrower than business correctness. It can help establish that particular file contents match the expected contents. It does not prove that the original extraction selected the correct records or that their meanings were preserved.
The receiving partner should therefore perform both technical and semantic checks.
Technical checks may establish that files are intact, expected columns are present and declared data types are interpretable.
Semantic checks examine whether identifiers link correctly, outcome definitions make sense, date ranges are consistent and the dataset genuinely supports the agreed purpose.
The distinction matters because a technically perfect transfer of an incorrectly assembled dataset is still a failed delivery.
An acceptance protocol should separate provider pre-release checks, recipient technical receipt and joint semantic acceptance. Before release, the provider signs off the extract's source snapshot, permission scope, manifest, schema version, exclusions and count reconciliation. The recipient then verifies authorised access, file integrity, schema loading and coverage. Finally, named domain reviewers compare sampled case histories, time ordering and outcome meanings against the task specification. A discrepancy should be assigned an owner and a disposition: correct and redeliver, accept with a documented limitation, or reject the affected subset. The contract should specify response windows and who may authorise acceptance; a matching checksum alone is never sign-off.
A worked delivery example
Assume a hypothetical equipment-maintenance business has 120,000 historical service cases across several systems.
An initial quality review excludes 10% as duplicates or invalid records, leaving 108,000. A further 20% lack sufficient diagnostic context, leaving 86,400. Rights and sensitivity screening then excludes 25% of that remaining population, producing 64,800 preliminary eligible cases.
These are sequential, illustrative filters, not empirical industry rates. The remaining cases still require task-specific validation.
For a restricted pilot, the company selects a stratified sample of 3,000 cases, covering relevant equipment families, years, failure categories and escalation levels.
The delivery manifest records:
- Dataset and transformation version.
- Approved scope and extract period.
- Expected number of cases.
- Schema and permitted field definitions.
- Individual file checksums.
- Exclusion and quality notes.
- Authorised recipient and retention period.
Suppose the buyer confirms receipt of all files and their checksums match. That establishes one important delivery control.
A separate validation then identifies that 120 cases refer to equipment versions absent from the accompanying product taxonomy.
The 120 exceptions represent 4% of the 3,000 sampled cases. That proportion must not be applied to all 64,800 preliminary cases without justification from the stratification and sampling weights. The company must determine whether the exceptions are legitimate historical products, mapping errors or incorrectly included records.
Until that issue is understood, delivery should not be regarded as fully accepted for the relevant use.
The parties could correct the mapping, exclude affected records with documentation or revise the agreed use if the limitation is material.
This is the purpose of staged validation: to identify consequential discrepancies before the dataset enters a production AI workflow.
5. From one-time exports to recurring data pipelines
A one-time delivery and an ongoing data partnership are different operating models.
An archive can often be prepared through a bounded project. Once the records have been extracted and validated, the company's immediate technical obligations may be limited.
A recurring arrangement requires a repeatable process that remains reliable as the underlying business changes.
New records may need to be selected, transformed, checked, packaged and transferred every week or month. Corrections to historical cases may also need to be communicated.
A recurring feed therefore introduces responsibilities beyond the original extraction.
The company must define how new records are identified, how updates are distinguished from duplicates and what happens when processing fails.
For example, a monthly delivery that simply exports all records created after the previous extraction date may miss historical records that were corrected later.
An incremental pipeline may need to track both newly created records and modifications to earlier records.
Pipeline reliability
A dependable process should record the period covered by each transfer, the extraction version, the number of processed records and any rejected entries.
If a delivery fails midway, the system should be able to resume or replay safely without introducing duplicate records or silently omitting part of the period.
This may require stable record identifiers, version numbers and reconciliation controls.
Schema changes also need agreed treatment. A newly introduced field may be harmless, but a changed outcome definition could materially alter the meaning of the delivered data.
Rather than allowing such changes to pass unnoticed, the provider and recipient should agree when notification, approval or renewed testing is required.
Corrections and deletion
Ongoing governance extends to corrections, withdrawal of permissions and the expiry of retention periods.
If a record is corrected in the provider's system, the recipient may need an agreed mechanism for identifying and applying that change.
Where a deletion request or contractual withdrawal obligation applies, the parties should understand which source copies, derived datasets, backups and permitted processing environments are affected.
Deleting a transferred file does not necessarily reverse its influence if it has already been used to train a model. Obligations concerning trained models or derivatives therefore require careful contractual and technical assessment rather than assumptions of easy reversibility.
The ICO's data-sharing guidance recommends documenting responsibilities across the sharing lifecycle, including retention, deletion, security and regular review 7.
These responsibilities are particularly significant in recurring arrangements because the parties are maintaining an operational relationship rather than completing a single technical exchange.
Commercial implications
A proposed refresh service must be priced as a continuing delivery obligation, not merely another copy of a pilot extract. Beyond initial transformation it requires schema monitoring, exception handling, approved rights screening, support and incident response. The provider should assign a recurring cost owner, specify a feasible refresh cadence and identify which changes trigger renegotiation or suspension. If recurring obligations exceed credible fees or staff capacity, a bounded archive or less frequent service may be preferable. Chapter 7 develops the broader archive-versus-feed economics; the delivery question here is whether the contracted service can be operated reliably.
Exhibit 4: Illustrative cost of moving from a pilot to a recurring feed
| Activity | One-time pilot | Recurring production feed |
|---|---|---|
| Source mapping | Initial integration | Maintenance after system changes |
| Data extraction | Controlled sample | Scheduled or event-driven extraction |
| Transformation | Defined pilot processing | Repeatable automated rules and exceptions |
| Quality control | Sample validation | Continuous or periodic reconciliation |
| Rights screening | Approved pilot scope | Controls for newly collected records and changed permissions |
| Transfer | One authorised delivery | Repeated, monitored deliveries |
| Governance | Defined pilot termination | Ongoing retention, correction, security and audit obligations |
| Commercial exposure | Primarily project-specific | Continuing operating and contractual commitments |
The table demonstrates why the cost structure of a recurring licence cannot be inferred from the cost of preparing its first sample.
A transaction that appears attractive on a one-off basis may require a higher recurring fee or narrower service scope to remain commercially viable.
Equally, automation can reduce repeated effort once a stable and well-governed pipeline has been established.
6. Management oversight: deciding what should be built and when
The executive investment question is whether a bounded, receivable data product can be delivered on agreed terms. Management should authorise an initial version with explicit cost and time ceilings, named signatories for information release and recipient acceptance, and a funded plan for handling rejected records. A successful one-off pilot does not automatically justify building a permanent pipeline; it must first reveal realistic rates of reconciliation, exceptions, security effort and recurring support.
Assign one accountable project owner, supported by explicit decision rights: the operational owner approves meaning and sampling, legal/privacy owners clear the scope, engineering releases the reproducible extract, security authorises the access channel, and the recipient accepts or rejects the documented package. No role should silently approve another role's conditions.
The executive review should focus on evidence rather than assurances.
Can the relevant records be identified and extracted without unnecessary exposure? Does the transformed dataset preserve the information needed for the intended task? Are definitions, provenance and exclusions documented? Can the recipient verify both technical integrity and business meaning? Can the process be repeated safely if a recurring arrangement is proposed?
Management should also consider what happens when the process does not work.
A technically feasible extraction may produce too few usable records. Privacy controls may remove context essential to the task. Historical mappings may prove unreliable. The cost of continuous governance may exceed the commercial opportunity.
These are legitimate reasons to revise or abandon a proposal.
A well-structured project should make such findings early, before the company commits to an expensive pipeline or broad licensing obligations.
Practical implications
A disciplined enterprise-to-AI delivery process follows a sequence of progressively stronger evidence:
Define the task → identify sources → establish permissions → extract and transform → document → transfer securely → validate → monitor or close.
Each stage should have an accountable owner, a defined output, an acceptance record and a decision about whether proceeding remains justified. Pilot closure also requires confirmation of recipient retention/deletion arrangements and disposition of unresolved exceptions.
The company should begin with the smallest controlled package capable of testing its commercial hypothesis. Only if that package demonstrates technical usefulness, appropriate permissions and acceptable economics should the parties consider scaling the arrangement.
For recurring deals, the delivery mechanism should be evaluated as an operating service, with costs and risks that continue after the first transfer.
The real product is not merely a collection of files. It is a defined and governed information asset that a recipient can understand, use, verify and, where contracted, receive consistently over time.
Sources and further reading
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework 1.0, Core. Voluntary framework covering governance, mapping, measurement and management of AI risks. https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
- UK Information Commissioner's Office. Pseudonymisation, anonymisation guidance. Explains the distinction between pseudonymisation and anonymisation and associated identification risks. Guidance is under review following the Data (Use and Access) Act. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
- UK Information Commissioner's Office (updated January 2026). A Guide to International Transfers. Explains the circumstances in which granting access to personal information may constitute a restricted transfer. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/international-transfers/a-guide-to-international-transfers/are-we-making-a-restricted-transfer/
- Gebru, T. et al. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92. Framework for documenting dataset provenance, composition, collection, intended uses and limitations. https://doi.org/10.1145/3458723
- Dempsey, K., Pillitteri, V. and Regenscheid, A. (2021). Managing the Security of Information Exchanges, NIST SP 800-47 Revision 1. https://doi.org/10.6028/NIST.SP.800-47r1
- Amazon Web Services. Checking Object Integrity in Amazon S3. Technical documentation concerning checksum-based object-integrity verification. https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html
- UK Information Commissioner's Office. Data Sharing Agreements, Data Sharing Code of Practice. Describes relevant responsibilities concerning scope, security, retention, deletion and governance. Guidance is under review following legislative changes. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/data-sharing-agreements/