Executive perspective
Executive perspective
Format describes engineering effort, not intrinsic economic value. Structured rows offer explicit fields, while narrative documents and media may contain the expert reasoning that makes a task learnable. Semi-structured records occupy the middle.
Relationships are often the scarce ingredient. A case joined reliably to its messages, decisions and outcome is more informative than either isolated source, but incorrect joins can contaminate both training and evaluation.
Extraction quality needs evidence. Optical recognition, transcription and document parsers introduce different error patterns; a clean-looking output is not proof of fidelity.
Preparation can defeat the business case. The relevant output is the number of eligible, interpretable cases after transformation, with the full cost of validation and rights controls included.
1. Three formats, three different kinds of friction
Exhibit 1. Structured, semi-structured and unstructured sources
| Format | What it looks like | Potential advantage | Typical difficulty |
|---|---|---|---|
| Structured | Rows and columns with defined fields, such as transactions, case statuses or CRM records. | Easier to filter, aggregate, validate and link when schemas and identifiers are stable. | May omit explanations, tacit expertise or the context behind a decision. |
| Unstructured | Documents, emails, conversations, audio, images, video or free-text notes. | Can capture explanations, nuance, expert judgement and real-world language. | May require OCR, transcription, classification, redaction, segmentation or human review. |
| Semi-structured | Logs, JSON, messages with metadata and exported application objects. | Often combines a predictable outer format with rich or variable content. | Fields and event formats can change; flexible contents may still need interpretation. |
The categories can overlap. A PDF invoice has a file format, document layout, text and potentially a table of line items. An application log might have fixed fields such as event_type and created_at, while its payload changes by event. A support ticket can have structured status fields plus free-text messages and attachments.
The data’s format is therefore not a quality score. A structured table can be duplicated, inaccurate or poorly documented. Unstructured content can be highly informative but still be unusable if its source, permissions or meaning cannot be established.
The distinction matters because different AI tasks consume data differently. A prediction system that estimates expected repair time may need stable categorical inputs and numeric outcomes. A support assistant may need natural-language descriptions and an authoritative procedure library. An agent evaluator may need ordered tool events with state changes, retries and verified end conditions. Choosing a 'preferred format' before choosing the task confuses the storage technology with the capability being bought.
Semi-structured formats merit particular care. A JSON field may be machine-readable yet lack a published vocabulary; a field called result may contain a string, nested object or absence depending on software version. Likewise, a PDF has visible headings and spatial layout, but converting it to linear text may lose which amount belongs to which invoice line. Transformation decisions should therefore be recorded as data lineage. W3C’s provenance model provides a general language for entities, activities and agents involved in producing artefacts; it does not validate their accuracy by itself.
2. The commercial value often appears at the join
Consider a support case. Its ticketing system may record category, timestamps, owner, status and resolution code. The conversation may show what the customer reported, which diagnostic questions were asked and how the issue was explained. Joined together, the records can show both what happened and some evidence of why.
A simplified linked example:
Exhibit 2. Illustrative linked support-case record
case_id | created_at | issue_category | message_id | message_text | resolution_code | verified_outcome |
|---|---|---|---|---|---|---|
| C-4812 | 2025-04-08 09:14 UTC | Payment failure | M-01 | “The card was charged, but the booking still shows unpaid.” | PAYMENT_SYNC | Yes |
| C-4812 | 2025-04-08 09:27 UTC | Payment failure | M-02 | “We found a delayed confirmation event and reconciled the booking.” | PAYMENT_SYNC | Yes |
The case_id links the narrative to the case; message_id and timestamps preserve sequence. But the join is only trustworthy if the provider can explain what the fields mean and how they were captured.
What can go wrong at the join?
- Unstable identifiers: one system uses a case number, another uses a customer ID that has since been reused.
- Different clocks: timestamps are stored in local time in one source and UTC in another, or represent different events.
- Many-to-many matches: a single customer conversation relates to several orders, but the link rule is undocumented.
- Missing context: the note refers to “the usual fix,” but the relevant procedure is not included.
- Retrospective edits: an outcome field is updated after the event, making it look as if it was known at decision time.
- Changed definitions: a resolution code was renamed or reinterpreted, but historical values were not remapped.
Record linkage and entity resolution are established data-integration problems: combining sources accurately requires explicit matching rules and quality checks, not just similar-looking IDs. Survey of entity resolution methods.
What good documentation includes
For each joined source, a buyer may need:
- field definitions, allowed values and units;
- identifier scope and whether IDs are stable or reusable;
- timestamp meaning, timezone and event ordering rules;
- source system and extraction date;
- join keys, match logic and known false-match rates;
- schema versions and documented changes;
- whether fields are original, inferred, corrected or manually entered.
NIST’s research-data framework describes provenance and metadata as important to assessing reliability and enabling reuse; without adequate documentation, even an important dataset can become difficult to interpret. NIST Research Data Framework.
Verify both matches and non-matches
A joined sample needs a measurable error rate, not merely a successful SQL join. Suppose 10,000 messages have been attached to work orders using a mixture of IDs and inferred customer/time windows. If an audit of representative strata suggests 2% are incorrectly linked, approximately 200 messages could carry the wrong outcome label. That illustration is not an observed rate. If false links cluster around escalations or rare failures, the damage is not proportional to 2% of cases: those cases may be exactly what the evaluator is supposed to assess.
Audit the full matching logic, including records that failed to match. A 98% match rate can conceal incorrect many-to-many matches, fabricated confidence or the omission of difficult cases. Review known matched pairs, deliberately mismatched pairs, record IDs reused after system migrations, time-window overlap and disagreements between systems. For any probabilistic linkage, retain match scores or rule identifiers, sample cases close to the decision threshold and describe the cost of false positive versus false negative links. An apparent increase in record count should never substitute for evidence that the relationships are correct.
3. Extractability is part of the asset economics
A company may hold useful information across databases, APIs, document stores, scanned files, email archives, audio platforms and operational applications. The practical question is not only “What data do we have?” but also “What work is required to make a safe, interpretable sample?”
Exhibit 3. Preparation burden across source formats
| Source | Common preparation tasks | Questions to investigate |
|---|---|---|
| Database or warehouse | Schema mapping, duplicate checks, null analysis, event filtering. | Are historical schemas preserved? Are status values and units consistent? |
| Documents and PDFs | OCR, layout parsing, document classification, table extraction and quality review. | Are scans legible? Do templates change? Can extracted fields be checked against source pages? |
| Email or support conversations | Thread reconstruction, speaker and timestamp handling, redaction, case linking. | Are quoted messages duplicated? Are attachments essential? Can sensitive content be excluded? |
| Audio or calls | Transcription, speaker separation, timestamp alignment and review. | Is audio quality adequate? Are accents, specialist terms or overlapping speakers handled? |
| Images or video | Classification, object or scene annotation, frame selection and rights checks. | Are labels reliable? Are images linked to the correct event or asset? |
| Logs or JSON | Event normalisation, schema-version handling, sequence reconstruction. | Do payload fields vary by event type? Are retries and failed calls distinguishable? |
OCR can make scanned documents searchable, while document-processing systems may extract text, layout, tables or fields into structured representations. But extraction does not guarantee that the result is correct; it creates a transformed dataset whose accuracy should be sampled and documented. Google Cloud Document AI overview.
A practical cost model should account for the work that is actually needed:
Preparation effort = extraction + linkage + transformation + redaction + annotation + validation + documentation
Not every archive needs every step. A clean API feed with stable identifiers may require little transformation. A mixed archive of scanned reports, email threads and changing schemas may need substantial manual review.
Small worked preparation example
Assume a hypothetical company has 12,000 scanned maintenance reports. A first-pass OCR process extracts text and tables, but some reports have rotated pages, faint scans or handwritten notes.
A responsible pilot might:
- sample 200 reports across years and document templates;
- classify them by scan quality and form version;
- compare extracted asset IDs, fault codes and dates with the original pages;
- quantify error types and identify fields that need human validation;
- estimate processing and review effort before scaling.
The 200-record sample is an illustrative test design, not a universal statistical threshold. The right sample depends on document variability, error consequences and the precision required. The purpose is to discover whether extraction is reliable enough for the intended task and to estimate the cost of correcting it.
Convert extraction uncertainty into a commercial estimate
Return to the 12,000 hypothetical scanned reports. Assume a first-pass study finds that one-quarter need manual correction after automated extraction. If correction averages £3.50 per affected report and fixed extraction/configuration work costs £14,000, direct preparation is 12,000 × 25% × £3.50 + £14,000 = £24,500. Suppose only 6,500 of the reports are ultimately eligible and linked to trustworthy outcomes. The illustrated direct cost per qualified report becomes £24,500 ÷ 6,500 ≈ £3.77, before legal review, delivery and buyer testing.
This ratio depends on the task. A report lacking a linked outcome may still be useful for text search, but it is weaker evidence for outcome-based evaluation. If manual correction rises to 50%, the same cost model becomes £35,000, or about £5.38 per otherwise unchanged qualified case. The appropriate decision is to test whether automation or a narrower document subset lowers costs without selecting away difficult cases. No market licence price can be inferred from the calculation; it is a preparation-cost floor for a particular hypothetical asset.
4. Quality and drift: the archive may not be one consistent dataset
Operational systems evolve. A CRM field may be renamed, a ticket workflow redesigned, a JSON payload expanded or a document template updated. If changes are hidden, an apparent trend may simply reflect a new definition.
A useful version ledger might look like this:
Exhibit 4. Schema version and comparability ledger
| Period | Change | Potential effect | Recommended treatment |
|---|---|---|---|
| Jan–Jun 2023 | case_status = "resolved" meant agent closed the ticket. | Closure may not prove the customer’s problem was solved. | Keep as administrative status; do not treat as verified outcome. |
| Jul 2023 | Workflow added supervisor approval. | “Resolved” now follows a second review step. | Record workflow version and distinguish approval from customer success. |
| Mar 2024 | issue_category values consolidated. | Historical categories no longer map one-to-one to current values. | Preserve original labels and provide a versioned mapping. |
| Sep 2024 | API began omitting empty fields. | Missing field may mean “not supplied,” not “false.” | Document null semantics and avoid converting absence to a negative label. |
Quality review should examine completeness, consistency, uniqueness, validity, timeliness, provenance and outcome strength. The right response is not always to “clean” the archive into one uniform format. Sometimes the honest and more useful package preserves historical versions and clearly marks breaks. False uniformity can erase meaningful differences.
Schema evolution is common in data systems; downstream users can be affected when sources add, rename or drop fields. Microsoft documentation on schema evolution.
Represent changing source systems honestly
Schema evolution is not only a software compatibility issue; it changes how an analyst should interpret behaviour. If a support platform replaces three escalation codes with one broad label, treating every historic and current code as the same class may hide a genuine reduction in precision. An unrecorded system change can masquerade as concept drift. Conversely, aggressive harmonisation may remove the very differences the buyer needs to study across time. Preserve the original labels and provide a separate mapping table with effective dates, transformations and a list of ambiguous cases.
Quality claims should report denominators and exceptions. '95% complete' needs a definition of complete for each candidate use: it may mean the customer ID is present, or that the case has all decision-time evidence plus a verified outcome. An archive can achieve excellent field completeness but lack legal authority for external use. A buyer will often favour transparent exclusions and measured limitations over claims that a mixed source was uniformly cleaned. Documentation schemes such as Data Cards and Datasheets for Datasets describe the collection history, composition and intended uses needed to interpret those limitations.
5. Executive review: is the combined asset usable?
Assess the intended use before deciding whether the format mix is valuable.
Data-readiness sequence
Define the task — What should the model learn, test, retrieve or predict?
↓
Select essential evidence — Which fields, documents, events or modalities are necessary?
↓
Establish links and meaning — Can inputs, decisions, actions and outcomes be connected and interpreted over time?
↓
Test extraction and quality — What can be automated, what requires review, and how will errors be measured?
↓
Document boundaries — Which records are missing, restricted, uncertain, outdated or governed by different definitions?
↓
Prepare a safe sample — Can a representative sample demonstrate usefulness without exposing unnecessary information?
Questions for an executive review
- Which fields, documents or events are essential to the intended capability?
- Can records be joined across systems using stable identifiers and interpretable timestamps?
- Do records preserve inputs, decisions, actions and outcomes where those relationships matter?
- What extraction, transcription, redaction or annotation work is required?
- How will schema changes, missing values, duplicates and uncertain links be detected?
- Does the data represent current operations, or several incompatible historical processes?
- Can a safe sample demonstrate quality and provenance without including unnecessary sensitive information?
- Does the expected buyer use justify the cost and risk of preparation?
Executive takeaway
A structured database is not automatically superior to messy text, and unstructured material is not automatically richer or more valuable. The strongest asset may be the relationship between them: an event record linked to the notes that explain it, the action taken and an independently verified outcome.
The commercial test is whether that combined asset can be made legible, traceable, repeatable and authorised for a buyer’s specific use at a preparation cost that makes sense.
The operating company’s decision
An executive sponsor should commission one small cross-system sample before a full export. Select representative operational periods and difficult cases, define the join keys and any redaction rules, retain the source-to-output mapping inside controlled access, and ask independent reviewers whether each assembled case still means what the original systems recorded. The output is not simply a spreadsheet of joined records. It is a proof of extraction fidelity, meaning, permissions, cost and repeatability.
The project should stop or narrow if the most valuable workflow relationships cannot be established without unreasonable privacy exposure, or if a critical identifier is no longer available. Where reconstruction succeeds, package the version ledger, linkage method, failure analysis and cost model alongside the resulting cases. Those materials make the asset testable and reduce the buyer's own integration uncertainty; they are not a substitute for evidence that the dataset improves a specific capability.
Sources and further reading
- NIST, Research Data Framework, data lifecycle and documentation. https://nvlpubs.nist.gov/nistpubs/SpecialPublications/1500-18/NIST.SP.1500-18r2.html
- World Wide Web Consortium (W3C), PROV-O: The PROV Ontology (2013). https://www.w3.org/TR/prov-o/
- Gebru, T. et al. (2021), Datasheets for Datasets, Communications of the ACM. https://doi.org/10.1145/3458723
- Pushkarna, M., Zaldivar, A. and Kjartansson, O. (2022), Data Cards, ACM FAccT. https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/
- Google Cloud, Document AI overview, technical product documentation on document extraction. https://cloud.google.com/document-ai/docs/overview
- Microsoft Learn, Schema evolution in lakehouse tables, implementation guidance. https://learn.microsoft.com/en-us/fabric/data-engineering/delta-lake-schema-evolution
- Binette, O. and Steorts, R. C. (2022), ‘(Almost) All of Entity Resolution’, arXiv:2008.04443 (revised preprint, 17 January 2022), reviewing record linkage, deduplication and entity-resolution methods. https://arxiv.org/abs/2008.04443