Data Monetisation 101 / Section 1 / Chapter 4

Section 01 · Understand the asset · Chapter 04

Five Years or Fifteen? How Data History Can Change AI Value

A longer archive can capture rare events, changes in operating conditions and outcomes that recent records cannot yet reveal. Its value to an AI buyer nevertheless depends on relevance, comparability, lawful use and the cost of making the history intelligible.

13 min read9 referencesSite last updated 9 October 2026

Executive perspective

Executive perspective

History adds value when it adds information. Longer archives can contain rare events, different market conditions and later outcomes that a short record cannot yet show. Another year of routine duplication may add much less.

Comparability is critical. System migrations, revised policies and changed outcome definitions can make older records misleading unless historical regimes are properly documented.

Mature outcomes can be distinctive. Long-horizon credit, customer and maintenance outcomes may only be observable in older cohorts, provided information available at the original decision point is preserved.

The additional years must earn their place. Rights, retention limits and the cost of recovering and cleaning old records should be assessed by period. The relevant comparison is the incremental usefulness of older cohorts for a defined buyer task, not a generic price per year.

1. What extra history can actually contribute

Time is valuable when it changes the information content of an archive, not when it merely increases its row count. A service business might resolve hundreds of ordinary tickets each day but experience only a handful of equipment failures that require multi-stage diagnosis. An insurance administrator could possess numerous straightforward claims and relatively few disputes with extensive investigation trails. In such cases, a longer period offers more opportunities to observe exceptions and connect them with effective or ineffective responses.

This distinction matters because AI buyers purchase suitability for particular tasks, not abstract age. A model intended to summarise current customer queries may benefit most from recent language, product versions and standard workflows. A system designed to support incident investigation may instead need unusual cases, failed first attempts, escalation paths and verified resolutions. The older material is useful only if it records the problem as it appeared to staff at the time, the choices available and the eventual result. A final ticket marked “resolved” rarely provides that sequence.

An important distinction is between event frequency and case quality. Longer observation can improve the chance of capturing rare events, but one complicated incident may generate dozens of duplicate emails and log lines. These do not constitute dozens of independent examples. Nor will an archive containing more historical crises necessarily represent the circumstances of a future crisis. Independent, well-labelled cases with distinct causal patterns are the meaningful unit of analysis.

Exhibit 1. The probability of observing at least one rare event

Qualifying opportunitiesAssumed independent event probabilityProbability of at least one event
1,0000.1%63.2%
3,0000.1%95.0%
5,0000.1%99.3%

Illustrative mathematics, not an empirical incidence estimate. Calculation: 1 − (1 − 0.001)^n, assuming a constant event probability and independent opportunities.

The exhibit explains a statistical mechanism through which additional observation might create coverage. It says nothing about how many different rare-event types appear, how useful the recorded responses are, or what a buyer would pay. Real cases may cluster because a single operational breakdown affects many records; the independence assumption then fails. Before claiming rare-event value, an organisation should deduplicate incidents and identify distinct failure modes, contributing conditions, actions and outcomes.

2. Fifteen years may contain several different businesses

A company can retain the same trading name for fifteen years while changing almost everything about how its records are generated. Customer identifiers may move between systems; eligibility criteria may change; a helpdesk may introduce self-service; a lender may revise risk policy; a maintenance team may replace equipment and diagnostic software. Apparent trends can therefore arise from a real change in customer behaviour, a change in recording rules, or both.

For AI, this creates an opportunity and a hazard. Properly labelled historical transitions can show how professionals adapt to unfamiliar operating conditions. They can also enable an evaluator to test performance separately in stable periods, market stress or major technology migrations. Conversely, mixing incompatible eras without disclosure can teach a system obsolete relationships or make performance measurements misleading. The technical term dataset shift describes changes between the data encountered during model development and those encountered later.

This is not merely theoretical. Guo and colleagues studied clinical prediction models across successive historical periods and reported a maximum absolute AUROC decline of 0.090 for sepsis prediction in their baseline temporal-shift comparison 2. AUROC is a measure of how well a model discriminates between cases with different outcomes. The study does not establish a rule about enterprise datasets or their commercial prices. It demonstrates that the passage of time can change predictive relationships enough to damage a model's performance. Research on machine-learning technical debt similarly identifies changing external conditions and hidden data dependencies as maintenance risks 3.

Historical data becomes more interpretable when it carries regime information: effective dates of policies, product and software versions, changes in outcome definitions, field dictionaries, and records of major migrations. Consider a drop in complaint rates after a CRM replacement. Without a map of discontinued categories and altered collection methods, neither a manager nor a model can confidently attribute the decline to better service. A timestamp establishes when something was recorded; regime metadata helps establish what the record meant.

Dataset documentation is therefore an economic input, not an administrative afterthought. The Datasheets for Datasets research programme recommends systematic descriptions of collection, composition, processing, intended use and maintenance 4. For a long operational archive, such documentation can make the difference between several expensive, disconnected extracts and a set of coherent historical cohorts that a buyer can evaluate.

3. The delayed-outcome advantage

Some of the most consequential enterprise outcomes take longer to observe than the underlying decision. A loan may not default or repay for years; a customer intervention may only be judged after another renewal cycle; an equipment repair can fail months after the technician closes the job; a code change can appear successful until later production incidents. Older records can contain mature labels that newer cohorts, however clean, cannot yet supply.

The valuable structure is often a traceable sequence: information available at decision time → decision or intervention → subsequent events → defined outcome. That sequence can support prediction, evaluation or analysis of professional judgement. But the archive needs reliable event identifiers, dated stages and clear outcome rules. A label assigned retrospectively can also introduce information leakage if it includes facts unavailable when the original decision was taken.

For example, a credit decision-assistance model should not be evaluated using a restructuring flag added two years after origination as though the underwriter knew it at application. Similarly, a service-resolution model should not receive the final repair code when asked to choose a first diagnostic step. Historical richness helps only when the original information boundary can be reconstructed.

Newer cohorts create another difficulty. If a company measures retention eighteen months after an intervention, a customer who entered the programme three months ago has not had enough observation time. Treating that customer's missing eighteen-month outcome as failure or success distorts the comparison. This is a form of right-censoring, long recognised in time-to-event analysis 5. The appropriate treatment depends on the analytical task and follow-up process; simply deleting all incomplete cases may introduce its own bias.

A manager assessing a fifteen-year credit archive should therefore ask whether older loans have complete repayment and recovery outcomes, which definitions were used in each period, and whether customer identifiers survive system changes. Years of nominal coverage cannot substitute for mature, attributable and legally usable outcomes. In some situations, five documented years with reliable resolutions may outweigh fifteen years of provisional statuses. In others, completed long-term outcomes are precisely what a shorter archive cannot replicate.

4. Diminishing returns, usable records and preparation economics

Extra history often provides diminishing average benefit once an archive already covers routine workflows well. Learning-curve research has explored how model error can decline as training datasets expand 6. Yet improvement should not be presumed smooth or guaranteed: a 2025 NeurIPS study found statistically significant non-monotonic or otherwise ill-behaved patterns in around 15% of the learning curves it examined 7. Neither study establishes a universal relationship between historical years and a dataset's licensing price. For operational data, the marginal contribution depends on the tasks and the examples added.

A first decade might cover thousands of repeated standard cases, while an additional year contains a product recall or major market disruption that creates distinctive evidence. Conversely, another five years of obsolete, poorly mapped records might cost more to prepare than any measurable benefit they offer. The economics are best tested at the level of usable, task-relevant cohorts rather than calendar years.

Exhibit 2. Five years versus fifteen: gross records can mislead

Assumed measureArchive A: five recent yearsArchive B: fifteen years
Case records per year150,000150,000
Gross records750,0002,250,000
Records passing structure and provenance checks95%50%
Usable records before outcome filter712,5001,125,000
Among usable records, outcomes mature and linkable80%30%
Cases passing both tests570,000337,500
Historical conditions coveredFewer regimesPotentially more regimes

Entirely hypothetical companies and assumptions. Calculation: gross records × structural pass rate × mature-and-linkable-outcome rate. The two percentage filters are assumed conditional in sequence, not measured from any actual seller.

Although Archive B has three times the gross records, it contains fewer fully usable cases under these assumptions. That is a data-readiness comparison, not a conclusion about value. A buyer of rare recession-era credit cases might still prefer Archive B; a buyer adapting software to today's support workflow might prefer A. Rights and security status remain separate eligibility tests for both.

The older archive's apparent disadvantage is also sensitive to preparation quality. If Archive B's structural pass rate rises from 50% to 70%, and the mature-and-linkable-outcome rate rises from 30% to 50%, the qualifying count becomes 787,500, above Archive A's 570,000. That does not mean a clean-up project can necessarily achieve either rate. It shows why cataloguing, linking and migration work can affect usefulness more than purchasing extra storage. Costs of achieving the improved rates must be compared with the benefit of the additional cases.

A commercial assessment should separate incremental technical usefulness from realised contract value. A hypothetical decision rule is: expected buyer benefit attributable to the older cohort, less data preparation, legal review, integration, privacy controls and ongoing support costs. Even a positive result is not a licensing price. Exclusivity, permitted uses, competition, liability, duration and negotiations also matter. No reliable market-wide “price per extra year” follows from this framework.

5. Rights, relevance and historical liabilities

A company may control the database that stores an operational record without owning unrestricted rights to disclose, train on or license everything inside it. Historical customer contracts might have different confidentiality provisions. Third-party intellectual property may appear in attachments. Personal information retained for one business purpose is not automatically available for another purpose merely because it has been stored for a long time.

For UK personal data, the Information Commissioner's Office explains that AI development remains subject to data minimisation and storage-limitation requirements: organisations should justify the personal data they need, its purposes and how long it is retained 8. Some ICO AI guidance is under review following legislative changes, so current advice and case-specific legal analysis remain essential. Pseudonymisation can reduce risks but is not the same as anonymisation, and neither operational ownership nor a historical retention practice by itself establishes licensing permission.

Older datasets can be particularly expensive to prepare. Legacy extracts may have inconsistent encodings, undocumented categories, damaged event links or inaccessible attachments. A business may need to reconcile acquired subsidiaries, identify former customers, suppress restricted material, classify commercially sensitive information and document what was deleted or transformed. Such costs should be estimated by tranche, not hidden inside an optimistic claim about data volume.

An effective first-pass classification distinguishes four states: (1) retained records; (2) technically retrievable records; (3) fit-for-purpose records with defensible provenance and outcomes; and (4) records legally and operationally suitable for the proposed use. Each state is narrower than the last. A company should avoid showing buyers raw historical samples until confidentiality, privacy, access and permitted-use controls are established. A limited metadata inventory and governed pilot can often answer feasibility questions more safely.

6. Demonstrate the value of older periods to a particular buyer

Different buyer tasks imply different optimal windows. Current-workflow assistance prioritises contemporary terminology, policies and interactions. Long-horizon prediction requires cohorts old enough to observe relevant outcomes and stable feature definitions. Evaluation and stress-testing may prize historically difficult, independently adjudicated cases even when they are unrepresentative of ordinary operations. A historical tranche can be valuable for one purpose and unsuitable for another.

Exhibit 3. Which historical periods suit which AI tasks?

Potential AI useHistorically useful evidencePrincipal diligence concern
Current process assistanceRecent, detailed action sequences and current definitionsObsolete procedures may teach incorrect actions
Long-horizon outcome predictionDated decisions linked to mature, verified resultsCensoring, changing policies and label leakage
Evaluation and stress-testingDistinct edge cases, regime changes, independent outcomesSelection bias and unrepeatable definitions

Interpretation: This task-by-period map is an analytical framework, not evidence of a buyer's acquisition policy. It distinguishes the marginal usefulness of history from the general dataset-quality tests in Chapter 3.

To test older history, a buyer and seller could agree a bounded question such as whether ten additional years improve detection of a defined equipment failure category. The comparison would hold the task and evaluation period constant, assess a recent-only dataset against a combined dataset, and measure task-relevant performance. Temporally ordered evaluation matters: ordinary random splits can give misleading results when future records contaminate training for earlier periods. Scikit-learn's technical guidance describes why time-series evaluation requires special care 9. The test should also examine performance in individual regimes, rather than mask a weak recent-period result with a strong historical average.

Such a pilot may show that extra years add little, expose a meaningful capability improvement, or identify only a small older slice worth curating. All three outcomes are commercially useful. The purpose is to avoid treating the largest extract as the best product, and to establish a credible answer to the buyer's core question: what could this additional history enable that a shorter archive could not?

Practical implications for operating companies

Start with a year-by-year inventory, not a headline count. Record expected and actual volumes, missing periods, changes in systems or policy, availability of linked outcomes, exceptional-event types and known contractual restrictions. Identify where definitions can be normalised and where periods should remain distinct. This is the evidence needed to decide whether deeper diligence is justified.

Next, select one or two buyer tasks for which historical depth plausibly matters. Prepare a small, documented sample or aggregate profile under appropriate controls, then examine whether additional years improve edge-case coverage, outcome maturity or robustness. Assign preparation and governance costs to each candidate tranche. Do not equate a technically interesting dataset with a permissioned, saleable asset, and do not make pricing claims before a credible use case and rights position have been established.

The executive conclusion is deliberately conditional: fifteen years can be more valuable than five, but only because of what the extra decade contains and what a buyer can responsibly do with it. The most valuable historical asset may be neither the complete archive nor the latest slice, but a well-documented selection of cases whose age brings genuinely new evidence.

Sources and further reading

  1. Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, NIST AI 100-1. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10
  2. Guo, L. L., Pfohl, S. R., Fries, J., et al. (2022). “Evaluation of domain generalization and adaptation on improving model robustness to temporal dataset shift in clinical medicine.” Scientific Reports, 12, 2726. https://www.nature.com/articles/s41598-022-06484-1
  3. Sculley, D., Holt, G., Golovin, D., et al. (2015). “Hidden Technical Debt in Machine Learning Systems.” Advances in Neural Information Processing Systems, pp. 2503–2511. https://research.google/pubs/hidden-technical-debt-in-machine-learning-systems/
  4. Gebru, T., Morgenstern, J., Vecchione, B., et al. (2021). “Datasheets for Datasets.” Communications of the ACM, 64(12), 86–92. https://www.microsoft.com/en-us/research/?p=563964
  5. Kaplan, E. L., & Meier, P. (1958). “Nonparametric Estimation from Incomplete Observations.” Journal of the American Statistical Association, 53(282), 457–481. https://doi.org/10.1080/01621459.1958.10501452
  6. Frey, L. J., & Fisher, D. H. (1999). “Modeling decision tree performance with the power law.” Proceedings of the Seventh International Workshop on Artificial Intelligence and Statistics, PMLR R2. https://proceedings.mlr.press/r2/frey99a.html
  7. Yan, C., Mohr, F., & Viering, T. (2025). “LCDB 1.1: A Database Illustrating Learning Curves Are More Ill-Behaved Than Previously Thought.” NeurIPS 2025 Datasets and Benchmarks Track. https://proceedings.neurips.cc/paper_files/paper/2025/hash/1543d6d5cb976e4f9fbfaedf2e257967-Abstract-Datasets_and_Benchmarks_Track.html
  8. Information Commissioner's Office (current guidance, accessed October 2026). Principle (c): Data minimisation, and How should we assess security and data minimisation in AI? Some guidance is under review. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-protection-principles/a-guide-to-the-data-protection-principles/data-minimisation/ ; https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-should-we-assess-security-and-data-minimisation-in-ai
  9. scikit-learn project (technical documentation). TimeSeriesSplit and Cross-validation: evaluating estimator performance. https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html