Data Monetisation 101 / Section 1 / Chapter 2

Section 01 · Understand the asset · Chapter 02

The Data Sitting Inside a 50-Person Company Could Be More Interesting Than You Think

A small company's distinctive data is usually a record of specialised decisions, not a giant database. This chapter shows how to find those records, test their usefulness and compare internal value with the cost and risk of licensing.

14 min read10 referencesSite last updated 9 October 2026

Executive perspective

Executive perspective

A 50-person company can have a narrow but unusually informative history of work: requests, constraints, expert judgements, exceptions and outcomes that never appear in public material. Headcount and storage volume do not reveal whether that history is useful. The asset question is whether a decision can be reconstructed, its result evaluated and its use permitted. Management should begin with one repeated unit of work, audit a sample and compare an internal application with any external buyer task. A small, controlled evaluation may be more credible than an ambitious claim about “training data”. Preparation, customer confidence and the firm's own competitive advantage can make external licensing unattractive even when the underlying records are technically strong.

1. The valuable thing may be the work, not the volume

Imagine a 50-person distributor of specialist industrial components. Each day it receives requests containing an equipment model, operating conditions, delivery deadline and partial description of a fault. Staff choose among compatible parts, check stock and supplier lead times, seek an engineer's approval for exceptions and learn later whether the chosen configuration worked. Over years, this process can create a record of how expert judgement was applied under constraints. The firm's website may show product descriptions; it will seldom show the ambiguity, rejected alternatives and feedback behind a difficult quote.

The commercial hypothesis is specific. A buyer developing a configuration or service assistant might need examples of how specialists resolve ambiguous requests, or a held-out set to test whether its system knows when to escalate. The company itself might use the same evidence to help new staff retrieve precedents. Neither possibility follows from the company's small size or the existence of a CRM database. The record must capture the information that made a decision rational at the time and a meaningful later outcome.

This is consistent with the OECD's observation that the value of data depends on its context of use and the knowledge extracted from it.1 It is an analytical principle, not a claim that small-company datasets command a premium. A specialist archive may be distinctive but commercially too narrow, too costly to clean or too entangled with customer rights to share. Data created incidentally for operations must be evaluated on those terms.

There is also a distinction between rare and useful. A company might be the only local distributor of an obsolete component, yet its records may concern a shrinking installed base with no willing buyer. Conversely, an ordinary looking workflow may have value because the firm consistently recorded rejected options and downstream results. The buyer's relevant alternative might be hiring experts to annotate cases, building a simulator or using its own records. The seller should be able to say what its archive adds to those alternatives: a wider range of faults, credible outcomes, a longer time series or a clearer permission trail. That comparison is more informative than calling the data proprietary.

Exhibit 1 — An illustrative decision record

FieldExample from a component requestWhy it matters
ProblemCustomer reports intermittent pump shutdownCaptures the initial task, including uncertainty
ContextModel P7, 2022 controller, high ambient temperatureDefines which products and instructions apply
Human judgementEngineer rejects a cheaper valve due to pressure ratingRecords an expert constraint, not just the final item
EscalationSafety review required above stated thresholdShows where an assistant must defer
Verified outcomeSelected part installed; no repeat incident during stated follow-upSupplies a bounded result, with its observation window
ProvenanceTicket, quote version, engineer approval and service confirmation IDsLets an auditor trace the example to its sources

This is a constructed example. The follow-up is a useful signal but does not prove the part alone caused the result. An actual extract would need the company to verify each field, define the observation window and exclude identifiers or third-party material as required.

2. Find a unit of work before opening the data warehouse

Small companies often have several partial views of the same event: a CRM note, quotation, engineer email, order, support ticket and return. Starting with a table name can hide the relationship among them. Start instead with a repeated decision: request → assessment → approval → action → observed result. Ask which systems record each stage, who can explain it and where the chain commonly breaks. One clear workflow is a better first candidate than every historical file.

For the distributor, an inventory might reveal that simple catalogue orders are highly regular but add little information beyond published specifications. Complex quotes may be richer, yet many never become orders; an absent outcome may mean the customer bought elsewhere, not that the recommendation was poor. A return code may reflect shipping damage rather than technical incompatibility. Such distinctions determine which cases are suitable for an AI task and which would teach a false lesson.

Data work is a management issue as well as an engineering issue. Sambasivan and colleagues' interviews with AI practitioners describe how neglect of collection and documentation can create downstream “data cascades” in high-stakes systems.2 Their study does not measure the value of this hypothetical distributor's archive. It helps explain why small errors in labels and context can persist through model development and evaluation. A company should assign a domain owner who can challenge what the data appears to say.

3. Test quality, coverage and transferability

The first audit should be a stratified sample, not a full clean-up project. Sample across years, product families, routine and exceptional cases, customers of different sizes, and both successful and disputed outcomes. For each case, ask whether another competent specialist could reconstruct the decision without interviewing the original employee. Record missing fields, copied text, duplicate cases, conflicting timestamps and policy changes. Define the intended population: all support requests, only approved quotes, or a specific product family. A completeness percentage without that denominator is misleading.

Quality and representativeness are distinct. A perfectly documented set of unusual failures may be excellent for testing escalation yet poor for predicting the average request. A long archive may reflect one supplier's discontinued products. Historical staff may have followed shortcuts that current policy forbids. Suresh and Guttag identify sources of downstream harm across collection, development and deployment; selection and measurement choices can matter long after records leave the source system.3 NIST's AI Risk Management Framework likewise calls for documenting data selection, representativeness and testing in the intended context.4

Transfer to another organisation is a separate question. A buyer may have different equipment, customer vocabulary, labour rules or safety thresholds. The seller should avoid describing a successful internal prototype as evidence that the same dataset will improve every buyer's model. A controlled test can compare performance on the buyer's target cases, record failure modes and define where human approval remains mandatory. NIST's 2026 evaluation work reinforces the distinction between performance on a fixed set of cases and expected performance on similar future cases.5

A further danger is that decisions and outcomes are observed selectively. Staff may log exceptional escalations in detail while routine correct quotes leave only an order number. Returns may be more visible than successful installations. A naïve model trained on this archive could learn that every ambiguous request is high risk, or that all unreturned products worked. The audit must report these missing pathways explicitly. Where possible, compare the sample with an independent operational denominator, such as the total number of requests received in the same period, and avoid treating “no complaint” as a verified success label.

Exhibit 2 — A sample-audit sheet for management

CheckReport the numerator and denominatorDecision implication
Complete trailCases with input, decision and outcome / cases sampledDetermines whether supervised examples are possible
Current relevanceCases governed by current products and rules / cases sampledSeparates current guidance from historical evidence
Expert agreementCases where reviewers agree on an acceptable action / cases double-reviewedReveals ambiguous labels and escalation needs
Sensitive contentCases requiring exclusion or access controls / cases sampledDetermines feasible delivery method
Segment coverageCases by product, year, severity and customer typeTests representativeness for the named task

The audit should report the sampling frame and uncertainty, especially when a small sample is used to infer archive-wide proportions. The table is a discovery specification, not a licence-ready quality certificate.

4. Match the records to the right AI use

The same archive can support several technically different activities. Current product instructions with clear version control may be best delivered to an assistant through retrieval, so staff can see the source behind an answer. Historical input–decision pairs may help adapt a classifier or routing tool, but only if reviewers can agree what a good decision was. Difficult, unseen cases may be most useful as an evaluation pack, testing whether a proposed assistant recommends an unsafe component or fails to escalate. Ordered approval records may help design an agent's permitted workflow, though old exceptions should never silently become today's rules.

The buyer question must therefore be concrete: “Can these records reduce wrong part recommendations on supported P-series pumps without increasing unsafe approvals?” This can be tested against a baseline such as the existing search process or a general assistant. An evaluation should specify the eligible case population, outcome categories, scoring rubric and treatment of ambiguous cases before looking at model results. In 2026, NIST described sequestered, blind testing as a way to reduce train–test contamination in its AI Technology Evaluation programme.6 A small company cannot necessarily replicate that infrastructure, but it can reserve cases, restrict access and document who saw them.

Documentation supports every use. Gebru and colleagues' Datasheets for Datasets proposes a structured account of a dataset's motivation, composition, collection and limitations.7 A concise version for the distributor would identify source systems, product coverage, time span, exclusions, known label defects, review method, access restrictions and contact owner. These details let a buyer assess whether a pilot is meaningful without receiving raw customer records at the first meeting.

5. Compare internal value with external licensing

The firm's first buyer may be itself. Suppose, illustratively, that it handles 8,000 quote requests annually, 35% of which are complex enough to benefit from better precedent retrieval. That is 2,800 relevant requests. If an assistant saves eight minutes on each adopted case while preserving quality, full adoption releases about 373 staff hours a year (2,800 × 8 ÷ 60). At a hypothetical loaded labour cost of £50 per hour, the gross capacity value would be about £18,667. At 60% adoption, it falls to 224 hours or £11,200. If ongoing tool, review and support cost £6,000 a year, the corresponding annual net capacity values are approximately £12,667 and £5,200. None is cash saved automatically: it counts only if the freed time can be redeployed or a measurable cost avoided.

Exhibit 3 — Illustrative internal-use sensitivity

Adoption of relevant requestsAnnual hours releasedGross capacity at £50/hourAfter £6,000 annual operating cost
40%149£7,467£1,467
60%224£11,200£5,200
100%373£18,667£12,667

Figures are rounded from 2,800 eligible requests and eight minutes saved per adopted request. If initial preparation and deployment cost £30,000, even the 60% scenario implies a long simple payback of roughly 5.8 years (£30,000 ÷ £5,200), before discounting or changing demand. That is a reason to pilot the workflow and measure actual adoption, not a reason to assume the data has no value. Faster response, fewer returns or improved customer retention could matter, but would need evidence and careful attribution.

An external licence has a different economic equation: contract consideration minus preparation, legal, privacy, secure delivery, support and expected risk cost. No published price list can resolve it for this company. A prospective buyer must name the task, required rights, access method and performance test before the seller can cost the work. A one-year evaluation licence and a perpetual, exclusive training licence transfer very different options. The latter may undermine the firm's internal advantage or future ability to work with other counterparties. Contract terms should address purpose, recipients, derivative use, model retention, refreshes, security, audit, deletion and liability.

Option value should be considered separately from the first fee. If the firm spends money to create a data dictionary and correct its outcome labels, that work may also improve stock forecasting, training and customer support. It can make later, permitted data products cheaper. Yet a buyer that demands exclusivity or unrestricted derivative rights could capture much of that future upside. Management should ask what it would give up under each licence term, who pays for refreshes, and whether it can withdraw records later found to be wrong or impermissible. A high headline fee may be inferior to a narrower contract that preserves these choices.

6. Permission, trust and the decision to say no

The archive is likely to combine material from customers, employees, suppliers and platforms. Its owner must check customer agreements, supplier terms, confidentiality commitments and the legal basis for any personal-data use. Removing direct identifiers does not necessarily make free text anonymous: unusual events, locations and combinations of attributes can permit identification. The UK Information Commissioner's Office advises assessing identification risk against the release context and the information a motivated person could combine with the data.8 Depending on that assessment, the feasible product might be aggregate statistics, a controlled evaluation, an internal-only system or no use.

The EU Data Act, applying since September 2025, creates specific access and sharing rights around data from connected products and related services, with protections relevant to smaller firms.9 It is not a universal permission to monetise any record held by a company. An industrial distributor's rights will depend on what generated the information, who contributed it and the terms governing the proposed use. Recent OECD work emphasises that different AI data collection mechanisms have different implications for developers, data subjects and rights holders.10 A single broad label such as “our data” conceals those differences.

Exhibit 4 — Choose the next action, not a theoretical valuation

Finding after sample auditAppropriate next stepWhat to avoid
Strong internal task, uncertain external rightsControlled internal pilotOffering a raw export
Clear rights, distinct buyer task, measurable outcomesRestricted evaluation with agreed rubricPerpetual rights before proof of value
Weak outcomes or shifting product definitionsImprove collection and labelsQuoting row count as quality
Sensitive cases dominate, limited mitigationRetain or aggregate; seek adviceTreating name removal as sufficient
Preparation cost exceeds credible benefitStop or narrow scopeSpending on a full data product first

Management can now make a practical decision. Appoint an accountable owner, choose one unit of work, audit a representative sample, document rights and sensitive fields, and run a small blinded comparison against the present process. Only then should the firm ask whether a bounded licence adds more value than internal use and whether the customer-trust cost is acceptable. A well-supported decision to keep the data private is a commercial outcome, not a failure to discover an asset.

The discovery exercise should end with a short decision memo, not a glossy data catalogue. It should name the tested task and baseline, show the sampled coverage and defects, record unresolved permissions, quantify preparation and ongoing support, and state the permitted next experiment. If the result is inconclusive, specify what evidence would change the decision and who will collect it. This gives a small management team a defensible stopping point and prevents an exploratory conversation with an AI buyer from becoming an open-ended commitment to supply data.

Sources and further reading

  1. OECD (2019), Enhancing Access to and Sharing of Data: Reconciling Risks and Benefits for Data Re-use across Societies. Scope: context-dependent data value and governance; not a valuation of small-company archives. https://www.oecd.org/en/publications/enhancing-access-to-and-sharing-of-data_276aaca8-en.html
  2. Sambasivan, N. et al. (2021), “Everyone wants to do the model work, not the data work: Data Cascades in High-Stakes AI”, CHI 2021. Scope: qualitative study of AI data practice. Author publication page: https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/ ; DOI: https://doi.org/10.1145/3411764.3445518
  3. Suresh, H. and Guttag, J. (2021), “A Framework of Potential Sources of Harm Throughout the Machine Learning Life Cycle”, EAAMO. Final published version in MIT repository: https://dspace.mit.edu/entities/publication/9fa559d2-a9ac-43c4-8db6-bf041eeedba5 ; DOI: https://doi.org/10.1145/3465416.3483305
  4. US National Institute of Standards and Technology (2023), AI Risk Management Framework 1.0, Core, especially Map and Measure. https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
  5. Keller, A. et al. (2026), Expanding the AI Evaluation Toolbox with Statistical Models, NIST AI 800-3. https://www.nist.gov/publications/expanding-ai-evaluation-toolbox-statistical-models
  6. NIST (2026), “Announcing NIST's Artificial Intelligence Technology Evaluation (AITE)”. Scope: example of blind evaluation design, not a direct model for a small firm's infrastructure. https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite
  7. Gebru, T. et al. (2021), “Datasheets for Datasets”, Communications of the ACM 64(12), 86–92. Author manuscript: https://arxiv.org/abs/1803.09010 ; DOI: https://doi.org/10.1145/3458723
  8. UK Information Commissioner's Office, “How do we ensure anonymisation is effective?”, guidance accessed 9 October 2026. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  9. European Commission, “Data Act explained”, accessed 9 October 2026. https://digital-strategy.ec.europa.eu/en/factpages/data-act-explained
  10. OECD (2025), “Mapping relevant data collection mechanisms for AI training”, OECD Artificial Intelligence Papers, No. 48. https://www.oecd.org/en/publications/mapping-relevant-data-collection-mechanisms-for-ai-training_3264cd4c-en.html