Data Monetisation 101 / Section 2 / Chapter 5

Section 02 · Understand AI demand · Chapter 05

Why Real Human Workflows May Be Valuable Training Data for AI Agents

An operational archive becomes more interesting to an AI developer when it records the task, the evidence available, the actions performed and a verifiable result. Yet reconstruction, permission and buyer fit determine whether that history is usable.

15 min read13 referencesSite last updated 9 October 2026

Executive perspective

Executive perspective

A workflow is more than its final answer. Linked requests, evidence, actions and results can be relevant to developing or evaluating AI agents that perform multi-step work and interact with software tools.

Outcome links improve the test. An independently checked resolution provides stronger evaluation evidence than an administrative “closed” status, although success alone does not prove that every intermediate action was correct.

Sequence integrity is the asset. Case identifiers, chronology, system state, policy versions and documented corrections can matter more than millions of isolated messages.

Useful does not always mean licensable. Reconstructing episodes across systems can be costly, and employee or customer information may carry restrictions. A narrow, rights-cleared task pilot is a more credible starting point than a general offer of raw activity logs.

1. What an agent needs that a conventional answer does not show

Most organisations have extensive records of completed work: resolved helpdesk tickets, approved invoices, credit memoranda, maintenance reports, software patches and customer-service transcripts. Such artefacts tell us something about performance, but often conceal the route to the result. An approved payment, for example, does not explain which exception triggered a review, what evidence the reviewer retrieved, what system was updated or why the payment was eventually released.

That distinction matters for AI agents: systems designed to pursue a task through successive observations and actions, potentially using software tools or interacting with people along the way. A conventional text assistant may be evaluated on the relevance of its answer. A tool-using agent must also choose valid operations, respect permissions, handle changing system states and recognise when to stop or escalate. ReAct, an influential research approach, combines reasoning and environment-facing actions and provides a technical explanation for why action sequences deserve attention 1. It does not establish that an employer's activity logs will automatically improve agent performance.

Consider an engineering-support case. The customer reports intermittent equipment failure; a technician checks a diagnostic dashboard, compares readings with past incidents, orders a test, rejects one explanation, schedules a replacement and confirms operation over subsequent days. A final ticket reading “sensor replaced” discards most of the information an evaluator would need to assess the decision process. The complete chain can instead distinguish initial conditions, available evidence, attempted actions, human interventions and later observations.

It is important not to mistake recorded actions for access to an employee's private thought process. A log may show that a technician opened a manual and changed a setting. It does not reveal the precise reasoning behind that choice unless a contemporaneous explanation was recorded. Adding invented rationales after the event would weaken, rather than enrich, the evidence.

Exhibit 1. A workflow evidence chain and the minimum useful context

StageIllustrative support-case evidenceWhat a buyer can assessFrequent weakness
RequestCase ID, timestamp, issue description, prior stateWhat task was assignedMissing or vague objective
EvidenceDiagnostic checks, retrieved documentation, system stateWhat information was available at that pointLater information accidentally inserted
ActionTool call, parameters, actor and permissionsWhether the operation was relevant and permittedAction recorded without parameters
OutcomeIndependent test, reopened case, customer confirmationWhether an objective success criterion was met“Closed” used as a proxy for success
CorrectionFailed attempt, override, escalation and subsequent resultRecovery and boundary handlingFailed steps edited out

Interpretation: The unit of analysis is a task episode, not an email or database row. A stable case identifier and ordered event history permit reconstruction; independent outcome checks add confidence without proving causation.

2. Research demonstrates the interest, not a universal data price

Published agent benchmarks offer tangible evidence that multi-step work is an important AI research problem. Mind2Web assembled more than 2,000 tasks from 137 websites across 31 domains, with crowdsourced sequences of web actions 2. Its significance for enterprises is methodological: recording real interaction trajectories can make tasks more realistic and transferable than isolated question-and-answer pairs. It does not imply that arbitrary browser histories are immediately suitable for training.

SWE-bench took another route. Its original release drew 2,294 software-engineering issues and associated pull requests from 12 Python repositories, asking models to produce patches that address real problems 3. Here, the test combines a request with a codebase and an externally checkable outcome. It illustrates why a buyer may care about the relationship between a task and its result, while also showing that rigorous evaluation may require runnable environments, not just static records.

Other benchmarks deliberately reproduce the operating complexity of software tools. WebArena tests long-horizon web tasks in functional sites; in its 2024 paper, a GPT-4-based baseline achieved 14.41% end-to-end task success against a reported 78.24% human benchmark 4. WorkArena uses enterprise-software tasks built around ServiceNow, addressing everyday knowledge work 5. And τ-bench examines agents in simulated user conversations with domain-specific tools and policies; its authors found reliability and rule-following problems across repeated trials 6. These figures describe particular research conditions and historical model baselines, not current industry-wide agent performance or demand for licensed enterprise data.

The commercial inference is narrower but useful: developers need ways to represent realistic tasks, evaluate permitted actions and measure completion. A company's workflow records might support those needs if they contain meaningful variation, recoverable context and appropriately verified outcomes. A buyer might instead prefer a synthetic environment, an internal collection process or a smaller curated sample. No public benchmark establishes a price per workflow episode or proves that third-party customers will purchase an archive.

Exhibit 2. Four distinct AI uses of the same operational history

UseWhat the workflow contributesEvidence neededLimitation
Behaviour imitation / fine-tuningExamples of which actions people took in specific statesOrdered expert demonstrations with tool contextHistorical human choices may be inefficient or unsafe
Evaluation and benchmarkingTasks, policies and expected results against which agents are testedIndependent success criteria, reproducible environment, withheld test casesA final label alone may miss unsafe intermediate actions
Retrieval / human assistanceRelevant precedents, fixes and procedures surfaced at task timeSearchable documents, provenance and appropriate access controlDoes not necessarily require training the model
Workflow redesign / internal analyticsBottlenecks, rework, exceptions and escalation patternsEvent completeness, timestamps, operational definitionsCommercial value may be internal, not licensable

Interpretation: A company should not label everything “AI training data”. The intended buyer task determines what must be extracted, annotated, restricted and tested.

3. Reconstructing the process is usually the expensive part

Operational evidence is rarely stored as a single coherent record. A customer request may enter a CRM; investigation takes place in a ticketing platform; technical checks appear in monitoring logs; approval occurs in a separate workflow; and the outcome is found in a later customer message. The business can therefore have all the underlying information and still lack an intelligible sequence.

A credible reconstruction begins with event-level identifiers and chronology. The data team needs to know whether case identifiers are stable across systems, whether timestamps use comparable time zones, which user or service performed an action, and what software and policy versions were in force. Events should be joined through defensible keys rather than speculative similarity matches. Missing links should be labelled as missing, not silently inferred. Where a customer outcome occurs weeks later, the measurement window and absence of follow-up must also be recorded: “no reopen event observed” is not the same as proof of permanent resolution.

The requirement becomes stricter when evidence might be used for testing. Evaluation tasks must preserve the information an agent would actually have known at the decision point. If a dataset exposes a future resolution note among the initial documents, it gives the agent an impossible advantage. Closely related cases spread across training and evaluation sets can also inflate apparent performance. Sensible splits may therefore separate customers, projects, incidents, time periods or whole organisations, depending on the task.

Provenance deserves the same care as the event chain. Dataset documentation work, including Datasheets for Datasets and Google's Data Cards, argues for recording collection methods, intended uses, limitations and relevant governance context [7,8]. For enterprise workflows, this becomes an operational question: who created the record, under what process, how was it transformed, and what was omitted during de-identification? Version histories and transformation logs can make the difference between a repeatable pilot and an opaque data export.

High-frequency routine work and rare exceptions should also be sampled deliberately. As Chapter 4 explains, older periods can improve coverage only if the historical task and outcome definitions remain comparable. A million near-identical password resets may add less task diversity than a much smaller set of difficult escalations. But rarity is not itself a guarantee of value: unusual cases can be poorly documented, legally sensitive or too idiosyncratic to generalise. A sample should reveal distributions of tasks, outcomes and interventions, not simply celebrate the most colourful incidents.

4. Failure, correction and outcome labels require careful interpretation

A failed attempt can be informative. It may show how a skilled worker recognises an unproductive path, rechecks an assumption or moves an issue to a specialist. Yet copying failed behaviour without a corresponding warning could teach an agent the wrong pattern. A useful record distinguishes “attempted”, “rejected”, “escalated” and “approved”, ideally with reliable timestamps and a reason code when available.

There is also a subtle causal problem. Human work logs record actions selected under a particular policy, by particular employees, facing particular customers. They do not show what would have happened had another action been chosen. If the toughest cases are systematically escalated to senior staff, outcomes cannot be compared naively with the easy cases handled automatically. Similarly, a successful payment or support ticket may involve several interventions, any one of which was decisive. Historical correlation should not be presented as proof that an action caused the outcome.

Research on process supervision offers a related, but distinct, insight. In mathematical reasoning tasks, Lightman and colleagues found that supervising intermediate steps could outperform feedback on the final answer 9. This is not a direct result about enterprise workflows or tool-using agents. It does, however, motivate the practical distinction between labelling each step and attaching a single case-closure flag. Both can help, but the evidential requirements differ.

Outcome definitions must be explicit. For software, passing tests may be a useful signal but not comprehensive proof of production correctness. For customer support, “ticket closed” may measure administrative status, whereas “not reopened within 30 days” measures something else and is itself sensitive to observation windows. In credit or fraud operations, delayed losses, overrides and policy changes may complicate apparently straightforward labels. Buyers need to know who verified results, which tests were run and whether any labels were independently reviewed.

Exhibit 3. Illustrative yield from a workflow archive (not a market estimate)

Assume a company has 120,000 historical support cases. The fractions below are hypothetical sequential retention rates: each is applied to records surviving the previous step. They do not indicate observed industry averages.

ScenarioLinkable across systemsOutcome reliably observablePermitted and fit after controlsUnique, sufficiently complete after deduplicationFinal candidate episodes
Constrained60%55%40%70%11,088
Base case85%70%60%80%34,272
Strong controls95%85%80%90%69,768

Formula: Candidate episodes = 120,000 × linkage fraction × observable-outcome fraction × permitted/fit fraction × completeness fraction. For example, the base case is 120,000 × 0.85 × 0.70 × 0.60 × 0.80 = 34,272.

Interpretation: The archive's headline count is a weak proxy for usable yield. If a hypothetical pilot costs £65,000 to extract, review, secure and document, the implied preparation cost per candidate episode ranges from about £5.86 (constrained) to £0.93 (strong controls), with £1.90 in the base case. These are seller preparation costs, not licensing prices or indications of buyer willingness to pay. Neither the fractions nor the budget comes from market evidence.

5. Rights, privacy and security can decide whether a dataset leaves the business

A company may possess an operational record without having unrestricted rights to licence every element. Workflows can contain customer correspondence, employee information, licensed third-party material, private source code, credentials, commercial secrets and contractual restrictions. Even if the company owns its software systems, it must still investigate underlying confidentiality obligations, intellectual-property rights, supplier terms and restrictions attached to the data's original collection and use.

Personal information presents an additional legal test. The UK's Information Commissioner's Office states that AI development and deployment require identified purposes and appropriate lawful bases for personal-data processing; its data-sharing code sets out fairness, accountability and security considerations [10,11]. Re-using a customer conversation to train another organisation's agent is not automatically permitted because the conversation was lawfully collected to deliver support. Depending on the processing, transparency, purpose compatibility, international transfers, contractual terms and data-subject rights may require attention. Legal assessment must be specific to the parties, jurisdictions and proposed use.

Removing names is not a complete answer. Ticket descriptions, dates, unusual events, free-text messages and combinations of seemingly harmless fields can still identify someone. The ICO distinguishes pseudonymisation from anonymisation: pseudonymised records remain personal data, even where identifiers are replaced or held separately 12. Similarly, an internal case identifier should not become an invitation to reveal a real person's identity in a buyer-facing file.

Controls can include field-level minimisation, masking and review of unstructured text, contractual restrictions, access-limited evaluation, secure processing environments and retention limits. Some use cases can be supported by derived task descriptions or simulated workflows instead of exporting raw conversations. Such transformations may reduce disclosure risk, but could remove crucial context or make outcomes impossible to verify. The feasible product is the one that balances task usefulness with provable permissions, not the one containing the most information.

NIST's generative-AI risk guidance recommends diligence around intellectual-property and privacy risks associated with training data 13. Executives should treat that as a practical procurement and governance concern. Without a rights map, provenance record and defensible control plan, the first commercial conversation may finish before the technical merits are considered.

6. Commercial test: make one narrow task credible before pricing an archive

A buyer is unlikely to begin with the abstract proposition that “our company has twenty years of workflows”. A more useful proposition specifies one repeatable task: diagnosing a particular class of equipment faults; responding to defined support requests; reconciling invoices under a written policy; or evaluating whether an agent follows an escalation procedure. The seller can then show the fields, decisions, permitted actions, success criteria and limitations that would support that task.

A practical pilot begins with a small, authorised sample selected across ordinary and difficult cases. A multidisciplinary team reviews linkage accuracy, observes a few complete episodes, checks label quality, flags sensitive information and compares the proposed sample with the buyer's stated use. If the buyer needs an executable test environment, static historical episodes may be only one ingredient. If the goal is internal retrieval assistance, the required transformation may be far simpler. In either case, test utility before constructing a large-scale export.

Commercial viability can be assessed as a conditional business case rather than a speculative record-price multiplier. The seller should estimate extraction engineering, legal review, security controls, annotation, ongoing updates and potential operational disruption. On the benefits side, distinguish measurable internal gains from externally negotiated licensing revenue, the latter of which is unknown until a buyer expresses concrete demand and terms. A useful management decision is often to stop a costly project early when rights are uncertain, task fit is weak or the data cannot be reconstructed reliably.

Exhibit 4. Proposed pilot acceptance gates

GateEvidence to produceStop or redesign signal
Task fitBuyer task description, required decisions, tool environmentNo specific capability or evaluation objective
Chain integrityReconstructed episodes, join methodology, missingness reportUnverifiable joins or pervasive gaps
Outcome reliabilityOutcome definition, independent checks, review sampleStatus flag routinely confused with success
Rights and controlsRights matrix, privacy assessment, approved access planContractual prohibition or unmanageable disclosure risk
Unit economicsCosted pilot, projected usable yield, update burdenPreparation cost cannot be justified by plausible use

Interpretation: These are episode-level pilot gates, beyond the portfolio screen in Chapter 3. Passing a gate does not establish saleability; it establishes a basis for proceeding to the next, more expensive step.

Practical implications for operating companies

Begin with the question “Which decisions and outcomes could an independent reviewer reconstruct?” rather than “How many logs do we hold?” Identify one operational area where procedures are reasonably stable, cases share identifiers and outcomes are not entirely subjective. Ask operational managers to demonstrate how a case develops across systems, then have data specialists verify that those events can actually be reproduced from the archive.

Next, build a sample inventory of tasks, systems, timestamp quality, exceptions, outcome windows and missing fields. Involve legal, security and information-governance teams before data is disclosed externally. Specify the intended use separately for training, evaluation, retrieval or internal process improvement; these purposes can require different permissions and preparation. Demand a concrete pilot evaluation plan from any prospective buyer, including how value will be measured and what happens to the data afterward.

The opportunity is real as a category of potential use, not a guaranteed market for every enterprise archive. Recorded operational work may contain demonstrations and test cases difficult to reproduce convincingly from generic text. But value emerges only when a buyer can understand the task, trust the sequence, verify the outcome and obtain the material on acceptable terms. The defensible asset is not the volume of digital exhaust; it is usable, rights-cleared evidence of how work was performed and whether it succeeded.

Sources and further reading

  1. Yao, S. et al. (2023), ReAct: Synergizing Reasoning and Acting in Language Models. ICLR. https://arxiv.org/abs/2210.03629
  2. Deng, X. et al. (2023), Mind2Web: Towards a Generalist Agent for the Web. NeurIPS. https://arxiv.org/abs/2306.06070
  3. Jimenez, C. E. et al. (2024), SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR. https://arxiv.org/abs/2310.06770
  4. Zhou, S. et al. (2024), WebArena: A Realistic Web Environment for Building Autonomous Agents. ICLR. https://proceedings.iclr.cc/paper_files/paper/2024/hash/4410c0711e9154a7a2d26f9b3816d1ef-Abstract-Conference.html
  5. Drouin, A. et al. (2024), WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks? ICML. https://proceedings.mlr.press/v235/drouin24a.html
  6. Yao, S. et al. (2024), τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/abs/2406.12045
  7. Gebru, T. et al. (2021), Datasheets for Datasets. Communications of the ACM 64(12). https://www.microsoft.com/en-us/research/publication/datasheets-for-datasets/
  8. Pushkarna, M. et al. (2022), Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. ACM FAccT. https://research.google/pubs/data-cards-purposeful-and-transparent-dataset-documentation-for-responsible-ai/
  9. Lightman, H. et al. (2023), Let's Verify Step by Step. https://arxiv.org/abs/2305.20050
  10. Information Commissioner's Office, How do we ensure lawfulness in AI? (Guidance; under review following legislative changes). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-do-we-ensure-lawfulness-in-ai/
  11. Information Commissioner's Office, Data sharing: a code of practice. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/data-sharing-a-code-of-practice/about-this-code/
  12. Information Commissioner's Office, Pseudonymisation (Guidance; under review). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  13. National Institute of Standards and Technology (2024), Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf