Training AI Agents on Corporate Archives Teaches Them to Lip-Sync
Google paid $10M for Spirit Airlines' emails and chats. Train agents on that, and they learn to mouth the motions of work, not to actually do the job.

Blame It on the Archive: The Record of Work Is Not the Work
Labs are paying millions for the messy record of how work gets done. Most of that record is the performance of work, and the enterprise win is to extract the real knowledge and fix the process, not to imitate the mess.
In 1990, Milli Vanilli won the Grammy for Best New Artist. Rob Pilatus and Fab Morvan had the hits, the choreography, the hair, and the trophy. What they did not have was their own voices on the record, because hidden session singers had done the actual singing.
The illusion held until the duo had to perform live. At a 1989 concert, the backing track skipped and looped the same line over and over while the two of them stood on stage with nothing to sing. Within a year the truth was public, the award was revoked, and the act became shorthand for credit that never matched the work.
Labs are now paying serious money for the corporate equivalent of that backing track. Google won a bankruptcy auction for Spirit Airlines' internal data, agreeing to pay $10 million for roughly 100 million emails and 500 million Microsoft Teams chats, pending court approval. Failed startups are selling their Slack, Jira, and email archives through data resellers for hundreds of thousands of dollars, and Forbes described the material as operational exhaust that labs now treat as premium training data.
The bet is that the record of how a company works can teach an agent to do that work. Train an agent on the performance, though, and you have taught it to lip-sync the job. The record of work is not the work, and almost everything that follows for an enterprise adopting agents depends on telling the two apart.
What the Archive Actually Contains
The labs are not buying this data on a whim. Agentic models need training material rich in context and feedback: documents, emails, tickets, and the back-and-forth that shows how a request turns into an outcome. Forbes reported that labs use these archives to build reinforcement learning gyms, simulated workplaces where agents practice the motions of office life. My own reading is that public web text has already given models most of what it can about language, and the next frontier is behavior inside organizations.
The problem is the curriculum. Somewhere, a reinforcement learning gym is teaching an agent that the correct response to a production outage is a 45-minute sync with no agenda.
Consider a vendor contract coming up for renewal, the kind of routine approval that passes through a handful of inboxes every quarter. The contract lives in one repository, the spend sits in the ERP, and the approvals are scattered across email and chat, so no single place shows the whole picture. A dozen messages and two approvals later, the renewal is signed and the task is marked done. To anyone reading the archive afterward, that looks like a clean, well-run procurement process.
Almost none of it was the work. The work was one person remembering that this vendor slips an automatic renewal clause into every draft, and that legal flagged an indemnity term last cycle that never made it into the standard template. The dozen messages were coordination, the cost of getting the right people to look before the signature, while the signed contract proves little, because the non-standard liability clause it waved through is what costs the company months later. An agent trained on that sequence learns the ritual of renewing a contract, not the judgment that catches what the template missed.

That scenario is not an edge case. It is the shape of most knowledge work. The archive is dense with performance and coordination: status updates, reply-alls, approvals, and the message that exists so someone can say they sent it.
The judgment that created value rarely gets typed out. The person who knows which vendor plays games with renewal terms does not file a ticket about it; she reads the draft and flags the clause. Experienced people compress years of exceptions into a glance, and the glance leaves no trace in Slack.
An archive records what people said to each other, not what they knew when they decided.
There is a second distortion, and it is worse. A closed ticket looks like success whether or not the fix held. A signed contract looks like diligence whether or not the clause was caught. The archive scores outcomes at the moment of closure, which is often the moment before the real outcome arrives.
Then there is the source. The archives on the market come disproportionately from companies that did not survive, which makes them literal records of what did not work. One commenter on the Forbes story asked the obvious question: why would anyone want a model trained on how to run startups that died?
The consequence for a leader is direct. A buyer or a deploying enterprise that treats the whole record as examples of competent work is training a convincing imitation, not a capable worker. The imitation will write the right messages in the right tone at the right time, and it will wave through the same liability clause, because nothing in the record told it that clause mattered. Defining what a correct result actually depends on remains the firm's own work.
Whose Knowledge Is It, Anyway
Spirit's archive did not belong to the pilots, gate agents, and analysts who wrote it. Lawyers quoted in coverage of the sale noted that employee work data is generally not private; it belongs to the company. The ombudsman review that favored the deal focused on consumer privacy and did not assess the risks to the employees whose messages were being sold.
This raises a question most organizations have never had to answer, because the knowledge never used to be portable. The business process and its records belong to the employer, and that employer can now share or sell them to a model lab, which means a workforce's accumulated judgment is suddenly an asset that can leave the building. The uncomfortable part is that the tribal knowledge worth having is exactly the part the records capture worst, so what gets sold is the mess, and what matters stays locked in people's heads.
The productive response is not to hope an agent will infer the judgment from the exhaust. It is to extract that judgment deliberately into artifacts that both people and agents can use, the unglamorous work of writing down how a thing is actually decided, what a correct result depends on, and which exceptions are real. A well-made document that states the rule, the edge cases, and the definition of done is worth more to an agent than a million messages, because it carries the judgment instead of the performance. That extraction is high-value work, and it should be recognized and governed as such, including who owns it and whether it is allowed to leave.
Start with governance. The decision about what tribal knowledge gets captured, where it is stored, and whether it ever leaves the organization is now a real decision with real stakes. Work records are never clean: they carry health disclosures in a sick-day message, salary details in an offer thread, and frustrations vented to a colleague at 11 p.m. Resellers promise to strip personally identifiable information, but scrubbing a name from a message does not scrub the context that identifies the person.
The durable method turns tacit judgment into structured, governed artifacts. That means a source-of-truth document for each decision that matters, a decision rule precise enough to test, and evaluation criteria that say what a correct result looks like before anyone builds against them. People query these artifacts when they need an answer, agents consume them as context, and the same artifact serves both.
This is context architecture treated as a discipline rather than a prompt afterthought. The team that does it well interviews the person who catches the renewal clause, writes the rule down, tests it against last year's contracts, and assigns an owner who updates it when legal changes the template. The team that does it badly pastes a policy PDF into a system prompt and hopes.
The extraction itself is expert work. The person who can articulate why a deal is risky is doing something harder than the deal, and if that effort is treated as a free byproduct, people will either refuse or do it badly. Recognize it, compensate it, and decide up front who owns the result and whether it is allowed to leave.
Fix the Process, Do Not Pave the Cowpath
Here is the move almost everyone skips, and it is the one with the highest return. A great deal of the friction in a work process exists only because humans are the ones running it. The handoffs, the status updates, the meeting to reconcile two systems, the follow-up to the follow-up, all of that is the overhead of coordinating people, and none of it is intrinsic to the task. An agent does not need to be reminded, does not need to look available, and does not need a meeting to agree with another agent.
So the largest gains from agentic AI do not come from teaching an agent to reproduce the existing process. They come from redesigning the process once the human-coordination friction can be removed. The strongest result is often not an agent that participates in the Slack thread and the meeting, but an agent that reads the data, applies the rule, and emits the result with no human channel touched at all.
Teaching an agent to imitate the old workflow does the opposite: it bakes the organization's dysfunction into the automation and calls it progress. The question to ask of any process before automating it is not how do we make an agent do this, but how much of this would simply disappear if people were no longer the ones coordinating it.
Most coordination is a tax on having people in the loop, and agents do not owe that tax.
In practice, the instruction is concrete. Before automating any process, map it end to end and mark every step as either task or coordination. Task steps transform information: checking a clause, reconciling a figure, applying a rule. Coordination steps move information between people: the reminder, the handoff, the status meeting, the approval that exists because two teams do not trust each other's systems.

Then remove the coordination that existed only because humans ran it. Some approvals remain, because a person needs to own the risk. Many disappear, because they were proxies for visibility, and an agent's actions can be logged and audited without anyone attending a meeting about them. Automate the redesigned process, not the original.
This is why so many early agent deployments underwhelm. Teams point an agent at the existing workflow, teach it to draft the status update and chase the approver, and then wonder why cycle time barely moves. They automated the performance of a broken process instead of fixing it, so the agent now runs the same relay race faster while still handing the baton to people who are in meetings.
The largest gain I have seen on an agent project did not come from the model. It came from deleting a weekly reconciliation meeting whose only purpose was to compare two exports that disagreed. Once the agent read both systems directly and flagged mismatches against a written rule, the meeting had no reason to exist, and the cycle shrank from a week to the time it took the data to refresh.
The model was the easy part. The redesign was the part that paid.
Personal Agents in a Human-Shaped Internet
Meta launched Muse in the US in September for adult users, running each agent and its data on a dedicated virtual machine. It connects to email, calendars, payments, and shopping apps, and it can open a browser, fill out forms, and negotiate on a user's behalf. It reached the top of Apple's free app chart within ten days, and Mark Zuckerberg pitched it as an agent that works around the clock for you.
The internet these agents act in was built for people. Its apps, its chat threads, its email, its meetings, and its messaging are designed around human attention, and the newest personal agents, Meta's Muse, OpenAI's Dots, and xAI's Grok Bot, are designed to operate through exactly those human channels on an employee's behalf. Pointed at the existing mess, that is what they learn to do: participate in the performance of work, send the standup message that says nothing to report, and produce the output that looks right.
The deeper problem for an enterprise is where the knowledge goes. A personal agent that holds an employee's work logins and acts across Slack, email, and internal systems is reading the organization's process, its exceptions, and its personal data, often on a vendor's cloud and, on several consumer products, into training unless someone opted out. That is the firm's tribal knowledge leaving through an app the firm never evaluated.
For a small business with low regulatory exposure, that trade can be perfectly reasonable, and the consumer agents are useful there. For a regulated enterprise, an employee's personal agent on work systems is an ungoverned export of the very knowledge the earlier sections said to protect.
Both halves of that judgment are true. For personal errands, family logistics, and a ten-person firm whose most sensitive data is a customer list, an agent that clears the inbox and books the travel is a sound choice. The trade between convenience and exposure tilts toward convenience when there is little to expose and no regulator watching.
The calculus flips inside a regulated enterprise. An employee who connects a personal agent to work email has granted a vendor's system access to contract terms, customer records, and internal exceptions with a consumer terms-of-service click. TechCrunch framed Muse as a test of whether people trust Meta with that much access; the enterprise question is whether the firm ever got a vote.
Treat that as a design problem. Agents that act on work systems should run under identities the organization issues, with scoped permissions, logged actions, and data that stays inside boundaries the firm controls. Anything else is a governance and data-ownership problem dressed up as a productivity feature.
This is where the threads meet. Pointed at human channels, a personal agent reproduces the performance of work and quietly exports the knowledge behind it. That is the opposite of extracting judgment deliberately and redesigning the process, because the firm's proprietary process and knowledge are the real asset, and they are leaking out one connected app at a time.
The Best Case Against This
The strongest case against all of this is the one the labs are betting billions on. It says that enough authentic operational data, at sufficient scale, will let agents absorb tacit knowledge the way public code repositories taught models to write software, and that deliberately extracting knowledge and redesigning processes is slow, costly, and sure to meet resistance from the people being asked to write down what they know. If scale solved coding, the argument goes, scale will solve accounting and operations too.
Scale learns what is common across companies, while value lives in what is specific to yours.
The reply is about what scale actually captures. Scaling laws are general, and they are good at what is common across many companies, while the value in knowledge work is specific to one company's context, customers, and systems, which is exactly the part that does not generalize. The archives on the market make it worse, because the companies most willing to sell are the ones that failed, so the data skews toward how dysfunction operated, not how value was created.
Extraction and redesign are slower, and they are also the only path that leaves the firm with something it owns, can evaluate, and can govern, rather than a rented imitation. The labs are right about one thing, and it proves the point: for repeatable, high-drudgery work, the win is an agent that reads structured data and emits a result with no human theater at all, which is an argument for redesigning the process, not for imitating it.
Where the Value Was Made
Milli Vanilli could hold the pose until the track skipped, and an agent trained on the record of your work can hold it too, right up until the moment the work is real. The market will keep pricing the record of work, and that price says nothing about where the value was made. The leaders who get real returns will be the ones who extracted the judgment into artifacts they own, fixed the processes whose friction was never the work, and governed the agents that act in human channels rather than letting them cosplay the old mess.
The record of work is not the work. Telling the two apart is the whole discipline, and it is a judgment no archive and no agent makes for you.