Executive Summary

On August 14, 2026, the internal data of the bankrupt Spirit Airlines went to a virtual auction, and Google won it with a $10 million bid. Mercor was named backup bidder at $7.5 million, and Micro1 came in after the deadline with $12.5 million. What sold was everything the company produced while it was still working: internal email, Teams conversations, source code, payroll and timekeeping, and employee records reaching back to 1986, bundled into a single lot.

The remarkable part of this deal is not the price. It is the schedule. A one-page asset schedule filed with the bankruptcy court lists system names, row counts, and the month each dataset begins, each line marked Included or Not Included. Until now, the value of a company's accumulated working records has been an estimate. This is the first time it has been published as an itemized list. And what turns that entire list into a sellable object is a single word: deidentified.

The nine-page objection the flight attendants' union filed with the court takes aim at exactly that word. The direction it aims is not the one most coverage suggested. The union does not argue that re-identification is possible. Paragraph 13 of the filing expressly disclaims that argument. What the union says is that deidentification was built to solve a different problem in the first place. It governs whether a record can be traced to a named person; it says nothing about whether the contents of that record are confidential. Read the schedule and the pattern is stark: consumer-facing items are almost all excluded, employee items almost all included. The privacy machinery was cut to a consumer shape, while the payload actually crossing the table is employee-shaped.

$10,000,000

Google's winning bid

$0.10 per email, $125 per mailbox

100 million

Internal emails included

80,000 Microsoft 365 accounts, from 2018 on

175,658

Employee records included

From August 1986. On the same schedule, 97.5M customer profiles are excluded

4,600

Crew Base population

The figure the union names. A decade of linked records describes these people

1

The Bid Is In; Nothing Has Been Approved

The auction was held over video. On Friday, August 14, 2026, the Spirit Airlines debtors put a single lot up for sale under court-approved bidding procedures. The lot was named "Deidentified Data." At 10:00:10 p.m. that night, a notice of auction results was docketed in the U.S. Bankruptcy Court for the Southern District of New York — Case No. 25-11897 (SHL), Docket 1463. The winning bidder was Google LLC at $10 million. Named as backup bidder was Mercor.io, a startup that supplies human-labeled data to AI developers, at $7.5 million.

There is a third party in the room. Micro1, another data-supply startup, offered $12.5 million after bidding had closed — $2.5 million above the winning bid, but submitted after the process had shut. That three companies valued the same asset at $7.5 million, $10 million, and $12.5 million tells you something on its own: there is no settled market price for this kind of object yet.

A missed deadline does not automatically kill the offer, either. Bankruptcy judges rarely reopen a finished auction, but the code leaves room for a late bid that puts more money in creditors' pockets. By Forbes's account, two questions go before the court together on September 9: the union's privacy objection, and Micro1's extra $2.5 million. Micro1's founder and CEO, Ali Ansari, argued that decades of real operational records are exactly what is needed to train more capable models, and that Google's price was far too low for data that valuable. That this comes from the buy side is a preview of something we will return to: the pricing argument caught fire before the privacy argument did.

Spirit went bankrupt twice. It filed its first Chapter 11 in November 2024, emerged in March 2025 and relisted, then filed a second Chapter 11 on August 29, 2025. The restructuring failed, and on May 2, 2026 the airline stopped flying altogether. Roughly 14,000 direct employees and about 3,000 contractor jobs — some 17,000 in total — disappeared at that point. The data being sold now is what that shuttered company left behind. The company is gone; the records are not.

One thing has to be clear before anything else. As this report goes out, the court has approved nothing. The sale hearing was originally set for August 19, was pushed to September 9 after the flight attendants' union objected, and no ruling has issued. The deidentification the agreement requires has not been performed either. Forbes put it precisely: Google has won the bid, not the data. The table below traces where this case started and how far it has come. The top two rows are not events. They mark how far back the merchandise goes.

Sources: asset schedule (attached to ECF 1463), AFA objection (Doc 1489), PPC Land reporting. The September 9 hearing is scheduled; no ruling had issued as of publication.
DateEvent
August 1986Oldest employee record in the HR system (UKG). This is where the merchandise begins.
May 2008Oldest revenue record in the reservation system (Navitaire)
August 29, 2025Spirit Aviation Holdings and five affiliates file a second Chapter 11
May 2, 2026Operations cease entirely. About 17,000 jobs affected
June 18, 2026Confidentiality agreement signed between Google and Spirit
June 22, 2026Court approves bidding procedures (Docket 1213)
July 7, 2026European Data Protection Board adopts new anonymisation guidelines. Six weeks before the auction.
August 14, 2026Virtual auction. Google wins at $10M; Mercor backup at $7.5M
August 17, 2026Objection deadline (4:00 p.m. ET)
August 18, 2026Flight attendants' union (AFA-CWA) files a limited objection (Doc 1489)
September 9, 2026Sale hearing scheduled. No ruling yet.

Adjournments themselves are routine; objections push hearings back all the time. What matters here is what the extra time preserves. The remedy the union asked for is still operable, because the deidentification work the agreement calls for has not begun — so if the court decides to attach conditions, a protocol reflecting them can still be designed. That is why the union wrote, near the end of its filing, that "the Deidentification has not yet occurred; the protocol remains to be designed."

Timeline of the Spirit Airlines data sale case Aug 1986 Oldest employee record Aug 29, 2025 2nd Chapter 11 filed May 2, 2026 Operations cease Jun 22, 2026 Bidding procedures approved Aug 14, 2026 Auction: Google wins $10M Sep 9, 2026 Sale hearing (scheduled · no ruling)
▲ Pebblous original diagram — six milestones in the case timeline (spacing is visual, not proportional to elapsed time)

One gap is worth stating as a gap. We were not able to review the full docket, so we cannot say whether any party other than the flight attendants' union objected, whether the debtors filed a response, or whether Micro1's post-deadline offer was ever submitted as a formal filing. This report can say that the union objected. It does not say the union was the only objector.

2

One Page of Asset Schedule

The real document in this case is the one-page table attached to the purchase agreement. Both agreements — Exhibit A for Google, Exhibit B for Mercor — carry the same asset schedule, which lays out every dataset Spirit holds as a row and marks each row Included or Not Included. The heading on that decision column reads "Google's Data Purchase Request." The buyer picked what to buy.

To anyone who works with data, the table looks familiar. System name, item name (roughly a table), row count, the month the data starts — one line each. It is an enterprise data asset inventory. The only unusual thing about it is that it was not built while the company was alive. Below are the rows marked Included. Some rows carry no count; those are Included all the same.

Assets marked Included. Sources: asset schedule (Exhibits A and B to ECF 1463), AFA objection ¶¶7–8, PPC Land reporting. Where the two sources overlap, the figures agree.
CategoryItemCountSystem · Start
ProductivityEmails100,000,000Microsoft 365, from 2018 (80,000 accounts)
ProductivityTeams items500,000,000Microsoft 365
ProductivitySharePoint items20,577,677Microsoft 365
ProductivityOneDrive items17,082,644Microsoft 365
ProductivityIT tickets667,563ServiceNow
SoftwareSource repositories (~30M lines of code)516Web · mobile · backend
SoftwareCommits (history, diffs, authors, timestamps)372,585git
SoftwarePull requests (descriptions, comments, review threads)43,170git
RevenueTransactions7,510,221,520Navitaire, from May 2008
RevenueCompetitor fare observations7,250,630,887Infare, from Jan 2021
RevenueBooking-curve observations3,530,769,431Navitaire · Kambr, from Jan 2023
RevenuePassenger Name Records (PNR)190,312,864Navitaire, from May 2008
RevenuePayment-processing transactions78,432,834Elavon, from Dec 2022
RevenueOnboard retail · base fares · refunds · vouchers29,000,147 and moreRetail in Motion · PROS, etc.
OperationsIrregular-operations reaccommodations3,000,347,472ITS, from Feb 2018
OperationsIndustry traffic stats · scheduled flights · flown flights74,163,945 and moreFrom Jan 2017
OperationsCrew pairings5,014,676From Jan 2021
OperationsCrew Base (headcount)4,600Base assignments
OperationsMaintenance parts and tasks · fuel tickets1,239,196 and moreTRAX · SkyMetrix
HREmployee records175,658UKG, from Aug 1986
HRPayroll records3,426,618UKG, from Jun 2016
HRTimekeeping (timecards)1,092,000Kronos, from Dec 2012
HREmployee tax forms148,018UKG, from Jun 2016
HRCorporate and crew training records · travel records · recruiting filesNo count listediCIMS and others
Finance · LegalCost accounting · invoices · vendor IDs1,849,736 and moreSAP · Coupa, from Jan 2019
Finance · LegalFinancial models · board reporting · deal pipeline · M&A fairness opinionsNo count listedCorporate development files
Finance · LegalLitigation records · contract redlines · employment agreementsNo count listedLegal files

Two things the table cannot hold. The first is how wide the software category actually runs. What crosses over is not only repositories and commits but issue trackers, continuous integration logs, automated test results, and code coverage reports. Exhibit A then widens once more: the software grant extends to source code and object code, architecture, software models, plugins, algorithms, libraries, compilers, subroutines, tools and application programming interfaces, together with technical specifications, user manuals and training materials. What moves here does not stop at the code. It is the full operating record of the engineering organization that wrote and ran it.

The second is the name the asset definition gives itself. As the union quotes at ¶7 of its objection, Section 1 of Exhibit A defines the Assets to include productivity and collaboration data and core business systems data — expressly including "employee behavior and productivity data" — plus workflow and process data including "HR and legacy operations data." You do not have to read the schedule row by row. The agreement has already written down what it is selling.

2.1Everything Excluded Is Consumer-Facing

Gather the rows marked Not Included and one property jumps out. Every one of them sits at a customer touchpoint. Customer profiles, loyalty miles, marketing send lists, call recordings, web analytics, regulatory complaints — the consumer side came out without exception.

Assets marked Not Included. Sixteen consumer-facing rows were excluded without exception.
CategoryItemCountSystem · Start
CustomerCustomer profiles97,500,000Navitaire, from May 2017
CustomerFree Spirit members50,200,000Loyalty program
CustomerAccrued miles44,000,000Loyalty program
CustomerSavers Club members · cardholders2,200,000 / 740,000Paid memberships
MarketingActive email addresses13,700,000Oracle Responsys
MarketingPost-purchase survey responses · campaigns · social posts1,446,033 and moreQualtrics · Sprinklr
Contact centerCall recordings30,865,471Calabrio · Cresta, from Jan 2021
Contact centerChat sessions15,784,473Quiq, from Jan 2024
Contact centerPhone numbers7,341,857Customer contacts
Web analyticsUnique visitors · searches · purchases (YTD)43,117,864 and moreGA4 · Databricks, from Jul 2024
RegulatoryDisability service requests · U.S. DOT complaints2,491,715 / 338,531Netracer, from Mar 2025
Asset schedule: what's excluded is consumer, what's included is employee and operations Not Included 11 rows · all consumer-facing Customer (profiles, loyalty) Marketing (email, campaigns) Contact center (calls, chat) Web analytics Regulatory & complaints Included 27 rows · overwhelmingly employee & ops Productivity & collaboration Software & engineering Revenue (PNR, fares, booking curve) Operations & maintenance HR (employee records, payroll) Finance & legal Privacy was designed for consumers. The payload is shaped like employees.
▲ Pebblous original diagram — the 27 Included and 11 Not Included rows regrouped by category (source: §2 asset schedule tables)

2.2"No Passenger Data Was Sold" Is Half True

The common read of this table is relief: passenger data is all out. That is half right. What came out is the passenger profile. What goes across is the passenger behavior log. The 97.5 million customer records carrying names and contact details are excluded — but 190.3 million Passenger Name Records and 7.51 billion transactions, accumulated in the same reservation system since 2008, are included. Two decades of who flew where, when, and for how much travels intact, with only the names lifted off. Add 7.25 billion competitor fare observations and 3.53 billion booking-curve points, and the entire pricing logic of an ultra-low-cost carrier survives in reconstructable form.

Nor is the excluded customer list being destroyed. Section 1(b)(ii) bars Spirit from selling the Assets to anyone but the buyer — and then opens one door. Spirit may separately sell its customer data list, "which includes individual traveler spend aggregated by year," to third parties in the hospitality or travel industries. The passenger list was not withheld from the market. It was routed to a different pool of buyers.

2.3The Delivery Formats Are Specified Too

That this schedule is a practitioner's document shows in the delivery terms. Microsoft 365 material is to be transferred while kept "within the native Microsoft 365 environment" — because exporting mail and files to CSV breaks thread structure and metadata. Non-Microsoft systems such as SAP, Navitaire and UKG are exported as CSV, JSON, or native SQL dumps. Repositories go over as one archive per repo, as bare git repositories or bundle files. Commit, pull request and issue metadata go over as JSONL.

JSONL is what a training pipeline reads a line at a time. A bare git repository keeps history and branches alive. These are not the delivery requirements of someone who wants to archive material; they are the requirements of someone who wants to feed it into processing immediately. The Mercor agreement carries an extra Exhibit C specifying that deidentified data be transmitted to the buyer's secure upload facility in each source system's native format. The Google agreement has no such exhibit — the delivery specification sits inside the asset schedule instead.

Set the schedule against the agreement's own drafting and a small seam shows. The body of Exhibit A says the Assets include "marketing data (campaigns, content, demand generation data and the like)," while the schedule that same Exhibit A points to marks the campaign and social rows as excluded. Because Exhibit A expressly defers to the schedule, the narrower schedule controls and the actual scope of the deal does not change. But a general clause and a row-level determination pointing in different directions inside one contract tells you the table was assembled in a hurry.

2.4The Counts Are Stated; the Quality Is Not Warranted

One question the schedule never answers: are the records behind those numbers accurate, and are they complete? Both agreements disclaim every warranty in capitalised text. The buyer takes the assets "AS IS" and "WHERE IS" with all faults, waiving claims about merchantability, fitness, non-infringement and the accuracy or completeness of anything disclosed. Spirit warrants only that it owns the assets and has authority to sell them, and even those warranties do not survive closing. What remains is a claim in fraud.

From a data practitioner's chair, that combination is what the deal amounts to. The schedule is a statement of volume, not a statement of quality. The figure "100 million emails" says nothing about how much of it is duplicated, what share is automated notification traffic, or whether a 1986 record in the HR system still parses against today's schema. The $10 million was set against that unverified state. At the moment a company's working records first acquired a public price, the price attached. The quality statement did not.

3

The Deidentification Machine

One word turns that entire $10 million list into a sellable object. The lot itself is named "Deidentified Data," and in the agreement deidentification is defined not as a property the assets already have but as a process they must pass through before delivery. Follow how that process was designed and you can see who the privacy machinery was built to protect.

The starting point is already strange. Section 1(a)(i) declares it the parties' express intent that the Assets contain no information linked or reasonably linkable to a consumer, which is to say no personal information. The very next subsection, 1(a)(ii), requires a process to strip that personal information out before delivery. Remove the thing you just declared absent. This is the sequence PPC Land flagged as unusual.

The mechanics run like this. Spirit first hands the files to a "Deidentification Agent." The agent removes or transforms data elements so the result cannot be linked to, or used to infer information about, a particular consumer. Only after that work is complete does the material move to the buyer. The question is which side of the table every lever on that process is bolted to.

Where control over the deidentification process sits. Source: sale agreement clauses as quoted in the AFA objection ¶¶9 and 15 and in PPC Land's reporting.
ItemClauseWho decides
Naming the Deidentification Agent§1(a)(ii)A party "acceptable to or designated by" the buyer
Standard for certification§3(c)The buyer's reasonable satisfaction
Right to review and comment on the process§4(c)The buyer. Spirit must consider those comments in good faith
Who pays§1(a)(iii)The buyer — with no reduction in the purchase price
Onward transfer to third parties§1(c)Buyer's discretion. No cap on hops, no duty to disclose recipients
No-deletion covenant§3(e)The debtors cannot pull material out even if they want to
Third-party beneficiary status§12None. Flight attendants have no way to enforce anything

Read the table down its middle and one direction emerges. The authority to decide how much deidentification, by what method, and to what line, sits with the side buying the data. Then §3(e) is layered on top: the debtors covenant that they have not "deleted, modified or removed" any portion of the Assets, apart from the deidentification itself, ordinary-course changes, and minimal removals to protect their own privilege. The union's conclusion about that combination is at ¶15 of its objection. The structure leaves the debtors no room to pull confidential employee information out, even if they wanted to.

The fifth row — onward transfer — carries one more device. Section 1(c) requires the buyer to commit publicly to maintaining and using the data in deidentified form and not intentionally associating it with any person or household. The same clause then supplies an alternate route: if the buyer has not made that commitment in a broadly accessible public document, such as a posted privacy notice, the agreement itself is deemed to constitute the public commitment required under applicable data protection law. Consumer privacy statutes that require a "public commitment" were written imagining a document anyone can read. Here, a private contract attached to a bankruptcy docket fills that slot. It is another view of what the union means when it grants that the protections in this agreement are real but all calibrated to one axis.

3.1One Line About Referential Integrity

Section 3(c), which sets the certification standard, attaches a condition that softens the deidentification. The process must run "while preserving referential integrity across the data set." To anyone who works with data, that sentence is an unambiguous instruction: keep the join keys between systems alive. The links that make an email, a timecard, a pairing and a payroll adjustment point at the same entity have to survive if you want to follow how one person's day unfolded.

The joins are what give the corpus its training value. A hundred million de-named emails is a large body of text on its own, but if you cannot tell which timekeeping record, which pairing, and which pay adjustment a message connects to, it is a pile of prose. Value comes from the relations between documents, not the count of them. And that is precisely the property the union is worried about.

One clause defines two things at once. Referential integrity is the property that makes this dataset worth training on, and it is the property that makes anonymisation fragile. The condition for data quality and the condition for privacy risk overlap inside a single line of the contract. Separating them means cutting the joins, and cutting the joins lowers the value of what the buyer bought.

3.2The Standard the Contract Chose, and the Paper Aimed at It

Who decides that deidentification was done properly? The agreement names two U.S. standards. One is the California Consumer Privacy Act, and the contract specifies it will use that standard whether or not the statute actually applies to these assets. The other is 45 C.F.R. §164.514, the HIPAA deidentification rule for health information — the provision that defines the Safe Harbor and Expert Determination pathways.

In February 2026, a team at NYU posted a paper on arXiv titled "Paradox of De-identification: A Critique of HIPAA Safe Harbour in the Age of LLMs." Its abstract lands on the same spot this contract stands on.

"Safe Harbor was designed for an era of categorical tabular data, focusing on the removal of explicit identifiers while ignoring the latent information found in correlations between identity and quasi-identifiers, which can be captured by modern LLMs."

Jiang, Liu, Cho, Oermann, arXiv:2602.08997 (2026-02-09), abstract

The scope needs stating exactly. This is a position paper, not an empirical study that settled anything with experiments, and its subject is clinical records, not corporate email. It does not carry over to internal business records as a result. The shape of the situation does carry over, though: the rule the contract chose as its warrant of safety is being criticised for its limits under precisely the condition this contract creates — the existence of language models to train on the material.

The unstructured nature of the text pushes the same way. Formal measures like k-anonymity operate over enumerable quasi-identifier columns. Narrative text has no columns to enumerate. That is the distinction behind the union's line at ¶12: "A process that removes identifiers from structured fields is a poor instrument for that content, and the Sale Agreement does not purport to review it." Against 100 million emails and 500 million Teams items, litigation files and employment agreements — material whose confidentiality lives inside the sentences — field-level masking is aimed at the wrong target from the start.

One more fact worth adding carefully. A 2022 study quantifying how language models memorize training data found that memorization scales log-linearly with model capacity, with how many times an example is duplicated in the data, and with the length of the prompting context. The authors reported that a 6-billion-parameter model memorizes at least 1 percent of its training dataset. Stop there. That is a lower-bound estimate for particular models and datasets, and it is not the claim that "if Google buys this, Spirit employees' emails will start falling out of a chatbot." The union does not make that claim either. What does connect is the underlying property — more duplication and longer context mean more memorization — and the shape of this asset: a single corpus that one organization built over two decades, with its parts constantly referring to each other.

3.3The European Standard Was Not Targeted

Six weeks before the auction, the European Data Protection Board adopted new guidelines on anonymisation, replacing the Article 29 Working Party standard in use since 2014. As reported, the guidelines set out three tests: records must not be singled out, must not be linkable to outside data, and — the newly prominent one — must not permit inference. That last test treats anonymisation as having failed if specific, meaningful conclusions about an individual can be drawn without singling out any record or linking to anything external.

A dataset that preserves referential integrity has exactly the property that test is aimed at. The Spirit agreement certifies against the California standard and a U.S. health regulation; it does not aim at the European one. There is a reason this contrast stops here: the basis for this paragraph is PPC Land's reporting, and we were not able to read the guidelines themselves. So the direction is recorded without article numbers or verbatim quotation.

3.4What the Same Company Argued in Other Venues

After the auction result circulated, a LinkedIn post by privacy attorney Alan Chapell — chairman of the board of the Network Advertising Initiative — put a different face on the case. His point was compact: "In a NY bankruptcy court, Google is about to purchase a trove of data from the Spirit Airlines bankruptcy. But don't worry they say, it's not personal data - it's been anonymized." Meanwhile, he wrote, "in two separate antitrust/competition venues, Google is OBJECTING to the provision of search data to rivals - effectively stating that it's all but impossible to truly anonymize data." That characterization is Chapell's reading, not a statement by Google.

The two proceedings he points at are on the record. In September 2025, Judge Amit Mehta ordered Google to share its Glue query-interaction system and RankEmbed training data with qualified competitors; Google appealed the resulting six-year judgment in January 2026 and asked the D.C. Circuit in May 2026 to throw out the liability finding entirely. In Brussels, the European Commission adopted binding Digital Markets Act decisions in July 2026 requiring anonymised query, click, ranking and view data to flow to rival engines, after finding that the company's first compliance offer had stripped out between 90 and 100 percent of unique search queries. The technical specification behind that remedy ran to 29 pages of field-level anonymisation requirements.

The two positions are not logically contradictory. Search query logs and internal corporate email are different kinds of material, and the interests of a party handing data over run opposite to those of a party receiving it. What remains is the shape: the same company argues the limits of anonymisation when it has to give data up, and assumes its sufficiency when it is receiving. Which is another way of saying that anonymisation is being treated less as a technical question than as a question of where you happen to be standing. Google did not comment publicly on the Spirit purchase in the court documents reviewed.

4

What the Union Actually Argued

Several outlets summarized the flight attendants' objection as a fear of re-identification. Read the filing and you find the opposite sentence. It is the first line of ¶13.

"AFA does not make a technical claim that any particular record can be re-identified, and it does not need to."

AFA-CWA Limited Objection, Doc 1489, ¶13

What the union contests is not the performance of the anonymisation technology. It is that anonymisation is aimed at a different problem altogether. Paragraph 2 defines the distinction in a single passage.

"Deidentification addresses whether a record can be traced to a named individual. It does not address whether the contents of the record are confidential. A flight attendant's disciplinary correspondence, a crew training deficiency, a leave or accommodation request, an internal Teams exchange about staffing or scheduling grievances, and a payroll adjustment history each remain sensitive employment information whether or not the employee's name has been stripped from it."

Same filing, ¶2

The distinction carries legal weight because the Bankruptcy Code already knows it. The union cites §107(b), which protects "confidential commercial information" regardless of whether any individual is identifiable. Confidentiality and identifiability are separate axes in the statute too. Every protection the agreement installs is graduated along the identifiability axis; the flight attendants' interest sits on the confidentiality one.

The union does not belittle the protections in the agreement. It grants that excluding personal information, excluding privileged material, requiring CCPA-based certification, and extracting a public covenant against re-association are real, and were not automatically owed. The problem is that all of them are calibrated to the same axis.

4.1What Survives the Removal of a Name

Paragraph 12 spells out what name-independent sensitivity looks like in practice. Even a pseudonymized dataset can reveal which crew bases generated the most grievances, how a small number of flight attendants scored in retraining, which employees were under investigation, what compensation adjustments followed which incidents, and what employees said to each other about management, staffing, and their own union. None of that sensitivity comes from a name. It comes from content and context.

Then, at ¶13, the union brings the referential-integrity clause back. With joins preserved, the Crew Base population is 4,600. "Where a small, highly structured population is described across linked operational and communications datasets spanning more than a decade, the risk that information about identifiable individuals or small identifiable groups can be inferred is not speculative, and the Buyer's covenant reaches only intentional association." Note where this sits in the filing, though: it is a supporting argument, not the spine. The spine of the union's case is confidentiality.

4.2Why the Bankruptcy Code Does Not Reach Employees

It is not that the Bankruptcy Code has no machinery for data sales. Section 363(b)(1) ties the sale of personally identifiable information to compliance with the debtor's privacy policy or to the appointment of a consumer privacy ombudsman. The problem lies in the definition. Section 101(41A) defines personally identifiable information as information provided to the debtor by an individual in connection with obtaining a product or service. An employee is not a consumer obtaining a product or service from an employer's mailbox.

Here the union's filing takes a step back. It does not argue that §363(b)(1) is triggered in this case. It concedes the opposite. And at ¶11 it writes that this is exactly the point: the framework the parties borrowed was built for customers, and the employment records this transaction actually transfers get nothing comparable in the way of review.

The sentence the union grounds in the schedule at ¶3 is the quantitative spine of that argument. Consumer-facing categories are designated almost entirely as excluded; categories under the "Team Member" heading are designated almost entirely as included. Hence: "The privacy architecture of this transaction is consumer-facing; its payload is disproportionately employee-facing." Employee data is more confidential than customer data, and receives less protection. Reduce the contrast to a pair of numbers and it reads: 175,658 employee records included, 97.5 million customer profiles excluded. Within one schedule, the same category — an identity-bearing profile record — got opposite verdicts.

4.3What the Union Actually Asked For

The union is not trying to stop the sale. Its opening sentence nails that down: it does not seek to disrupt the debtors' sale process, unwind the auction, or prevent the estate from monetizing data assets. The request runs in two branches. First, exclude all flight-attendant-related information from the assets. Second, if the court lets the sale proceed, extend to former Spirit flight attendants at least the same level of protection the deal gives consumers. Five conditions hang off that second branch.

  • The deidentification process should include a review or segregation protocol aimed directly at confidential employee and labor-related information, rather than relying on the removal of consumer identifiers alone — covering personnel, payroll, tax, grievance, disciplinary, investigative, medical, leave, accommodation, training, performance-evaluation and union-related records.
  • For employee email, Teams, OneDrive and SharePoint content, apply measures appropriate to free-form communications, and ensure the no-deletion covenant is not read to block removal of employee-related content that the sale order requires.
  • Bar the winning bidder from using these assets to analyze, profile, evaluate, score, or draw conclusions about individual flight attendants or identifiable groups, and bar attempts at re-association.
  • Impose the same restrictions in writing on any onward transfer of employee-related assets, make them enforceable by the debtors or their successors, and notify the union of which categories were transferred onward.
  • Any further protective measures the court deems appropriate.

Can a court attach such conditions? The union says yes and offers three grounds. First, §363(b)(1) permits a sale outside the ordinary course only "after notice and a hearing," and the court must find an articulated business justification, considering "all salient factors pertaining to the proceeding" and the interests of the parties affected (In re Lionel Corp., 722 F.2d 1063, 1071 (2d Cir. 1983)). Second, the same authority that permits approval permits approval on terms — 11 U.S.C. §§105(a), 363(e). Third, it cites In re Trans World Airlines, 322 F.3d 283 (3d Cir. 2003), where employee-related rights constituted "interests" subject to §363(f) and to the court's treatment in the sale order. In the union's phrasing, conditioning approval is an ordinary exercise of that authority and is routinely how courts in this District reconcile a value-maximizing sale with the legitimate concerns of parties who are not at the negotiating table.

The union adds that these conditions would not reduce the sale price, reopen the auction, or unsettle the winning bid — because "the Deidentification has not yet occurred; the protocol remains to be designed; and the Buyer already bears its cost without any reduction in the Purchase Price." It also argues from irreversibility: once decades of employment records have been delivered, no later order can meaningfully claw them back, which makes a measured safeguard now better than litigation afterward.

5

Is $10 Million Expensive?

Until now there was no way to say whether $10 million is expensive or cheap, because no company had ever priced its own working records in public. This auction leaves the first reference point. There is more than one way to read it, and the answer changes with what you hold it against.

The simplest method is division. Be careful with the denominator. Add up every count on the schedule and you clear 20 billion records, a total dominated by high-volume operational telemetry like the 7.25 billion fare observations and 3.53 billion booking-curve points. Put machine-emitted observation logs and human-written emails in the same denominator and the unit price stops meaning anything. Dividing only by human-made units is the honest move.

Bids converted to unit prices. Micro1's $12.5M was offered after bidding closed.
BidderPricePer emailPer mailboxPer employee record
Google (winner)$10,000,000$0.10$125.00$56.93
Micro1 (post-deadline)$12,500,000$0.125$156.25$71.16
Mercor (backup)$7,500,000$0.075$93.75$42.70

Ten cents an email. A hundred and twenty-five dollars a mailbox. Whether that reads as large or small is not yet answerable. Placed next to the price tags already known in the data licensing market, it looks like this.

Comparison set. The reported ~$70M/year Reddit–OpenAI arrangement is an estimate and is left out of this table.
BenchmarkAmountSpirit's $10M equals
Reddit's data license with Google~$60M/yearAbout two months of that deal
Reddit's quarterly data licensing revenue
(Q1 2026 disclosure)
$36M27.8% of one quarter's licensing revenue
Reddit's advertising revenue, same quarter$549M1.8%
The New York Times' annual cost of producing journalism
(stated by its CEO on a podcast)
~$2B/year0.5%
Converted at $85/hour for expert data work117,647 hoursAbout 56.6 person-years
Converted at $200/hour for senior work50,000 hoursAbout 24.0 person-years

The last two rows turn the sum into human time. Everything one airline produced over two decades of work costs the same as somewhere between 24 and 57 years of a person working as an expert labeler to earn that sum. One to two working lifetimes. For that, an entire company's operating record changes hands.

How big — or small — is $10 million? A log-scale comparison Google's winning bid — this deal $10M Reddit–Google annual data license ~$60M Reddit's quarterly ad revenue $549M The New York Times' annual journalism budget ~$2B The horizontal axis is logarithmic — the real gap is far larger than the bar lengths suggest.
▲ Pebblous original diagram — $10M compared to other data deals and budgets on a log scale (source: §5 comparison table)

5.1The Same Sum Weighed Differently on Each Bidder

Look from the bidders' side and another angle opens. Both challengers supply human-generated data to AI developers, and Pebblous covered both of them in July in our report on the expert data labor market. Hold the revenue figures from that report against these bids and the felt weight is nothing alike.

Mercor's annualized revenue was around $2 billion as of June 2026. Its $7.5 million backup bid is 0.375 percent of annual revenue — about 1.4 days of sales. Rounding error, near enough. Micro1 sits in a band climbing from $125 million to $300 million annualized. Its post-deadline $12.5 million is 10 percent of annual revenue at the low end, and still 4.2 percent at the high end. That is a bet that leaves a mark on a quarter. The gap between the two bids on the same asset is $5 million; the gap in what that money means to each company is far wider.

5.2What Did the Buyer Actually Buy?

Does training agents on corporate operating records actually improve them? The academic ground under that question is thin. The closest literature is a 2024 benchmark study that evaluated LLM agents on tasks simulating a real workplace. The authors observed that agents were notably weaker at administrative and financial work than at software development, and explained it this way.

"We hypothesize that part of the reason lies in the fact that current LLM development is heavily based on software engineering abilities, such as coding, due to several high profile benchmarks that measure this capability (e.g. HumanEval, SWE-Bench) as well as the abundance of publicly available training data related to software. On the other hand, administrative and financial tasks, are usually private data within companies, not readily available for training LLMs."

TheAgentCompany, arXiv:2412.14161 (NeurIPS 2025 Datasets & Benchmarks Track), body text

The verb matters: the authors hypothesize, they do not report having confirmed. That sentence is an interpretation of an observation, not an experimental result. We found no controlled study showing that training agents on internal corporate logs improved performance. What exists is an observation that agents are weak where public data is missing, and a hypothesis about why. So there is no basis for calling what Google bought a verified performance gain. What it bought is a kind of corpus nobody else has. If the hypothesis holds, that corpus fills exactly the gap it names. If it does not, this becomes a $10 million experiment.

The buyer itself has said almost nothing. One sentence, relayed by Forbes citing Bloomberg Law — that the data will improve its products and AI models — is the whole of it. The agreement filed with the court contains no clause requiring the buyer to disclose what it will use the assets for, and no clause limiting that use. The reason the union separately asked the court to prohibit profiling, evaluation and scoring of individual flight attendants or identifiable groups is that the agreement as written contains no such restriction.

5.3A Benchmark Set by a Distressed Seller

The argument that broke out on social media after the result went to price. Privacy came second. Trent Krupp wrote: "This was actually a steal I think. Street value for this might be 3-5x what GOOG payed." Which raises the obvious question of why no other large lab bid $30 million. The answer offered was resale: "You sell the data multiple times. Mercor should have bid over $10m to win it, and then sell it 3-5 times to frontier labs." Matt Mickiewicz summed the whole exchange up in four words: "new asset class unlocked."

There is an inversion buried in that conversation. The party that handed the market its benchmark was a seller in no position to negotiate the value of its own asset. A §363 sale in Chapter 11 moves assets free and clear of liens and claims, on a schedule measured in days, with a three-day objection window. There is no time to bid the price up and no option not to sell. The number that came out of that process is now the only public comparable any organization can point to. Companies pricing their own working records from here will start not from what a seller with leverage received, but from what a seller in the act of shutting down received.

This does not stop with one case. In the Chapter 11 of the genetic testing company 23andMe, the California Attorney General moved against the sale of residents' genetic information, and a £2.31 million penalty imposed by the UK regulator was stayed pending the bankruptcy outcome. Hughes Satellite Systems filed for Chapter 11 on August 2, 2026. Spirit's structure runs the opposite way from these: consumer records come out, and what a company left behind goes in. The mechanism, though, is the same. PPC Land put it exactly: the distinction between a consumer database and a corporate one is a drafting choice made by the seller, not a statutory boundary.

In June, Pebblous mapped the shift in AI data licensing from one-off dumps to real-time access. This deal runs against that current. With no entity left to sell real-time access, all that remains is a closed static dump that will never grow again. A dead company's data has no updates, and without updates there is no subscription. You sell it all at once or not at all.

6

Enron Left a Door Open; This Contract Closed It

This is the second time a collapsed company's internal mail has become research data. The first was Enron. On March 26, 2003, the Federal Energy Regulatory Commission released more than 1.6 million messages exchanged by some 150 Enron executives between 2000 and 2002, as a byproduct of a regulatory proceeding. The trouble was that the files were posted in an unusable format. Leslie Kaelbling at MIT bought a workable copy from a government contractor for $10,000 and cleaned it up for research; the version Carnegie Mellon distributes covers about 150 people and roughly 500,000 messages. For the twenty-odd years since, that has been essentially the only public corpus of real email.

It is tempting to line up $10,000 against $10 million. Check one thing first. The $10,000 was the price of a usable copy of files that had already been made public, bought from a contractor — not the price of data rights bought from the Enron estate. The two transactions are different in kind, so reading "a thousandfold increase" as the same sort of price movement gets it wrong. The comparison holds only this far: in 2003 an asset of this character could be had for $10,000; in 2026 it cost $10 million.

The sharper contrast is not about price. In the Enron corpus, employee deletion requests were actually honored. Carnegie Mellon's distribution page notes that some messages were deleted "as part of a redaction effort due to requests from affected employees." FERC itself removed Social Security numbers and banking records after receiving complaints. The people in the data had a door to knock on, and when they knocked, it opened.

The Spirit agreement designed that door out. Section 3(e) bars the debtors from deleting anything; Section 12 denies flight attendants third-party beneficiary status. There is no one inside the agreement to receive a request, and no one recognized as entitled to make one. Material released in 2003 as the byproduct of a regulatory proceeding is more open to the people in it than material sold by contract in 2026.

Carnegie Mellon's distribution page leaves one sentence for researchers who use the dataset.

"In using this dataset, please be sensitive to the privacy of the people involved (and remember that many of these people were certainly not involved in any of the actions which precipitated the investigation.)"

Carnegie Mellon University, Enron Email Dataset distribution page

Spirit's 4,600 flight attendants and 175,658 employee records did not bankrupt the company either. What they left behind is timecards, training completions, and Teams conversations. Where those records will sit twenty years from now, and in whose model weights, they have no way to know — and no door to ask through.

7

Are Your Own Working Records in Sellable Shape?

Everything up to here is another company's story. The question this case puts to yours is less comfortable. Could your organization produce, today, the one-page table Spirit assembled in a few weeks inside a bankruptcy proceeding? If so, how many rows would it have, and could you fill in the count and start month for each one? Failing that question means something before it means you are unprepared to sell: it means you do not know what you have or how much of it.

7.1What Your Retention Policy Is Actually Holding, and for How Long

Spirit's email extraction window runs eight years, 2018 through 2026. That is not the product of some unusual policy. Microsoft 365 moves mail to archive automatically after two years unless a policy says otherwise, and the retention period most commonly adopted for compliance purposes is seven years. Spirit's eight years sits just past that ordinary window. Which is to say: simply following a standard retention policy accumulates a corpus this size. The 100 million emails on the schedule are not the result of a company deliberately collecting something. They are eight years of a decision not to delete — or of no policy at all.

A common misreading is worth heading off here. Divide 100 million emails by 80,000 accounts and eight years and you get 156 messages per account per year, or 0.43 a day. Do not use that as a benchmark for your own company. The reason is not the obvious one. It sits an order of magnitude below surveys reporting that a working account handles low hundreds of messages a day and tens of thousands a year. But those surveys count all work email, outbound and inbound, marketing and spam included. Move to a benchmark that isolates internal mail and the picture flips. A 2026 study drawing on 2 billion internal corporate emails across roughly 11 million employees found that employees receive 14 internal messages per month; Spirit's 13 per account per month lands almost exactly on it. Whether Spirit's number is unusually low depends entirely on what you compare it to.

Which leaves one conclusion. The number on the schedule is not mailbox traffic; it is the output of an extraction scope defined by the sale agreement. Whether thread deduplication, a folder restriction, or an internal-only filter drove it is nowhere in the documents. That indeterminacy is itself the practitioner's point. A count printed on a schedule is not the whole ledger. It is a contractually defined subset, and the definition of that subset does not appear in the table. The same question follows you when you build your own inventory: what does that number count, and what did it leave out?

7.2The Contractual Status of Employee Communications

The second question is about what your employment agreements and workplace policies actually cover. Do those documents contemplate that communications employees and contractors leave in company tools may outlive the company itself? Nearly every employer has monitoring language for the period of employment. Language governing what happens when those records pass to a third party as an asset after the company is gone is rare.

It is also worth seeing that data flows outward whether or not the company sells anything. Privacy firm Incogni examined ten workplace apps in March 2026 using their Google Play Store disclosures and found that the average app collects 19 data types and shares about 2 of them with third parties. Gmail collected the most at 26 types, followed by Teams at 25 and Zoom Workplace at 23. Among those studied, Workday was the only app that does not support data deletion requests. Before you get to the question of a company's data assets, there is already a layer that tool vendors are harvesting.

7.3What If This Happened Under Korean Law?

The gap in the U.S. Bankruptcy Code comes from defining personally identifiable information around the consumer. Korea's Personal Information Protection Act (PIPA) draws that line elsewhere — a useful comparison, since it is one of the stricter regimes any multinational has to satisfy. Three points stand out.

First, the transfer rule is not narrowed to consumers. Article 27 governs the transfer of personal information arising from a full or partial business transfer, merger, or division, and does not limit data subjects to consumers. Both transferor and transferee must notify data subjects, and the notice must include what steps a data subject can take, and how, if they do not want their information transferred. Where individual notice is impossible without fault, substitute methods such as a website posting for at least 30 days are permitted — a provision that becomes practically decisive when the company has already shut its doors. The point: were the same asset transfer to happen in Korea, employees would be among those entitled to notice.

Second, pseudonymized data is bound to specific purposes. Under Articles 28-2 and 28-3, pseudonymized information may be used without the data subject's consent only for statistical compilation, scientific research, archiving in the public interest, and similar purposes, and combining pseudonymized data held by different controllers is permitted only through an expert institution designated by the Personal Information Protection Commission. A Spirit-style deal — selling a pseudonymized internal corpus to an AI company — would first have to clear the question of whether training a model counts as scientific research. The U.S. agreement has no such gate at all.

Third, it collides head-on with the destruction duty. Article 21 requires personal information to be destroyed without delay once the retention period expires or the processing purpose is achieved. Winding up and bankruptcy are the textbook cases of a purpose disappearing. We found no regulatory interpretation or case law reconciling that duty with a demand to sell the data as an asset.

This section has places we could not fill. Within the scope of this research we found no Korean precedent for data being sold as a standalone asset in rehabilitation or bankruptcy proceedings. We cannot conclude there is none — only that we did not find one. Whether a receiver or trustee holds the status of a personal information controller, and how the succession of employment relationships interacts with the transfer of personnel records, are likewise unresolved. And that vacancy may be this section's conclusion. If a Korean company lands in the same position, the rules that would answer it are not settled yet.

7.4Four Questions You Can Answer Now

There are questions you can prepare answers to while the rules are still unsettled. Map Spirit's schedule onto your own organization and these are what remain.

  • Inventory. Can you produce one page listing system name, item name, row count, and start month? If not, is the obstacle tooling, access rights, or people?
  • Retention. For each system, what is the retention policy holding and for how long — and is that a value someone set, or a default?
  • Contractual status. Do your employment agreements and workplace policies say what status employee and contractor communications have independent of the company's survival?
  • Who the deidentification protects. When you design pseudonymization or deidentification, does the documentation state who it was designed to protect? Bolting a consumer-grade standard onto employee data reproduces exactly the gap Spirit has.

The fourth item comes with a trade-off attached. Cut referential integrity and privacy risk falls — so does the analytical value of the data. They are two ends of the same lever, and where you set it is not something technology decides. Policy does. In the Spirit agreement, the hand on that lever belonged to the company buying the data.

8

Why Pebblous Is Watching

What this case put into the world is not only a price tag. It is a schedule — one page carrying system names, row counts, start dates, include/exclude determinations, and extraction formats. In form, it is an enterprise data asset inventory produced under pressure by a bankruptcy process, which is the same artifact an AI-Ready Data assessment sets out to produce. The only difference is that Spirit built it after the company died. Which makes the question for readers simple: can you build that page while you are still alive?

From a data quality standpoint, the center of this case is the proviso attached to §3(c). The condition imposed on the Deidentification Agent was "while preserving referential integrity." Live joins are what make the material meaningful as training data, and that identical property is the one the union is worried about. The core attribute of data quality and the core attribute of privacy risk overlap inside a single clause. If you accept that the structure of training data carries into a model's internal representations, what sold here is not a pile of email. It is the structure of how one company did its work for twenty years.

Most companies cannot say what is sitting in their own Microsoft 365 tenant, or how much of it. This case shows where that ignorance turns into cost. A company without a list of its own data cannot negotiate over it, protect it, or price it. Spirit's schedule exists because the buyer asked for it — which is why the determination column is headed "Google's Data Purchase Request." Whoever holds the list first writes the terms.

And the list carried a price but no quality statement. The buyer received counts and no warranty whatsoever as to accuracy or completeness, and $10 million was set against that unverified state. What creates leverage when data gets priced is, in the end, the ability to state the condition of your own data. Only a company that can put in writing what it has, how much, how far it can be trusted, and where the holes are gets to write terms at that table.

Pebblous is not in the business of selling data. We are in the business of getting data into a usable state. This case is the first public instance of that state being converted into an actual asset value — and simultaneously the case in which a union's court filing marked where the conversion breaks down. The place where the price attaches and the place where accountability is missing sat in the same document.

Editor's Note. This report is based on court filings public as of publication (the AFA-CWA objection, Doc 1489), PPC Land's reporting quoting the sale agreement clauses, U.S. Securities and Exchange Commission disclosures, and the primary text of each academic paper cited. The sale hearing is scheduled for September 9, 2026; no ruling had issued as of publication, and the deidentification the agreement calls for has not been performed. We were not able to review the full docket, so whether anyone other than the flight attendants' union objected is unverified, and the description of the European Data Protection Board guidelines records only the direction, based on reporting rather than the source text. Among the academic sources, the HIPAA Safe Harbor critique is a position paper concerning clinical records, and the statement about private corporate data in the agent benchmark study is the authors' own hypothesis.

R

References

Court and Regulatory Documents

  • 1.Association of Flight Attendants-CWA, AFL-CIO. "Limited Objection to the Proposed Sale of the Deidentified Data." In re Spirit Aviation Holdings, Inc., Case No. 25-11897 (SHL), Doc 1489, S.D.N.Y. Bankr., 2026-08-18. — Primary text of the union's argument (¶¶2, 3, 10–15, 17–18). The basis for all of Section 4
  • 2."Notice of Auction Results and Scheduled Hearing for the Deidentified Data." Same case, ECF No. 1463, 2026-08-14. — Auction result and asset schedule. Not read directly; clauses quoted via PPC Land's reporting and the AFA filing
  • 3.U.S. SEC EDGAR. Spirit Aviation Holdings, Inc. (CIK 1498710), Form 8-K Exhibit 99.1, Monthly Operating Report. — Primary disclosure of financial condition during the bankruptcy
  • 4.Personal Information Protection Act of Korea, Articles 21, 27, 28-2 and 28-3 (Korea Law Information Center). — Basis for the Korean-law comparison in Section 7.3
  • 5.European Data Protection Board. "Guidelines 02/2026 on Anonymisation," adopted 2026-07-07. — Source text not obtained. Section 3.3 records direction only, based on PPC Land's reporting

Academic

Reporting and Research

  • 10.Rijo, Luis. "Google wins bankrupt Spirit Airlines data for $10 million." PPC Land, 2026-08-29. — Reporting that quotes the agreement clauses and every row of the asset schedule. The main basis for Sections 2 and 3
  • 11."The Immortal Life of the Enron E-mails." MIT Technology Review, 2013-07-02. — How the $10,000 purchase came about, and the regulator's later redactions
  • 12.Incogni Research Lab. "Workplace apps are watching, keeping tabs, and sharing what they learn." 2026 (data collected 2026-03-20). — Workplace apps collect an average of 19 data types
  • 13.Maharishi, Meghna. "Why Flight Attendants Are Fighting Google's Purchase of Spirit Airlines Data." Skift, 2026-08-27. — Only the lede and summary were accessible behind the paywall. Used to confirm that deidentification has not been performed
  • 14.Carter, Sandy. "AI Data Wars Begin As Google, Mercor And Micro1 Bid For Spirit's Data." Forbes, 2026-08-23. — Micro1's post-deadline bid and the issues before the September 9 hearing; the buyer-side statement (quoted from Bloomberg Law). ⚠️ The same article's summary that the union "fears potential re-identification" conflicts with ¶13 of the objection, and this report does not adopt it
  • 15.PoliteMail. "2026 Internal Email Benchmark Report" (sample: 2 billion internal corporate emails, ~11 million employees). — The internal-only email benchmark in Section 7.1 (14 messages received per employee per month)

Related Pebblous Reports

🔗 Related reading — the people who wanted to buy this data

Who Mercor and Micro1 are, and what they pay for expert data, is covered in The Expert Data Labor Market. The broader move from one-off dumps to real-time access is mapped in our AI data licensing report — and this deal is the exception that runs backward against it.