Executive Summary

One model now carries two prices. Muse Spark 1.3, which Meta published on September 2, is sold under a standard rate and a contributor rate. Both rates point at the same model with the same context window. The one thing that separates them is a data handling condition. Let Meta use your prompts and the model's responses to train its next model, and the price drops.

It drops by a lot. A million output tokens costs $4.25 under the standard agreement and $0.20 under the contributor agreement. TechCrunch put the average discount across the pricing model at about 95%. The discount is not uniform, though. The steepest cut lands on cache reads rather than on output, which is what the coverage led with.

The gap matters more than the discount rate. Multiply the difference between the two rates by your team's monthly token volume and you get the price Meta has attached to your prompts. How far that number can be trusted, and what an organization trying to book data as an asset can actually do with it, are the questions that follow.

Key Numbers

Sources: TechCrunch (2026-09-03), public price listings for Muse Spark 1.3 and Muse Spark 1.3 Contributor (checked 2026-09-05)

21.25x

Discount on output tokens

$4.25 per million becomes $0.20. The 21x figure in the coverage comes from this tier

75x

Discount on cache reads

$0.15 per million becomes $0.002, the steepest cut of the three tiers

~95%

Average discount per TechCrunch

A single figure covering unequal tiers, so a real invoice moves with usage mix

$980/month

Gap for a mid-sized team

Pebblous estimate assuming 500M input and 100M output tokens monthly, or $1,050 against $70

1

The Same Model Now Carries Two Prices

Most AI tools let you opt out of sharing your usage with the model provider. Usually that is one checkbox and it costs nothing either way. Meta has put a price on the checkbox. Muse Spark 1.3, built for coding and other agent workloads, is listed at a standard rate and a contributor rate, and choosing the second one sends your prompts and the model's outputs into the training set for Meta's next model.

Tier by tier, the published listings read as follows. Input falls from $1.25 to $0.10 per million tokens. Output falls from $4.25 to $0.20. Cache reads, which cover re-reading material the model has already processed, fall from $0.15 to $0.002. The context window is 1,048,576 tokens on both rates, and both accept the same input types: text, images, video, files, and audio. Nothing was trimmed from the product to make the cheaper number possible.

As multiples, input gets 12.5 times cheaper, output 21.25 times, and cache reads 75 times. The "21x" that ran through the coverage is the output figure. Line the three up and the biggest cut sits somewhere else.

The steepest discount was not on output tokens Contributor pricing as a multiple below standard pricing, Muse Spark 1.3 Input $1.25 → $0.10 12.5x Output $4.25 → $0.20 21.25x Cache read $0.15 → $0.002 75x The 21x cited in the coverage is the output tier Web search calls cost $2.50 per 1,000 on both rates. The discount lands only where data changes hands
▲ Pebblous original diagram | Source: public price listings for Muse Spark 1.3 and Muse Spark 1.3 Contributor (checked 2026-09-05)

Not every line item was discounted. Web search calls cost $2.50 per 1,000 on both rates. The price only moved where tokens change hands, which is to say where your prompts and the model's responses actually travel. The listing itself tells you what Meta is buying.

Meta's own pricing guide describes the tier this way: it "lowers the barrier to entry for prototyping, testing integrations, and scaling experiments where training on your data is acceptable." The conditional sits in the middle of the sentence. The judgment about whether your data is fit to train on is handed to the customer, and that hand-off is written into the price description.

The tier description on the price listing draws the permission wider than the coverage did. It says prompts and outputs "may be used to improve Meta's products," which is broader than the development of future models that the reporting described. The same listing recommends the tier for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. The standard rate is described as being for long-running ones. The context window, the input types, and function calling are identical on both, so the product is the same and only the suggested use is smaller. Anyone weighing what this contract gives away should start by comparing those two sentences.

2

Why Meta Had to Put Money on the Table

Meta has had a rough time obtaining training data lately. An internal initiative to track employees' computer usage and feed it into training, launched earlier this year, drew wide internal criticism and was paused in June. When TechCrunch asked about the new pricing model, the company did not respond. No sentence in the article ties the two events together as cause and effect. The sequence remains, though: a route closed inside the company, and then a route that pays for the same material opened outside it.

The external route appeared where the internal one stopped Meta's paths to training data, in order Internal route Employee PC tracking plan Early 2026 Paused after backlash June External route 1.2 Contributor ships Aug 21 Standardized as 1.3 Sep 2 The reporting never calls these cause and effect This diagram lists only the sequence — it does not claim the paused plan produced the paid tier
▲ Pebblous original diagram | Source: TechCrunch reporting (2026-09-03), public OpenRouter price history (checked 2026-09-05)

The article also explains why agent logs in particular. Mario Zechner, the developer behind the open source harness Pi, told TechCrunch last month: "The reason we saw a big jump in [coding agent] capabilities between April 2025 and October 2025 was that Claude Code, by default, would store all your coding agent sessions and use them for reinforcement learning training." What grew the models was not more web text but real working sessions.

That supply runs out the moment agents are pushed beyond software engineering. Most professional work is complicated and leaves almost no digital trace. An accounting review, an equipment inspection, and a contract negotiation do not end up in a log, so there is nothing to evaluate the tool against and nothing to improve it with. That is the reason to buy the traces even at a cost.

The price was already low before the discount. Mark Zuckerberg told Bloomberg on July 9 that Muse Spark 1.1, released that month, was the first time Meta charged businesses for access to one of its models. "Since this is not an open source model, this is I think the first time that we're doing a real serious API," he said, adding that "the pricing is going to be very aggressive and attractive." Bloomberg put that API pricing at around 25% of what OpenAI and Anthropic advertise. Those standard rates have not moved since. The contributor rate is a second discount stacked on top of them, so pulling in developers with a low price and paying for data with a lower one now sit as two layers of the same listing.

The competitive picture points the same way. Anthropic's Fable and Mythos models, released the same day as Muse Spark 1.3, came with lower costs for processing cached tokens, and OpenAI's latest models took major price cuts at the end of July. Falling prices are the industry norm. What Meta did differently was write the consideration for the cut into the price list as a line of its own.

3

Training Data Used to Be Priced After the Fact

Until now, a number got attached to training data mostly after something went wrong. Copyright settlements, court-ordered accountings, datasets handed to creditors in bankruptcy proceedings. A price was set, but always after the fact and usually by a third party rather than the owner. Sony Music and Warner Chappell asking a court to order an accounting of Claude's training data belongs to that lineage, and we covered that complaint here.

It is not that no price existed. The price existed and simply was not written down anywhere. Arvind Narayanan, a computer science professor at Princeton, put his finger on exactly this point, noting there is good evidence that large companies do not want their data used for model training: "They stick with token-billed Enterprise plans even though the subscription-based consumer plans like Claude Max and ChatGPT Pro are discounted by 10x-20x or even more! (The main difference between the plans is data retention + enterprise IT governance)"

If those companies were paying 10 to 20 times more, they were already paying for their own data. The payment just never appeared as a line item on an invoice. It was dissolved into contract terms and plan selection, which made it impossible to book and impossible to compare against anyone else. This pricing model changed how the price is recorded, not whether it exists. Two lines sitting next to each other under the same model and the same performance turn the gap into a number you can see.

So how long have those two lines been there? The contributor rate did not appear with this model. Walk back through the public listings and the same numbers are already on Muse Spark 1.2 Contributor, published on August 21: $0.10 for input, $0.20 for output, $0.002 for cache reads, not a digit different from today. The standard rate has held even longer. It has been $1.25 and $4.25 since version 1.1 on July 16, across three model releases. There was also a time without the second line at all, because 1.1 shipped without a contributor tier.

The contributor rate hasn't moved since August 21 Muse Spark price history by version 1.1 · Jul 16 1.2 Contributor · Aug 21 1.3 · Sep 2 Standard $1.25 input · $4.25 output — unchanged across three versions Contributor No tier $0.10 input · $0.20 output — same as Aug 21 What changed twice was the model, not the price Not a launch-day number but one that held through a third release
▲ Pebblous original diagram | Source: public OpenRouter price history — Muse Spark 1.1 (2026-07-16), 1.2 Contributor (2026-08-21), 1.3 (2026-09-02) (checked 2026-09-05)

A launch-day number designed to draw attention would be hard to measure anything against. A number that stayed put while the model changed twice is a different object. Anyone moving data onto a balance sheet needs a price that holds still more than a dramatic discount.

So calling this listing the first price tag on training data goes too far. The price was there all along and companies were already living with it. What is new is that the price is stated in advance, in public, with every other condition held constant. That gives us one comparable number, and comparability is where asset valuation starts.

4

What Did Meta Price Your Prompts At?

The arithmetic is one multiplication. Take the per-tier price gap and multiply it by your team's monthly token volume. What comes out is the money you pay each month to keep your data to yourself, and read the other way, the revenue Meta gives up each month to get your prompts.

(standard rate − contributor rate) × monthly token volume = the monthly price on your prompts

Assume a team that spends 500 million input tokens and 100 million output tokens a month, which a mid-sized organization running agents over its codebase can reach.

Estimate for a team using 500M input and 100M output tokens per month

  • 500M input tokens: $625 standard, $50 contributor. Gap of $575
  • 100M output tokens: $425 standard, $20 contributor. Gap of $405
  • Total: $1,050 standard, $70 contributor. Gap of $980

The usage assumption is ours; the rates come from public listings checked on September 5, 2026. Cache reads and web search calls are excluded.

Annualized, that is $11,760. It is what this team pays over a year to keep its prompts, and equally what Meta has priced one team's working records at. Down at the level of a single call it gets more concrete. One agent turn with 50,000 input tokens and 2,000 output tokens runs about $0.07 on the standard rate and about $0.005 on the contributor rate.

Three caveats govern how far you can trust the number. First, one buyer set it unilaterally. It is not what your data is worth; it is what Meta is willing to pay at this moment. Second, the price never looks at the content. A prompt from a team handling regulated work and a prompt summarizing an internal wiki are priced identically if the token counts match. Third, ownership does not change hands; you are granting permission to train. Your data stays with you after you hand it over, but once it is inside the next model there is no taking it back.

The value of the multiplication is not that it settles a price. It puts a comparable floor under a discussion that has been running on instinct in meeting rooms. Whether your prompts are worth more than $980 a month is still your call, but at least there is now something numeric to make the call about.

5

For Teams Trying to Book Data as an Asset

What changes inside a company once the price is visible? Narayanan suggested it could push large companies to be more diligent about which data is truly proprietary and which could be shared with model providers. A two-box decision about keeping everything or giving everything away turns into a decision made item by item.

Splitting the decision item by item starts with an inventory, a written record of what your organization is currently sending to models. Without one there is no way to separate the prompts that carry proprietary knowledge from the ones that could safely be shared, and without that separation the choice of pricing tier becomes a condition applied to the whole team at once. One team switching to the contributor rate to save on its bill can ship the organization's know-how along with it.

Three questions are worth writing down while this pricing is fresh.

  • Is anyone measuring how many prompts go to models, from which workflows, and at what volume? Without the second term, the multiplication never starts.
  • Is there a written standard separating what may leave the building from what may not? Without one, the choice of pricing tier becomes the data policy.
  • Does the saving from the price gap get compared against the value of the data given up, in the same meeting? If finance owns the rate and security owns the data, that comparison happens nowhere.

Editor's Note

The wall Pebblous runs into most often while diagnosing data quality is that there is no yardstick for putting a value on it. We can show that quality improved. What that improvement is worth is harder to say, because there is nothing to compare it against. This price list does not supply that yardstick. But a public number is now attached to data given up for training, and it has not moved since late August, which gives organizations trying to translate their data into money one axis to measure by.

Thanks for reading. The original reporting is at TechCrunch, and the rates were verified directly against the public price listings. If you run this multiplication against your own team's usage, tell us what number came out.

Pebblous Data Communication Team
September 5, 2026

R

References

Primary Sources — Public Pricing Pages

  • 1.OpenRouter. (2026). "Meta: Muse Spark 1.1." — First publication of the standard rate card ($1.25/M input, $4.25/M output). This version had no contributor tier (page returns 404).
  • 2.OpenRouter. (2026). "Meta: Muse Spark 1.2 Contributor." — First appearance of the contributor tier (Aug 21, 2026): $0.10/M input, $0.20/M output, $0.002/M cache read — unchanged through 1.3 Contributor.
  • 3.OpenRouter. (2026). "Meta: Muse Spark 1.3." — Standard rate card: $1.25/M input, $4.25/M output, unchanged since 1.1.
  • 4.OpenRouter. (2026). "Meta: Muse Spark 1.3 Contributor." — Contributor rate card. States "Prompts and outputs may be used to improve Meta's products."

Industry & Press