If you’re based in New York, you can now get your apartment cleaned for free. The cleaner shows up wearing a cap-mounted camera that captures everything: how they load your dishwasher, which drawer holds your knives, how to use your fancy coffee machine. Every wipe and fold becomes training data, sold to labs to train household robots. And it doesn't stop at chores…last week the same company started offering free private chefs. Doctors may be next - we've spent time with a company applying the model to healthcare, where the visit costs nothing and everything the clinician does - the questions, the tests ordered, the reasoning in between - trains bioscience models.

As demand for training data accelerates, we expect frontier labs to increasingly subsidize more and more of everyday life. Every human task, activity, and interaction becomes a dataset someone will pay to capture.

As models improve, so do the demand, complexity, and cost of data

Three years ago, the industry-wide assumption was that as models get smarter, they'll need less data to improve. In reality, the opposite is happening. Frontier labs are spending an estimated $10-15B per lab on data, in a market that’s grown from $3.8B in 2024. It’s commonly referred to as the J-curve of training data: early on, the market looks like a commodity racing to the bottom on price (LLM data in 2024, egocentric robotics data right now). Then it hits an inflection point. Models master the generalist work, and the only data that still moves the frontier is more complex, higher quality, and harder to source. As task complexity increases, so does the price labs will pay.

Take Mechanize, for example, arguably the most in-demand RL coding data provider right now. Rather than following the Mercor/Scale contributor model, they are hiring full-time engineers on $250-300k salaries plus meaningful equity. The result? Allegedly, labs are paying Mechanize $10-20k for the same long-horizon coding task they’d pay Mercor closer to $2k, because full-time engineers produce higher quality data than gig workers. It's the same reason non-technical employees at human data companies (SPLs) are getting poached with $1M signing bonuses two years out of undergrad. Any marginal model improvement can drive billions in incremental revenue for the labs.The money is not the constraint. The constraint is that the most valuable data has yet to be extracted.

The most valuable data is the hardest to reach

Now that the public internet is scraped, the most valuable data is whatever is hardest to collect. This is why Meta started recording employees' keystrokes and mouse movements to train the company's own coding models, with an estimated 6,500 engineers effectively doubling as training-data creators. Meta sits on data no other lab can access: thousands of employees doing best-in-class work inside internal systems, every day. To Zuck, it's an obvious move: use the data you already sit on to give your model an advantage.

In most cases, the valuable data sits far outside labs’ reach. Healthcare is one of the clearest examples. End-to-end clinical data is among the most valuable datasets in AI and among the hardest to touch. HIPAA requires de-identification before patient data can be used for training, and the records themselves sit fragmented across hospital EHR systems. A market is already forming to get it out: Truveta aggregates de-identified records from 120 million patients across 30+ health systems, and companies are now paying hospitals directly to hand over data for AI training. The players who can deliver clinical data end to end become increasingly interesting to the labs.

A brain-computer interface (BCI) goes one step further. Invasive BCI companies like Neuralink and EEG wearable companies like Sabi and Alljoined are building instruments to capture currently nonexistent datasets. Just by reading the electrical activity off your scalp, they can begin to decode what you're seeing, thinking, and feeling in real time. This dataset - a live readout of what's happening inside your head - is completely unprecedented, because no device has ever produced it. It gets even more interesting when you pair the brain signal with a camera or smart glasses: you capture what someone sees alongside how their brain reacts, so you know not just what they're looking at, but what they think and feel as they look at it. Neither of these is for sale on the open market, which makes it more attractive to the labs. Though it’s not yet clear whether data this far out on the frontier actually moves model performance.

What this means for the consumer business model

For twenty years there was one dominant business model behind free consumer products: ads. Get users, sell attention. That loop built Google and Facebook. Today, monetizing data collection presents a real alternative to a free consumer product built on ads. Ads sell consumers’ attention; data collection sells tacit knowledge and skills. Imagine a cellular plan that's free because your carrier sells your anonymized voice and messaging data to labs. It sounds far-fetched until you remember you've been trading attention for free social media since 2004. Shift will not be the last company to fund a free product this way. The flywheel: free product → more usage → more and better data → labs pay for it → the product stays free.

And it’s repricing the gig economy

The same trade is reshaping the gig economy too, expanding what it means to be a gig worker while changing which jobs are actually worth doing. Sunday Robotics put skill-capture gloves in hundreds of homes, paying everyday consumers (they call them "Memory Developers") to do their normal chores, folding laundry, loading the dishwasher, while the gloves record every touch, grip, and fine motor movement. DoorDash launched a standalone Tasks app that pays couriers to film everyday activities for AI training in between deliveries; one example task is filming your own hands washing five dishes while wearing a body camera. The median dasher earns $11.63/hour all-in doing deliveries, while complex video tasks pay $10-25 each and take a fraction of that time. Kled built a consumer platform where anyone can sell the data they already produce: camera roll photos go for about $1 per 100, food shots sell for $5 each, and one frontier lab offered $1,000 for a set of 20 selfies.

Which raises the philosophical question on which the whole model rests: how willing are we, humans, to accept upfront cash to train systems that eventually displace us?

How long do the labs keep paying?

Through the 2010s, venture dollars subsidized everyday life. Uber ran up more than $30B in cumulative losses keeping fares below the true cost of a ride. Then the free money ended, and consumers ate the difference: rideshare prices rose 92% between 2018 and 2021.

Today the labs are running the same playbook across a much bigger surface area, and there are three ways it ends: (1) demand for data plateaus, which likely means we’ve reached major research breakthroughs (i.e. recursive self-improvement or Flapping Airplanes on data efficiency), (2) the labs' own economics force the engine to stop, or (3) enterprises’ lab spend decreases dramatically as training proprietary in-house models becomes more cost effective and performant.

Take the second option. In 2026, both Anthropic and OpenAI filed for an IPO. Anthropic hit a $47B revenue run rate and reportedly turned a $1B+ quarterly profit, while reports put OpenAI around $25B in revenue while losing $21B a year. When these companies start trading on the public market, data spend stops being a strategic asset and becomes a quarterly line item. Unlike compute (which reads as infrastructure, like capex), human data is pure opex that has to re-justify itself every earnings call. Today, a handful of companies' cost lines ripple through the entire economy.

The third scenario might actually be the base case. A multibillion-dollar enterprise with decades of proprietary IP can now take a cheap open-source model, train it in-house on its own data, and keep both the data and the returns, instead of paying frontier-model prices for capabilities that were never built for their workflows or data. Every enterprise that makes that switch shrinks the revenue that funds the labs' data budgets. That said, for this to make a meaningful dent, enterprises would need to move a majority of their AI usage in-house, which is hard to imagine today.

So how long does this last? The demand signals say we're early. RL environments are just being built, robotics data collection has barely begun, and the labs' data budgets keep growing fast enough to double Mercor's valuation in nine months. But a decade of demand and a decade of subsidy are not the same thing, and the gap between those two is where this space gets repriced.

What survives when the subsidy ends?

The open questions we keep coming back to at Torch:

How long do labs keep funding everything?

What breaks first, demand or their balance sheets?

Who's sitting on a dataset they don't fully value? Who owns a source that labs structurally can't reach?

Will businesses that are purely data arbitrage - collection schemes with no real product underneath - survive an AI winter?

Will the winners be the companies that used this window to build something people actually want, where data is the second business model rather than the only one?

If you're building here, or you have a strong opinion on any of this, we’d love to hear from you.