The race among AI labs to build smarter models has quietly created a gold rush in a less glamorous corner of the industry: the data those models actually learn from. Snorkel AI, a seven-year-old startup that builds training data sets and simulated environments for AI companies, has raised a $350 million Series E at a $3.5 billion valuation — nearly triple the $1.3 billion valuation it held just 17 months ago.
From data-labeling software to a full data supply chain
The round was led by Insight Partners and S32, with existing backers Addition, Lightspeed, Greylock, GV and Wells Fargo also participating. What’s notable is less the size of the check than how fast Snorkel’s business model has shifted underneath it. The company originally sold software that automated data labeling — the tedious work of tagging raw data so machine learning models can learn from it. Last year, it pivoted to something more ambitious: delivering fully completed data sets directly to customers, an approach it calls data-as-a-service. Rather than running as a pure marketplace connecting AI labs to human experts, Snorkel now blends its own software and models — which generate training data synthetically — with subject matter experts working alongside that automation.
That shift appears to be paying off in a big way. Snorkel says its annualized revenue run-rate now stands at $375 million, an 18-fold jump over the past 12 months — growth the company attributes directly to AI labs’ seemingly bottomless appetite for high-quality training data as frontier models get harder and more expensive to improve.
Snorkel isn’t alone in this boom
The training-data sector has become one of the least-visible, fastest-growing corners of the AI economy in 2026. Mercor’s gross annualized revenue has climbed to $2 billion. Handshake crossed the $1 billion mark earlier this year. And as reported earlier, Micro1 scaled to a $500 million gross run-rate. Taken together, it’s a clear signal that as large language models increasingly hit the limits of what’s learnable from the open internet, AI labs are willing to pay heavily for curated, expert-generated data instead.
One caveat worth flagging for anyone sizing up these companies: firms like Mercor, Handshake and Micro1 typically pay out roughly 60% to 70% of their gross revenue directly to the domain specialists doing the underlying work, meaning their real net revenue is substantially smaller than the headline figures suggest. Snorkel’s accounting works a bit differently — because it sells finished datasets and reinforcement learning environments rather than raw human labor, payments to its expert contributors are booked as cost of goods sold rather than folded into the top-line revenue number the company publicizes.
The company’s origins
Snorkel launched commercially in 2019, the product of four years of research led by co-founder and CEO Alex Ratner and his team at a Stanford University AI lab. That academic-to-commercial arc — and the seven years it’s taken to reach this valuation — stands out in a funding environment where plenty of AI startups are reaching billion-dollar valuations within a year or two of founding, a contrast that underscores just how much the underlying market for AI training data has accelerated recently, even for companies that started well before the current boom.
Why this matters
Snorkel’s tripling valuation is a useful proxy for a broader shift happening beneath the AI industry’s more visible headlines about chatbots and foundation models: as the easy, freely available internet data gets exhausted, the companies that can supply high-quality, expert-curated, and increasingly synthetic training data are becoming critical infrastructure for the entire sector — valuable enough that investors are now backing them at valuations that rival many consumer-facing AI products themselves.

No responses yet