On September 22, 2026, US-based AI data company Snorkel AI announced it had raised $350 million in a Series E round, pushing its valuation to $3.5 billion. Insight Partners and S32 co-led the round. The valuation jumped sharply compared to the company's previous funding round, but to understand this announcement, we need to look at what the company is actually selling to customers now. It has shifted from selling software that lets customers build their own training data to a business that deploys experts and AI to deliver finished data. It's worth separating the demand driving up the valuation from the quality of revenue that hasn't yet been made public.
What the $3.5 Billion Valuation Reflects
In May 2025, Snorkel AI announced a $100 million Series D round, putting its valuation at $1.3 billion. About 16 months later, the company raised $350 million and its valuation rose to $3.5 billion. From $1.3 billion in May 2025 to $3.5 billion in September 2026, Snorkel AI's announced valuation grew by roughly 2.69 times. This figure comes from dividing the two publicly disclosed dollar valuations—$3.5 billion divided by $1.3 billion. This is the price investors assigned at the time of funding; it does not mean revenue or profit grew by the same multiple.
The new funding will go toward expanding what the company calls its "production pipeline for agent data." According to the company's announcement, it plans to increase investment in vertical- and enterprise-focused AI, broadening the types of data and domains it handles. Some funds will also support research on public benchmarks. CEO Alex Ratner told Reuters that the funding will be used to hire researchers and engineers and to expand the company's enterprise and government business.
What the funding round demonstrates is that investors have put money behind the expansion of this business. How much customer contracts persist, and how much profit remains after paying experts and computing costs, are separate figures that still need to be verified.
From Customer-Operated Software to Finished Data Delivery
Snorkel's starting point was software designed to reduce the labor of building training data. In the traditional Snorkel Flow product, a customer's own staff used rules and models to label large volumes of data and shape it into a form usable for model training. Rather than manually inputting expert judgments one case at a time, the idea was to reuse rules that captured those judgments. It was a business of selling customers the tools to build their own data.
In May 2025, the company announced two new products: Snorkel Evaluate, designed to measure a customer's AI system, and Expert Data-as-a-Service, which provides finished data—produced with expert involvement—for evaluation and additional training. According to this latest press release, the company began full-scale operation of the latter in September of that year. This means the 2026 Series E round is not the unveiling of a new service. Rather, it represents funding raised after roughly a year of expanding the data-delivery business.
The change looks significant from the customer's perspective. When customers operate the tools themselves to build data, the responsibility for gathering experts, designing tasks, and inspecting results remains with the customer. When customers order a finished product instead, Snorkel takes on that process. As the unit being sold shifts from software usage to datasets and evaluation environments, Snorkel's own production capacity and quality control become the deciding factors in its competitiveness.
What Experts and Dedicated AI Are Building
Training AI to handle advanced coding tasks requires more than just lining up model answers. You need a development environment the AI can operate in, tasks that need to be accomplished, and scoring criteria that define correctness and quality. You also need to prevent the AI from exploiting loopholes in tasks midway through—earning a high score without actually completing the work. What Snorkel calls "Data 2.0" refers to this work of combining complex tasks, environments, and scoring criteria. This label is the company's own framing of the market.
According to Ratner, human experts create the seeds and constraints for tasks, and dedicated AI expands on them. The AI searches for areas where models struggle, adjusts tasks accordingly, and inspects submissions. Where necessary, it routes items back to the appropriate experts for review. The goal isn't simply having AI scale up human-created data—it's a cycle where expert corrections also feed back into improving the next generation of inspection AI.
The company claims that its internal pipeline for code-agent data and environments achieved more than 50% greater efficiency in quality control and more than 15 percentage points higher inspection accuracy compared to a baseline combining human review with off-the-shelf large language models. However, the public materials don't sufficiently disclose the sample size of tasks used in the comparison or the measurement conditions. These figures alone don't tell us about the quality of all delivered data or the actual performance improvements in customer models.
What Public Benchmarks Reveal About the Difficulty of Evaluation
As an example of work that can be verified externally, there's Senior SWE-Bench, which Snorkel built together with researchers. According to the company, it consists of 100 tasks derived from change requests across 12 real-world software development repositories, with 50 tasks made public and the remaining 50 kept private. Making some tasks public makes it easier to scrutinize the task design, while keeping others private guards against the risk of AI scoring high simply by memorizing questions and answers.
When tasks are designed to closely resemble real-world work, evaluation can't just be about whether something "worked." Judgments are also needed about whether code changes fit with existing design, or whether they introduce unnecessary complexity. The company's public benchmark serves as a concrete example showing that the value of finished data lies not just in the number of tasks, but in how the scoring is designed and operated. That said, the public benchmark and the private datasets delivered to customers are not the same thing. The former's public release doesn't directly confirm the quality or contractual outcomes of the latter.
Making Sense of the $375 Million Annualized Revenue Run Rate
Ratner stated this time that the business has grown more than 18-fold since launching the data-delivery service, and that by the week of the announcement, its annualized revenue run rate had exceeded $375 million. Reuters, in a same-day report, used the looser figure of "more than $350 million." The CEO's stated figure indicates a more precise level, but the measurement date and aggregation method haven't been disclosed. The 18-fold multiple also can't be used to back-calculate confirmed revenue from the prior year, since the denominator and scope aren't clearly defined.
Annualized revenue run rate is a metric that extrapolates the pace of revenue at a given point in time out to a full year. It doesn't mean the company actually earned $375 million over the past 12 months. Whether contracts persist, whether the customer base grows, and whether the same unit pricing can be maintained will all affect future full-year revenue. The company hasn't disclosed revenue by customer, retention rates, or gross margins this time around.
Even revenue figures among companies labeled "AI data companies" can't simply be lined up for comparison. TechCrunch pointed out that at companies that broker expert labor, a large portion of what they receive from clients goes right back out to pay those experts. Snorkel, by contrast, reportedly told TechCrunch that it sells finished datasets and reinforcement-learning environments, and that payments to experts are booked as cost of goods sold. When revenue is structured differently, comparing the size of annualized figures alone can't tell you how much of the business, or profit, each company actually retains. What's needed for comparison is gross margin after subtracting delivery costs from revenue, and how sustainable that margin is.
What Comes Next: Quality and the Repeatability of Revenue
According to Reuters, Snorkel expects to become profitable sometime in 2026. What we can confirm right now is the company's outlook, not an actual profit record. Whether it can retain profit after paying experts—while simultaneously scaling hiring, research, and data production capacity—will be the real financial test going forward.
The other open question is quality. An announcement that internal quality control has gotten faster demonstrates process efficiency, but whether customers actually get better models on unfamiliar tasks needs to be measured separately. Investment in public benchmarks helps build out that measurement methodology. Still, for the private data delivered to customers, only once reproducible evaluations and a track record of contract renewals are demonstrated can the rising revenue run rate be judged as a durable strength rather than a temporary surge.
