Intelligence, which operates Design Arena—a platform for comparing AI-generated outputs—has raised a $7.9 million seed round. Users select their preferred option from outputs generated by multiple models, receiving superior generated results in return. The company's business involves selling the accumulated selection data to model companies for use in evaluation and improvement. According to Index Ventures, annual recurring revenue (ARR) grew from $5 million to $60 million in six months. In the design field, where correct answers are difficult to score automatically, the company has placed a mechanism for continuously collecting human "preferences" at the center of its business.

AD

$7.9 Million Raised, ARR Grows 12x in Six Months

On August 3, 2026, Index Ventures announced its participation in Intelligence's $7.9 million seed round. According to TechCrunch, Index Ventures led the round, with Conviction, A*, and Valkyrie also participating as investors. The co-founders are Grace Li and Kamryn Ohly, who in 2025 were computer science students at Harvard University.

The starting point was an AI game engine the two built during a hackathon. While the model could generate playable games, there remained inconsistencies in appearance and playfulness that humans could immediately notice. So they created an experiment that displayed generation results for the same prompt side by side, letting users select their preference. When a friend posted it on Reddit, thousands of visitors came overnight, and users reportedly began using it not to view benchmarks, but as a tool for obtaining good outputs across multiple models.

Index Ventures' investment announcement explains that Design Arena has been used by 5.5 million people, grew ARR from $5 million to $60 million in six months, and has become profitable with a team of 10. ARR is a metric that converts current recurring revenue into an annualized figure; it does not mean the company has already generated $60 million in sales. The number of customers, contract duration, and revenue concentration have not been disclosed. Still, the figure showing a free consumer-facing service converting into enterprise revenue in such a short period is the most eye-catching material in this round.

If the company's explanation holds true, ARR calculates to a 12x increase in half a year. Intelligence has already expanded its scope beyond design. In Prediction Arena, AI models are given actual funds to trade in prediction markets, testing whether they can forecast real-world events. The concept of placing evaluation targets within actual usage is shared with Design Arena.

Free "A vs. B" Comparisons Become Enterprise Data

On Design Arena, when a user selects a prompt and the type of generated content, four models respond to the same instruction. Two options are shown at a time with model names hidden, and a total of five votes are collected—including comparisons among winners, comparisons among losers, and a final ranking decision. Each vote carries equal weight and factors into win rates and Bradley-Terry model calculations, determining the ranking of the four models. By concealing brand names, users can compare the outputs themselves while suppressing preconceptions based on brand recognition.

What users gain is an experience where they can have multiple models generate websites, images, videos, games, and more from a single screen, then select the best option. There's no need to vote solely for the purpose of evaluating models. The act of searching for a desired output naturally generates comparison data. Through this design, Intelligence can continuously collect real-world use cases, instructions, and selection results—rather than fixed evaluation questions prepared in advance.

Because the same prompt is given, comparisons are limited to preferences among the displayed outputs. On the other hand, this ranking is not a comprehensive procurement evaluation that includes price, response time, data retention, and safety. If companies use these rankings for model selection, they need to separately verify the conditions required for their own use cases.

Companies, for their part, can have their own models participate in this flow and use what users actually chose as material for improvement. In an interview with TechCrunch, Li explained that the company's first contract with a major AI research lab was secured immediately after launch. Index Ventures states that Google used the platform to test Gemini 3 before its public release. OpenAI, Thinking Machines Lab, xAI, and Meta reportedly also used it for evaluation. Under this business model, the more free users increase, the more comparison targets and prompts can be expanded. Intelligence explains that it sells this signal to companies and has reached $60 million ARR, but has not disclosed the customer count or contract structure needed to judge the causal relationship between the two.

AD

Not One Overall Leader, But Preferences by Region and Use Case

Even when human votes are collected, a single "best model" common to everyone is not automatically determined. A paper published in April 2026 by University of Chicago researchers analyzed 115 users out of 13,383 Chatbot Arena users who had made 25 or more comparisons. The Spearman correlation between individual rankings created using the Bradley-Terry method and overall rankings averaged only 0.043, with 57% of users falling below 0.1. Under the ELO method, the average was 0.432—showing substantial variance between individual choices and aggregate rankings.

This research dealt with text responses and did not directly verify Design Arena's rankings. However, for services using the same pairwise comparison and aggregate ranking approach, the problem of collapsing user differences into an average cannot be ignored. In design, the colors, information density, and layout considered favorable can vary depending on purpose, region of use, and industry. Simply showing an overall leader may miss the answer a specific customer actually wants.

Li told TechCrunch that preference changes can be tracked by continent and over time based on login information. The weakness of aggregate rankings can turn into a different kind of value in an enterprise-facing business. This is because if it's possible to break down not just which model ranks highest overall, but which user, for which use case, chose which output, model companies can pinpoint specific areas for improvement. The value of the data Intelligence accumulates depends not only on the volume of votes but also on whether this segmentation can be reliably reproduced.

Data Responsibility Grows Along With Growth

Design Arena's privacy notice, updated on June 7, 2026, states that inputs such as prompts and generated outputs—including images and videos—may be shared with AI technology providers. Recipients use this data for evaluation, benchmarking, and model improvement/development, and it may also be used for training. Design Arena takes measures to remove usernames, email addresses, and similar information before sharing, but personal information that users have written into prompts or images may not always be removable.

In other words, the convenient model comparison screen and enterprise data collection arise from the same action. If users input confidential information or third parties' personal information, that content could be passed on to model providers. To maintain the convenience of freely trying multiple models, it's essential to clearly indicate the scope of sharing and enable users to judge what information is safe to submit.

There are also limits to how representative the evaluations are. Design Arena treats all votes with equal weight, without filtering or editorial adjustment. Models with fewer than 50 comparisons are excluded from the main graph, and results are displayed as provisional until they typically reach around 200 comparisons—but if voters are skewed by region or use case, the rankings will reflect that composition. An increase in vote count and sufficient inclusion of the user segments companies need are two separate conditions.

Intelligence has expanded its evaluation scope to include Prediction Arena, but has not disclosed how the $7.9 million will be used. Meanwhile, the customer composition, renewal rate, and churn rate behind the $60 million ARR remain unknown. Whether a business that feeds human preferences back into model improvement can sustain itself over the long term will be confirmed through contract renewal figures and whether evaluation results remain stable even when broken down by region and use case.