Hands holding a stylus above a tablet screen showing two side-by-side design templates.

Human Judgment: Intelligence Lands $7.9M to Fix Bad AI Design Choices

As co-founder Grace Li tells it, her company started just a few weeks before college graduation in 2025. She joined a handful of college friends to make their own artificial intelligence game engine work. The underlying software models could generate functional games without issues, but none of the resulting titles felt fun to play. That problem raised a bigger question for the team: how do you teach an algorithm good taste?

The team realized early on that automated benchmarks could never replace raw human judgment. They needed a reliable way to get real feedback from actual people at scale. That realization led directly to the creation of Design Arena. The tool grew fast, reaching over 5.3 million users around the globe. As it turned out, tech companies building machine learning tools desperately needed scaleable human feedback, and many showed a willingness to pay big money to get it.

Li noted that collecting human preferences solved the main bottleneck holding back model improvements. About a week after launching, the team closed its first major commercial deal with a top research lab. On Monday, the company behind Design Arena, now officially named Intelligence, announced a $7.9 million seed funding round led by Index Ventures. Conviction partners Sarah Guo and Mike Vernal participated in the round alongside A star, Valkyrie, and several individual investors.

For everyday consumer users, Design Arena looks and feels like a specialized model routing hub. It features a chat window for prompts, complete with dropdown controls to select specific formats, images, websites, or visual styles. Once you enter a request, the system presents a series of side-by-side choices comparing different outputs. Users rank those options from best to worst based on personal preference.

While consumers enjoy the free routing features, the primary business value lives on the enterprise side. Participating artificial intelligence labs treat the platform as a continuous stream of direct human feedback. Everyday users do not care which specific company built the underlying model. They simply want the cleanest, best-looking visual output possible. Their voting patterns provide clear direction on what real people prefer, giving developers actionable data to refine future software releases.

Frontier research labs view that human rating data as vital. Li shared that the platform currently generates $60 million in annual recurring revenue, locking in its position as a primary source of human evaluation data for the tech industry. Because users log in to download their generated media, Intelligence tracks how aesthetic preferences change across different countries and regions. For instance, users in Asia frequently favor maximalist design styles compared to Western audiences.

Collecting real human feedback complements automated benchmarking tools, which often fall victim to manipulation or gaming. However, crowdsourced feedback platforms still face market risks. Competing tool Yupp shut down despite raising $33 million, proving that crowdsourcing alone does not guarantee long-term survival. On the flip side, evaluation platform LM Arena raised $150 million in Series A funding earlier this year. For now, Intelligence plans to expand its core voting features and capture a larger share of the fast-growing human evaluation market.