Design Arena Secures $7.9M Seed Round Led by Index Ventures.
Design Arena Raises $7.9M to Teach AI Models What Good Taste Looks Like:
How a college side project became a $60M-ARR pipeline of human judgment for frontier AI labs.
$7.9M: Seed Round Led by Index Ventures
5.3M: Users Ranking AI Output
$60M ARR: Current Run Rate
Intelligence, the company behind the AI evaluation platform Design Arena, has closed a $7.9 million seed round led by Index Ventures, with participation from Conviction, A*, Valkyrie, and others. The raise formalizes what's become one of the AI industry's more unusual success stories: a tool for measuring good taste that grew out of a broken game engine.
Co-founder Grace Li has said the company started just weeks before her 2025 graduation, when she and a group of college friends were building an AI game engine. The models could produce functional games, but none of them were actually fun to play — which raised a harder question than the team expected: how do you even measure whether a game is fun in the first place?
Their answer was that there was no substitute for real human judgment, so they set out to collect honest feedback at scale. That effort became Design Arena, now used by 5.3 million people worldwide.
1: From Side Project to Frontier-Lab Vendor:
The team quickly discovered they weren't the only ones who needed this kind of feedback. AI companies broadly were struggling with the same bottleneck — they had models capable of generating design work, but no scalable way to know if the output was actually good. Design Arena filled that gap, and within about a week of realizing it, the team had closed its first major deal with a frontier AI lab. Growth followed quickly from there.
2: How the Platform Turns Opinions Into Data:
For everyday users: Design Arena functions like a sophisticated model router. Users type a prompt into a ChatGPT-style interface, choose a format — websites, images, and roughly a dozen other visual categories — and then get shown a series of head-to-head comparisons, ranking outputs from best to worst until a clear winner emerges.
For enterprise clients: The real business value sits on the other side of the platform. Participating AI labs treat Design Arena as a constant stream of feedback for their media-generating models.

The Hidden AI War
Nobody Is Telling You About
Our latest documentary deep-dive into the geopolitical struggle for machine intelligence dominance. Explore the two paths of AI development: open source vs. closed architecture.
Most users don't care which underlying model produced a given result — they just want the best output — which makes their rankings a clean signal of what people actually prefer, undistorted by brand loyalty.
Because users have to log in to see their results, Intelligence can also track how design preferences shift across regions and over time — Li has pointed out, for instance, that web dashboards in Asia tend to trend toward a more maximalist visual style. That kind of human-sourced signal is a meaningful complement to automated benchmarks, which scale more easily but are also easier to game, as last week's Hugging Face breach demonstrated in dramatic fashion.
Automated benchmarks can be gamed at scale — real people ranking real output is much harder to fake.
3: Not Every Human-Feedback Startup Survives:
Human evaluation isn't a guaranteed winning business model on its own. Yupp, a competing platform, shut down earlier this year less than a year after launch, despite raising $33 million from a16z crypto's Chris Dixon, landing frontier-model customers, and reportedly reaching more than 1.3 million users — it simply couldn't turn that traction into a sustainable business.
Even so, other players in the space are thriving. LM Arena, which applies a similar human-ranking approach to text-based model responses, raised a $150 million Series A in January, only four months after formally launching its paid product — a sign that investors still see real staying power in human-judged AI evaluation, even if not every entrant makes it.
Great Models Still Need Real Judgment Behind the Wheel:
Design Arena's rise makes a broader point that applies well beyond frontier labs: even the best AI output still needs a human judgment layer to know if it's actually good.
That's the same philosophy behind Otherworlds AI's Agent+ platform — AI agents built to fit how your business actually works, refined with real feedback rather than deployed and left alone.
Whether you need a ready-to-launch Agent+ deployment or a custom enterprise AI build tailored to your workflow, our team can help you get it right the first time.
Visit otherworldsai.com to talk to us.







