The users are not the customer. That is the clever part.
What they built
CoArena is a live arena where anyone can take a real computer task, race two of the world's top computer-use models against each other on it, and judge which one did it better. The company's own line is the sharp part: every battle becomes something the labs cannot build themselves, an honest test.
Prateek Jannu, co-founder and chief executive, has a master's from Stanford and a degree from Purdue, and invented Solo-Synth-GAN. Nitish Kovuru, co-founder and chief technology officer, previously launched Coasty, a computer-use framework that reached 83 per cent accuracy on OSWorld. The site is coarena.ai.
Evaluating models is a real field now. LMArena proved the crowd-voting format works and became the score people actually quote. Braintrust, Langfuse and LangSmith help teams evaluate their own systems. Scale AI ran private evaluations for labs, Epoch AI tracks capability trends, and HELM came out of academia. On computer use specifically, OSWorld and Browserbase supply the environments. Most of them sell evaluation to the person doing the evaluating.
The point: users come free, the challengers pay
Start with why anybody shows up.
You bring a real task. Two of the best computer-use models do it for you, side by side, and it costs you nothing. That is worth turning up for on its own. You get the work done, you get it done twice, and you get to see which approach handled your particular job better. People come for the free part and for the convenience of running two models at once without setting anything up.
Now look at who is on the other side.
Every new model that appears needs two things it cannot buy easily: somebody to try it, and a credible result to point at. Being ignored is the default outcome for a new model, however good it is, because the attention all sits with a handful of names.
CoArena can put a challenger into the ring against Claude or ChatGPT, in front of real users doing real work, and give it a fair chance to win. That is exposure and proof in one move, and it is worth money to the company being exposed.
So the users get something free and useful. The model makers get distribution and evidence. And the second group pays, because the second group is funded and the first group never would have.
Why the timing makes this easy
The reason this works now is the state of the market.
Model building is booming and crowded. New labs, new open-weight releases, new fine-tuned specialists, new national efforts, all launching into a world where a handful of incumbents hold nearly all the attention. Every one of them has the same problem: nobody is trying it.
In a boom, anyone who can offer exposure is welcomed rather than resisted. The newcomers need it badly, and unlike most startup customers they are well funded and can pay.
That is a rare and comfortable position. You are not persuading a reluctant buyer to change a habit. You are offering the thing they are already desperate for, to people who have just raised money specifically to get noticed.
It also means the arena gets better as the market gets more crowded. More challengers means more matches, more matches means more results, and more results means the scoreboard matters more.
What else could be sold into this fight
This is my opinion, not something the company has said.
If the insight is that a crowded model market will pay to be seen and compared, then CoArena is one product out of several, and the others are open.
Be where new models launch. Every new model needs a first thousand users on day one. Whoever becomes the default place a model debuts holds the introduction, and introductions are valuable in every industry that has ever had them.
Sell the switch, not the choice. Companies stay on a model they have outgrown because moving prompts, tools and evaluations across is painful. Make switching trivial and every challenger will pay you, because the thing standing between them and a customer is inertia rather than quality.
Certify, do not rank. A ranking has one winner and everybody else loses. A certificate saying tested and verified at this task can be earned by many, which is a far better business, and it is what safety standards bodies have always understood.
Run the bake-off for the buyer. Enterprises choosing a model today do it badly, with a spreadsheet and a fortnight. Run that properly as a service and both sides pay: the buyer for a real answer, the vendors to be in the running.
Sell the losing results. A model that fails a task publicly has just been handed the most valuable thing it can get, which is a precise description of what to fix. The report is worth more to the loser than the trophy is to the winner.
The common shape is worth stating plainly. When a market is crowded, loud and well funded, the reliable business is not joining the fight. It is selling something to everyone who is in it.
Why it is bigger than it sounds
The short description is a benchmark site. The business model is the interesting half.
- The free side is genuinely useful. Two top models do your real task at no cost, which is a reason to turn up even if you care nothing about benchmarks.
- The paying side is desperate and funded. New models need someone to try them and a result to point at, and neither can be bought easily anywhere else.
- A boom makes the sale easy. You are not changing anybody's habit, you are offering exposure to people who just raised money to get noticed.
What to watch
The measure is whether any model maker has paid to be in the arena. That single transaction proves the business, and everything before it is a popular website.
The second is volume, because crowd-judged results only mean something in bulk. A handful of votes is an opinion. Tens of thousands is a measurement.
The third is whether a lab ever cites a CoArena result in its own announcement. That is the moment it stops being a site and becomes the scoreboard, and the paying side gets much easier from there.