They built a machine to mark students. It turned out the machine does not care whether a person or a model wrote the answer.
What they do
HackerEarth is a platform for developers, and it now describes itself as an evaluation layer for humans and AI. It began the way LeetCode did, as somewhere to practise, and practice, mock assessments and mock interviews are still free for developers.
More than ten million developers use it. Around three quarters are in India, roughly twenty per cent in the United States and the remainder spread across everywhere else, a distribution that simply follows where developers are.
The revenue comes from the other side. Large companies, many of them in the Fortune 500 and including Google, use the platform to evaluate developers, because a job posting can draw a hundred thousand applicants and no human team can read that much code. Candidates complete a coding assignment and the platform marks it automatically.
That marking has had to change. Writing code is no longer a novelty when a model can do it, so the platform added a sandbox with agentic coding support, comparable to what a developer would use day to day. The task used to be to write a sorting algorithm. Now it is to build a working website inside the sandbox in fifteen or twenty minutes, and what gets assessed is the quality of the prompts and whether there is any logical thinking behind them, or whether the candidate is typing at random.
HackerEarth was founded in Bangalore in 2012 by Sachin Gupta and Vivek Prakash, and is now headquartered in the Bay Area. Vikas Aditya, one of its key investors, took over as chief executive, with Gupta moving to chairman, and runs the company from California. The engineering team of more than a hundred people is in India. The company is self-sufficient and not raising, though it would take strategic money to expand.
Technical hiring is a well populated market. HackerRank is the closest comparison, with Codility, CoderPad and Karat taking different parts of the same problem. On the other side, model evaluation has become its own field, with LMArena crowd-voting the answer and SWE-bench the benchmark most coding models quote. Running competitions is a third business again, where Devpost and Kaggle have been at it for years. Very few companies sit in all three at once, and almost nobody arrived at the third by building the first.
The point: one engine, and both things it can grade are exploding
The asset here is not the question bank or the developer community. It is an engine that can look at a piece of code and say how good it is, at scale, without a person reading it.
That engine was built to screen job applicants. But code is code, and the same engine works just as well pointed at the output of a model. HackerEarth now uses it to evaluate frontier models and publishes benchmarks from the results.
The interesting part is that they score along separate dimensions rather than producing one number. Five models can all write functionally correct code while one is fast, another is memory efficient and a third is secure.
That distinction matters commercially. A trading application you run behind your own firewall needs performance and can worry less about hardening. Something exposed to the public inverts that completely. At the moment there is very little public visibility into whether one model writes more secure code than another, and an enterprise choosing where to send a workload needs exactly that.
Now look at the human side, because it is moving at the same time and in the same direction.
The word developer is being redefined underneath them. Sales people, marketing people and finance people are now expected to use coding agents. None of them were ever going to sit a test about sorting algorithms, and all of them can be assessed on whether they can direct an agent to build something sensible. The company's estimate is that this is roughly a tenfold expansion of who they can sell to.
That is a rare thing to be handed. Their addressable market grew by an order of magnitude because the world reclassified who counts as a developer, and the product that serves it is the product they already had.
Vibe Code Arena, and what it is really collecting
Alongside the evaluation business they run Vibe Code Arena, where developers build applications, compete, and enter innovation challenges that run every week or two. HackerEarth funds the prizes and, notably, pays for the model tokens the entrants consume, which is why it tracks model pricing daily.
The prizes are modest. A current AI innovation challenge carries around two thousand dollars. Two days after it went live it had drawn two hundred and fifty submitted projects, against roughly four thousand seven hundred on the platform in total. Each one includes the full source code and a working preview.
This part is opinion rather than anything the company has claimed. What is accumulating there is not a competition archive, it is a searchable body of things people chose to build, with the code attached and a grade already applied. Aditya's own suggestion is that an incubator could sponsor a challenge for a few thousand dollars and meet the top ten entrants, or simply be told which fifty of the four thousand ideas are worth a conversation. That is a deal flow business hiding inside a hackathon, and the sponsor is paying for the privilege of generating it.
Why it is bigger than it sounds
The short description is a coding assessment company. What it owns is more general than that.
- The grader is indifferent to authorship. An engine that judges code does not need to know whether a human or a model produced it, which is how a hiring company became a model benchmarking company without building anything new.
- Scoring by dimension beats a single number. Correct is not one thing, and an enterprise picking a model for a workload needs to know about security or memory specifically.
- The market redefined itself in their favour. Everyone in every department is now expected to direct an agent, which multiplies who can be assessed without the product changing.
What to watch
The measure is whether any model provider or large enterprise pays for the dimensional benchmarks. Publishing them builds a reputation, and the business only exists when somebody buys the report.
The second is whether non-developers actually get assessed. The tenfold expansion is real only if a company starts testing its sales team, and that is a change in corporate habit rather than in technology.
The third is the token bill. Funding the compute for every entrant is a generous way to grow an arena and a cost that scales with exactly the thing they want more of.