Treasure hunting for the AI age is a startup in its own right, and this is one.
But finding the treasure is the half you can see. The half worth thinking about is that whoever finds it opens it first.
What they built
Petrarch brings internal company data to the frontier AI labs, starting with manufacturing and industrial businesses. Its sources are codebases, project files and payment records, obtained from bankruptcy courts. It de-identifies that material and prepares it for model training.
Bankruptcy is the part that makes the supply possible at all. A functioning company will not sell its internal files at any price, because they hold its methods, its customers and its mistakes. A company being liquidated has no such objection, and a court is obliged to convert everything into money for creditors. Manufacturing is the right first target because models are conspicuously weak at physical-world work, and that knowledge was never posted anywhere to be scraped.
The founders are Ian Lee, chief executive, previously a fellow at Meridian AI building training data and go-to-market infrastructure for AI-native spreadsheets in financial services, and a technical advisor at Scale AI. Samuel Hahn, chief operating officer, previously led go-to-market at Magier AI, a Techstars company building agentic redaction and data-compliance products, and was a venture analyst at Mainstreet. Sudhish Swain, chief technology officer, previously built computer vision software for identifying medical equipment, deployed in California hospitals, as head of software engineering at Clean Sweep.
Who else is in this field
Supplying training data has become one of the largest businesses in AI. Scale AI built it into a giant. Surge AI, Appen, Labelbox, Snorkel AI and Turing all work the labelling and expert-data side. Gretel and Mostly AI generate synthetic data instead. Shutterstock, Getty Images and Reddit discovered their archives were worth more to labs than to their original customers. Common Crawl supplied the open web that started all of it.
Every one of them sells what it holds. None of them is in the business of finding something nobody else can reach, and then being the first to look inside.
The point: the finder reads it first
Petrarch will open these archives before any lab does.
It will know what is in them. Which industries left behind the richest records. Which kinds of problem appear again and again. Where the gaps are. What a model trained on this material would suddenly become good at, and what it still would not.
Nobody else can know any of that, because nobody else has seen the inside of the vault.
Being the first handler of a treasure is a different asset from owning it. You know what can be built on top of it before the people you are selling to know it exists, and that is a head start nobody has to grant you and nobody can take away.
Think about what that turns into. A supplier who has read ten thousand industrial codebases knows precisely which tool is now buildable and was not buildable last year. They know it months before the lab that bought the data finishes training on it, and years before the market notices. Selling the data is the revenue today. Building the thing the data makes possible is the far larger business, and it is sitting there unclaimed.
That second step is not in the company's stated mandate. It may never be. But it is the natural consequence of the position, and it is exactly the sort of thing that gets written into a mandate later, once somebody realises it was always there.
Reach uniquely, then build uniquely
Strip out the specifics and there is a strategy here worth carrying away, and it is the reason this file exists.
Step one: reach uniquely. Get to a source of value nobody else can reach, for a structural reason rather than a clever one. Structural matters. A clever advantage is copied by the next clever person; a structural one holds because the alternative is genuinely blocked for everybody else.
Step two: build uniquely. Use what reaching it taught you to build the thing that sits on top, which you can see and nobody else can.
Most companies only ever run step one, and treat step two as somebody else's opportunity. They sell the ore and watch other people build with the metal.
There are still a great many troves waiting to be found. Most will be less dramatic than a bankruptcy court full of industrial records, and the pattern holds for all of them: get there first, learn what is inside, and build what only the learning revealed. In an era when everybody has the same models, what you have reached and nobody else has is close to the only durable difference left.
The question that decides it
The honest caveat is the interesting one, and it is the thing to put to the founders directly.
Data businesses are very often contractually barred from using what they broker. Licence terms can reserve derived works to the buyer. The de-identification promise that makes this legal at all may restrict what Petrarch can do with the material itself.
So the question is not whether the second step is valuable. It obviously is. The question is whether the deals being signed today quietly give it away.
That the two founders with data compliance and redaction backgrounds are the ones writing those contracts is the most encouraging detail available. They of all people should know what a licence gives up.
Why it is bigger than it sounds
The short description is training data for frontier labs. The position is worth considerably more than the product.
- The finder reads it first. Opening a trove before anyone else means knowing what can be built on top of it before the buyers do, which is a head start nobody grants and nobody can take back.
- The second step is where the money is. Selling the data is today's revenue. Building what only the data made visible is the larger business, and nobody has claimed it.
- It is a repeatable strategy. Reach somewhere nobody else can reach, then build what only that reach revealed. There are more troves out there, and this is how you take one.
What to watch
The thing that decides how big this gets is whether Petrarch ever ships something built on what it learned. A tool, a benchmark, a model of its own, anything using knowledge only a first handler could have. The day that appears, this stops being a data broker and becomes the first worked example of a strategy other founders will copy.
Before that, the near-term measure is whether a frontier lab buys the same kind of data twice. A single purchase proves curiosity. A repeat purchase proves the material actually improved a model at something it was previously bad at.
And the legal route has to hold as volume rises. Buying a bankrupt company's data through a court is clean in principle, and the questions get harder at scale: what the customers and employees inside those files ever consented to, and whether de-identification stands up when somebody tests it properly. Getting that right is the whole company, and the founders appear to know it.