Comparison

AI training data companies, compared.

Five categories of supplier sell to the same buyers. They are selling completely different things, and only one of them buys what an operating business holds.

The five categories.

Coverage of this market tends to treat every supplier as interchangeable. They are not, and the difference decides who a given company should be talking to.

Annotation and labeling
Firms such as Scale and Surge apply human judgment to material a buyer already holds. Labeling, rating, ranking, and feedback. They sell work, not data.
Expert marketplaces
Firms such as Mercor and micro1 recruit specialists to produce training material to a brief. They sell access to people and their time.
Data marketplaces
Platforms that list datasets from many sellers and take a cut of each transaction. The seller competes on volume and gives up control of the destination.
Rights holders
Publishers, archives, and media owners licensing catalogs they already control. Large, well documented, and mostly already picked over.
Data labs
Firms that source material from operating businesses, verify provenance, clear rights, prepare it, and supply the labs as a principal. Polyshares is one of these, focused on enterprise operating records.

Why most of the market cannot help an operating business.

Three of those five categories have nothing to offer a company sitting on ten years of Slack and a CRM.

An annotation firm is not buying data. An expert marketplace wants to hire your people, not license your record. A marketplace will list you, which is not the same as paying you, and leaves you competing against everything else uploaded that quarter.

That leaves rights holders, which describes publishers rather than businesses, and data labs. For an ordinary operating company, a data lab is the category that actually writes a cheque.

What a serious buyer checks.

Provenance first. A clean chain showing the seller had the right to grant the license, established before price is discussed.

Then whether personal information can be properly removed, and how much of the record is genuinely unpublished. Anything already scraped from the open web has been free for years.

A supplier that skips these steps is not going to be able to place the material, whatever it pays you on paper.

Questions

Questions about the market.

What are the main types of AI training data company?

Five. Annotation and labeling firms, expert marketplaces, open data marketplaces, rights holders licensing their own catalogs, and data labs that source and clear material from operating businesses. They sell different things to the same buyers.

Which type buys data from ordinary companies?

Data labs. Annotation firms sell work performed on data a lab already has, expert marketplaces sell people's time, and marketplaces list whatever sellers upload. A data lab is the category that acquires existing records from businesses and pays for them.

How do AI labs actually buy training data?

Rarely by opening procurement against individual small businesses. They contract with a limited number of suppliers who can guarantee provenance, handle clearance, and deliver at a usable standard.

Can a company list its data on a marketplace instead?

It can, and it usually means competing on volume in a queue, giving up control of where the material ends up, and paying a cut to the platform. A principal buyer takes the other side of the trade instead.

Where does Polyshares fit?

We are a data lab focused on enterprise operating records. We license the record directly from the company that produced it, pay between $100K and $2M in a typical deal, anonymize it, and supply the AI labs.

Inquiries

Speak with a Managing Partner.

Five questions tell us whether there is a market for what your company holds. You will get a straight answer either way.

Check your data