Skip to content

project

Common Crawl

Nonprofit organization that crawls the public web and publishes free, open datasets of web pages, widely used to train AI language models.

Known aliases

  • CCBot

Relationships

No evidence-backed relationships are recorded.

Current stories

product6 publishers

Microsoft measured a 93 percent click-through drop to nytimes.com from its own answer engine

A newly unredacted brief in the New York Times' case quotes a Microsoft director calling AI scraping the largest theft of labor in human history. The underlying exhibits are still sealed, so the quotes arrive without their original context.

Perspective Coverage

6 publishers
Builder
Builder 20%
Operator
Operator 44%
Investor
Investor 36%

Reality

Evidence60
Adoption
Insufficient
Hype gap+30
Incentives72
Confidence64