AI Firms Acquire and Digitize Millions of Books for Training Language Models
Artificial intelligence companies are purchasing millions of physical books, including rare and out-of-print titles, to scan and digitize their contents for training large language models. After scanning, the physical books are often destroyed or recycled. This practice, highlighted by court documents and reports, has sparked copyright lawsuits and raised concerns about the loss of scarce literary works. Companies like Anthropic, Meta, OpenAI, and Google have engaged in these efforts to obtain high-quality training data under legal frameworks such as the first-sale doctrine.
First-hand measurement across 2 sources
We measured how 2 outlets covered this story. No outlet gave this story a measurable political slant — there is no left–right reading to report. Overall sentiment is neutral (51/100). Lens Score 49/100.
Outlets measured: economictimes, economictimes. See how each one headlined and framed the same story in the source comparison below.
AI Analysis
Sentiment was consistent across outlets (50–52/100), indicating broadly factual reporting rather than editorialising.
Coverage timeline
economictimes broke this story on 30 Jul, 02:08 pm. Other outlets followed.
