News

AI Firms' Book Acquisition Practices Raise Questions About Training Data Ethics

A developing story suggests that major AI companies have been purchasing vast quantities of physical books to use as training material for their language models, with the process reportedly outsourced to third-party middlemen who acquire titles discreetly.

The practice highlights the ongoing debates around data sourcing in AI development. Training large language models requires enormous amounts of text data, and publishers have expressed concerns that their copyrighted works are being used without compensation or clear permission.

The reported method of using intermediaries to acquire books reportedly allows AI companies to accumulate large libraries while maintaining a degree of separation from the direct acquisition process. According to reports, many of these physical books are subsequently destroyed after their content has been digitized and processed for training purposes.

This approach differs from companies that have explicitly negotiated licensing agreements with publishers or authors. The lack of transparent agreements has drawn criticism from authors' groups and publishers who argue that creative works deserve protection and compensation regardless of how they are ultimately used in AI training pipelines.

The broader AI industry continues to face legal challenges over training data practices, with multiple lawsuits pending from publishers and authors alleging copyright infringement. The specifics of how training data is obtained—whether through direct licensing, web scraping, or purchasing physical books—remain central to these ongoing legal and ethical discussions.

Sources