News

Court Rules Scraping Search Results for AI Training Falls Under Fair Use

A US court has dismissed Google's attempt to stop others from scraping search results to train artificial intelligence systems. The ruling represents a significant development in the ongoing legal debate over what data AI companies can legitimately use to develop their models.

The decision could have far-reaching implications for the AI industry, where training data has become a critical and contested resource. Companies developing large language models and other AI systems have faced increasing scrutiny over how they obtain the vast amounts of data needed to train their systems.

Google had argued that scraping its search results constituted copyright infringement and violated its terms of service. However, the court found that using search results for AI training purposes qualifies as fair use under copyright law, a doctrine that permits limited use of copyrighted material without permission for purposes such as criticism, comment, news reporting, teaching, scholarship, or research.

This ruling may encourage AI companies to continue developing their systems using publicly available web data while also potentially prompting copyright holders to seek new legislative remedies for their concerns about unauthorized AI training.

Sources