Artificial Intelligence
Estimated Quarter of Llama AI Sources Came From Books
Lawsuits alleging the makers of AI models infringed on copyrights during the training process are piling up. A new proposed class action lawsuit brought against Meta by major publishers in a Manhattan federal court will see a first hearing in September. Previously, groups of authors had sued tech companies, but in the cases of and had not succeeded in showing that their copyrights had been breached by the training process itself.
However, the latter case found that Anthropic had stored a whopping 7 million pirated books to train Claude, which was ruled illegal. The finding also shows the sheer volume of works that were fed to large language models.
, CEO of T1U.ai, shows that between five major LLMs, Meta's Llama is estimated to have been trained on approximately a quarter of book material as of 2025, actually much higher that the estimated share for Claude. Comparably many books are also thought to have been used on the training of ChatGPT and Google's Gemini. The calculation based on company disclosures, model documentation and additional analysis shows that Claude used a high share of academic papers, while Gemini was trained on a majority of web crawl data. Grok by Elon Musk's company xAI, maybe predictably so, used a higher share of social media input.
Description
This chart shows estimated data source distribution for major AI models as of 2025.
Related Infographics
Any more questions?
Get in touch with us quickly and easily.
We are happy to help!
ƽ Content & Design
Need infographics, animated videos, presentations, data research or social media charts?