Here's an uncomfortable thought experiment: if you've ever published a book, an article, or even a really detailed blog post, there's a decent chance it helped train the chatbot you now worry might replace you. Nobody asked. Nobody paid you. That feels like it should be illegal. According to the lawyers actually litigating this stuff, "should be" and "is" are doing a lot of heavy lifting in that sentence.
Think of it like this: a novelist spends decades reading everything they can get their hands on, absorbing style, structure, and voice, then sits down to write something original. Nobody accuses them of stealing from every author on their bookshelf. Courts are now being asked whether a machine doing something that looks similar, at a scale no human could match, deserves the same benefit of the doubt.
Last year, Judge William Alsup handed down one of the first major rulings on AI training and copyright, ordering Anthropic to pay $1.5 billion to settle claims brought by a group of writers. On its face, that reads like a landslide win for authors. It wasn't, at least not in the way most headlines suggested.
Alsup didn't rule that training AI on copyrighted books was illegal. He ruled that training was lawful. What he penalized Anthropic for was how it got the books: pirating them from illegal shadow libraries rather than acquiring them legitimately. In his opinion, Alsup compared the way a large language model ingests text to how an aspiring writer studies literature, writing that Anthropic's models trained on existing works "not to race ahead and replicate or supplant them, but to turn a hard corner and create something different."
Cathy Gellis, an attorney specializing in intellectual property and technology law, told TechCrunch the ruling actually favors AI companies more than the headline dollar figure suggests. "I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work," she said. Her math is blunt: a $1.5 billion penalty barely registers against a company projected to bring in roughly $200 billion a year by 2028.
Gellis also points to the legal principle underneath the ruling: "Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work." Training a model on a book, in this framing, is closer to reading it than photocopying it.
Nearly every one of these disputes comes down to fair use, the legal carve-out that lets people use copyrighted material without permission for things like commentary, education, or parody. Courts weigh factors including the purpose of the use, how much material was used, and whether it damages the market for the original work.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch that judges are landing on a rough dividing line: training that directly competes with the original work tends to lose, while training that doesn't compete tends to survive. "What's tending to win is if what you're doing is you're training on somebody's property because your purpose is to directly compete, then the courts will frown on it," Henderson said. "If what you're doing is not going to compete, then the courts are tending to find ways that it will be okay."
He points to Thomson Reuters' lawsuit against Ross Intelligence, a research firm that copied Thomson Reuters' content to build a competing AI-based legal research platform, as the clearest example of the losing side of that line. Courts found Ross's use wasn't transformative because it didn't serve a meaningfully different purpose than the original product it was scraping from.
Part of why this is such a mess is structural. US copyright law hasn't had a major update since 1976, decades before anyone imagined a model trained on hundreds of millions of books and articles. Judges are stuck retrofitting 50-year-old statutory language onto a technology its authors couldn't have conceived of. "Everybody is very worried right now because the law is all over the place," Henderson said, "and it's because of this question. They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question."
This isn't limited to text generation. The same uncertainty runs through the entire AI and copyright landscape, including whether AI-generated output itself can be copyrighted at all. US copyright law only protects works created by humans, and the Copyright Office has already denied registration to AI-assisted art when the human contribution was judged too minimal. That's a separate legal fight from the training question, but it comes from the same root problem: a legal system built for human authorship trying to make sense of generative AI models built on transformer architecture that only became mainstream after 2017, when the current AI boom really took off.
The uncomfortable truth is that "trained on your book without asking" and "illegal" are not the same thing under current US law, and courts so far seem more interested in punishing how companies acquired material than whether they used it to train a model at all. That's a meaningful distinction for any company doing this work: pirate your training data and you're exposed, license or scrape from legal sources and you're on much firmer ground, regardless of how the authors whose work built your product feel about it. Until Congress rewrites copyright law for the AI era, or a higher court draws a brighter line, this is going to keep getting settled case by case, judge by judge, with authors mostly left arguing after the fact.
This post contains no sponsored content or affiliate links. All quotes and case details are sourced from TechCrunch's original reporting, linked above.