Fair Use or Foul Play? The High-Stakes Fight Over AI Training Data

 

AI versus IP. This battle of acronyms has heated up in our federal courts over the last 18-24 months. And this summer we have seen a pair of rulings that start to resolve some big issues, while creating new ones.

The two cases I referenced above are all about how AI companies use copyrighted material to train large language models (LLMs). Both cases were brought by groups of well known authors—including Sarah Silverman, Michael Chabon, Ta-Nehisi Coates, and others—who argued that their books were used without permission to train AI systems. The defendants? Two of the biggest players in the space: Anthropic, the company behind Claude, and Meta, the makers of LLaMA.

The courts are being asked to figure out: Is it legal to train AI models on copyrighted books without a license? And if so, under what conditions?

Let’s break down what happened in each case, and then walk through what these rulings mean. We’ll look at what they say and, just as importantly, what they don’t say!

 

The Anthropic Case: A Split Ruling

In the Anthropic case, the authors claimed that the company had used their copyrighted works without permission to train Anthropic’s Claude model. The court broke this up into two issues:

1 Lawfully Acquired Works

The court found that using books that Anthropic had obtained legally (say, by purchasing them) could fall under fair use. The judge went so far as to call

the use “exceedingly transformative”—meaning that training a machine to understand language is quite different from reading or copying a book. Transformative use is one of the key factors in a fair use analysis, and courts often view uses that serve a new and different purpose more favorably. So if you buy a million books or CDs and train your LLM with that IP, this case says that practice is legal.

Pirated Content

Here’s the part that will drive lawyers (and everybody else) a bit crazy. Anthropic had also stored over 7 million pirated books, and the court ruled that storing and using those pirated books is not protected by fair use. That part of the case will move forward to determine whether damages are owed. The idea here is that fair use might not apply when the original content was acquired illegally.

So the message here is: Provenance matters. How you got the content matters just as much as what you did with it.

This distinction could have broad implications for companies developing generative AI. It means that even if the end use is transformative, if the training data was obtained through illegitimate means, there could still be liability.

What about scraping IP for training? What about IP you have access to through a subscription? These are all open issues that have not been decided yet by this or any other court.

The Meta Case: Lack of Harm

In the Meta case, the authors alleged that Meta used their books to train its LLaMA model. The court sided with Meta. Why?

The judge said the plaintiffs didn’t show that Meta’s use of their books caused any economic harm. This is important because, after the Supreme Court‘s 2023 decision in Warhol v. Goldsmith, courts are heavily focusing on whether the allegedly infringing use harms the original market for the work. While this factor was one of many considered by a court when assessing fair use, the Warhol case made this factor the most important, and judges have been applying it accordingly.

In other words, fair use isn’t just about how different the new use is—it’s also about whether the copyright owner loses out.

The court concluded that the authors hadn‘t provided enough evidence of market harm, so Meta prevailed on its fair use defense. That doesn’t mean training AI models on copyrighted works is always legal, but it does mean that without evidence of concrete harm, a fair use defense might stand up in court.

In my opinion, this creates a serious and substantial hurdle for authors and other rights holders. Many plaintiff IP holders may find it difficult to prove harm if the AI model doesn’t generate outputs that directly compete with or replicate their works. Proving economic harm might require expert analysis, market data, and clear evidence that AI-generated content is replacing human-authored material.

Take, for example, Disney.

Even if ChatGPT can produce outputs of Buzz Lightyear, that output alone (I think we can all agree) would have no impact on the market for Buzz Lightyear or Toy Story in any way.

If someone took that image and made merchandise out of it, or created a short film with it, then sure, that would be a problem, but that‘s a different use than training, and that would be a pretty blatant case of copyright and/or trademark infringement.

Yet, it would be very hard to show that the training alone (or even the training followed by similar output) would have a provable impact on the market for Disney IP. This is one reason why I think the Meta ruling could be an ominous one for IP holders.

 

What Do These Rulings Actually Mean?

First off, neither of these are Supreme Court or appellate rulings, so they don’t set binding precedent across the country. In other words, they can be very persuasive to future courts, or they can be disregarded altogether. That said, they’re still hugely influential, especially given how early we are in this wave of litigation and how respected each of these judges is. These decisions are like early markers, giving us a sense of where the legal winds may be blowing.

Here are a few takeaways:

Fair use can apply to AI training. These rulings suggest that using copyrighted text to train a model can be fair use, especially if the model isn’t reproducing or replacing the original work. Courts seem to be recognizing that using books to teach a machine how language works is fundamentally different from reprinting or redistributing those books. But how the content was obtained matters. If the content was pirated or scraped from sketchy sources, that could lead to liability. Even a transformative use can be undercut by an unlawful acquisition.

Evidence of harm matters more than ever. Creators will need to show that training an AI on their work actually caused them financial harm, which isn’t easy if the AI isn’t replacing their work or hurting sales. AI training is different from copying. Courts are recognizing that using text as input to teach a model is not the same as republishing that text.

That distinction is helping companies like Meta and Anthropic argue that they’re not exploiting the works in the traditional sense.

There may be a path forward for both sides. If creators and companies can figure out how to license content effectively— perhaps through collective licensing systems—this legal tension could ease over time. My guess is larger enterprises like media companies and publishers will enter into significant licensing deals with the tech companies (much like The Washington Post and NYT have recently done), whereas smaller creators and creatives will need to rely on the courts to police their IP, enforcement of which might only be possible where they can present evidence of harm.

 

Big Open Questions

These rulings don’t resolve everything. Not even close. Here are some things we still don’t know:

What types of content can legally be used for training? Is it okay to train my LLM by using blog posts? News articles? Unpublished manuscripts? What about content that is publicly available but not intended for machine consumption?

Is scraped content legal? If a model was trained on text pulled from websites without permission, does that violate copyright law? Does the presence of terms of service on a website affect the legality of scraping and use?

Does quantity matter? Does it make a legal difference if you train on 100 books versus 10 million? At what point does the volume of use affect whether something is still considered fair use?

What about purpose? Is there a difference between training a model for academic use versus for a commercial chatbot? Does it matter whether the company using the data is non-profit or for-profit?

Do derivative outputs raise new issues? If an AI generates text that is very similar to a copyrighted work, is the training still fair use? Or does the nature of the output introduce new liabilities?

We also haven’t seen rulings yet on outputs—i.e., what happens if the model spits out paragraphs that are too close to the original work? That’s a whole separate legal debate, one that may hinge more on traditional copyright principles of substantial similarity than on the mechanics of training. But what if the output is substantially similar and causes no economic harm. Fair use?

Confused yet? It’s ok, everybody is!

 

So Who‘s Winning?

In the short term, companies like Meta and Anthropic can take some comfort. Courts seem to be saying that training on lawfully acquired materials is likely to be legal, and that as long as you’re not reproducing the original work or harming the market for it, you may be covered by fair use.

That means, at worst, there‘s a pathway to market for AI companies, even if it requires negotiating licenses in some situations. It suggests that AI development won’t grind to a halt under the weight of copyright claims.

For authors and artists, these rulings are harder to swallow. Unless you can show real economic harm from AI training, it’s going to be tough to win in court. Larger publishers and media companies with licensing deals and data might have a better shot.

This is where the disparity in resources and infrastructure becomes clear. A single author might struggle to gather the evidence needed to prevail. But a large organization with a robust legal team and access to analytics might be able to build a much stronger case.

In some ways, this moment feels like a reversal of the typical narrative. Usually, it’s the big companies enforcing IP rights against smaller players. Now, the biggest tech companies in the world are arguing for broader fair use, while individual creators and rights holders are looking for protections.

It’s a little like Bush v. Gore in that sense—where the usual ideological lines got scrambled, and the decision was as much about power and structure as it was about principle. In that 2000 case, conservative justices relied on federal due process rights while liberal justices advocated for state authority—a reversal of the typical roles. Here, it’s usually the creative class invoking fair use, and the big corporations pushing back. But now it’s Google, Meta, and Anthropic leading the charge for fair use, and writers and artists asking for stricter interpretations of copyright law.

It’s a reminder that in law, context and incentives often shape arguments more than ideology.

 

Final Thoughts

We’re still in the very early innings of the AI/IP legal battles. These cases are shaping how courts think about the boundaries of fair use in the age of machine learning, but they leave a lot unresolved. They are clarifying some foundational issues, while raising new ones.

For now, here’s what you can count on:

If you’re building or advising an AI company, provenance of your training data matters.

If you’re a rights holder, gathering real evidence of economic harm will be key in putting a stop to training, if that’s your desire.

These decisions are not the final word. There will be appeals, new lawsuits, and possibly even legislative action. But they give us a glimpse of how courts may draw the lines between innovation and ownership in the years to come. Everything I’ve written here in the last 2000 words could be completely wrong in 12-18 months once the appellate courts have weighed in. That said, it’s an interesting start, and one that warrants some analysis and observation.

Let’s keep talking about it in plain English, with open minds, and a shared curiosity about what comes next.