The first appeals court to rule on AI training didn't care much that a model sat in the middle. It cared what the model was for.
ROSS Intelligence shut down its legal research platform in 2021, citing the cost of being sued by Thomson Reuters. Five years later, on September 29, the company lost the first AI-training copyright case ever decided by a US appeals court.
In Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., the Third Circuit affirmed partial summary judgment for Thomson Reuters. It held that 2,243 Westlaw headnotes were original enough to be protected by copyright, and that ROSS had not made fair use of them when it trained a system meant to compete with Westlaw. The opinion, written by Judge Tamika Montgomery-Reeves, was filed under seal and released the next day.
I've spent years arguing that AI governance fails in the gap between what a company documents before deployment and what its systems actually do once they're running. I did not expect an opinion about legal headnotes to make a version of that argument.
But that is roughly what this court did. It declined to treat training as a sealed technical step and looked instead at what the training was for.
The limits matter, so let me state them once. ROSS did not build a generative model. Its system wrote nothing. It took plain-language legal questions and returned passages from existing court opinions. The material it copied was editorial work written by Westlaw's staff, and the product it built was designed to replace Westlaw. The court itself called this "no more than an ordinary copyright case." Anyone reading it as a ruling on ChatGPT is reading something that isn't there.
What is there is a method. And methods travel further than holdings.
What ROSS Took
Westlaw headnotes are short summaries of the legal points in a judicial opinion. They sit above the opinion and link to the passage they summarize. Thomson Reuters tells its editors to keep each one to 800 characters where possible, to stay close to the court's language, and to make sure each one stands on its own for a reader who hasn't seen the opinion.
The opinions themselves are public law. Nobody owns them. The headnotes are different, the court said, because someone had to decide which points of law mattered and how to say them.
That clears copyright's originality bar, which is famously low. The court was careful not to go further. It expressly left open whether a headnote that copies a court's words verbatim would qualify, because none of the 2,243 did.
ROSS started as an idea from three University of Toronto computer science students in an IBM Watson competition: legal research without Boolean search strings. The platform they built drew on roughly ten million judicial opinions and returned the passages it judged most responsive to a question.
To train it, ROSS hired LegalEase Solutions, which worked with a subcontractor, Morae Global. They produced about 25,000 training memos. Each one posed a legal question and offered four to six passages from court opinions, graded great, good, topical, or irrelevant. The memo writers built the questions from Westlaw headnotes because, in a phrase the court quoted from the record, the headnotes were "an easy way" to frame them. The great answers were most often the very passages Westlaw had linked to those headnotes.
The 2,243 figure has an origin worth knowing. ROSS's own expert identified a batch of 2,830 memo questions that closely tracked Westlaw's headnote language while differing significantly from the underlying opinions. The district court went through that batch and found 2,243 where no reasonable juror could conclude the headnotes hadn't been copied.
None of this was a gray zone. ROSS ran ads comparing itself to Westlaw, priced its service "in line with" Westlaw, and its executives hoped it would serve as a substitute. Some law firms switched.
Two footnotes are worth reading closely. In the first, the court describes ROSS trying to get into Westlaw with a law-firm investor's credentials after being told that violated Westlaw's terms, an employee asking about an account while posing as a solo practitioner, and another using a student account while hiding that he worked for a competitor. To whatever extent good faith still matters in fair use, the court said, it cut against ROSS.
The second is shorter and, for anyone who outsources data work, more uncomfortable. On appeal, ROSS did not contest that the copying done by LegalEase and Morae Global was attributable to it. The labeling was outsourced. ROSS didn't argue that the liability had been.
Training Didn't Buy a New Purpose
The heart of the opinion is the first fair-use factor, and specifically whether ROSS's use was transformative.
ROSS's argument was intuitive. Westlaw shows headnotes to lawyers. ROSS never showed them to anyone. It used them inside a training pipeline, and its users only ever saw text from public court opinions.
The court allowed that this intermediate step "arguably" presented a slight difference. Then it looked past it. Westlaw uses headnotes to help researchers find relevant law. ROSS used the same headnotes to build a platform that helps researchers find relevant law. Same ultimate purpose, the court said, which made the use "minimally transformative, at best."
That is the passage AI companies should read twice. Training is a technical operation. Fair use asks a legal question about purpose. Putting a model between the copy and the product did not change the question.
ROSS leaned on two lines of precedent, and the court distinguished both. Google Books scanned entire libraries, but it built a search tool whose purpose differed from reading, one that could even send readers toward buying the books. ROSS's tool didn't send anyone to Westlaw. By its own admission, it aimed to replace it. The software cases, from Sega to Google v. Oracle, permitted copying that was necessary to reach unprotected functional elements and make programs work together. ROSS had ten million free judicial opinions it could have used to write its training questions. It used the headnotes because they were faster.
The court's answer fits on one line: "Unlike necessity, ease is not a justification for copying." Anyone whose training-data strategy rests mainly on convenience should sit with that line for a while.
The other factors followed. The headnotes were published and largely factual, which tilted the second factor slightly toward ROSS. On the third, ROSS pointed out that it had taken only 0.08 percent of Westlaw's 28 million headnotes. The court was unmoved. The question is whether more was taken than necessary, and with the opinions freely available and no transformative purpose, it was. Each headnote, the court added, is a complete work in itself. The fourth factor, market harm, went to Thomson Reuters because ROSS was building a substitute in Westlaw's own market.
ROSS also argued public benefit: its copying would widen access to the law, the ruling would stall AI development, and AI matters for national security. The court noted that the opinions were already free, that ROSS charged roughly what Westlaw charged, and that ROSS had offered no evidence for either of the larger claims. Some AI may raise national security concerns, the court wrote, but that does not give a company "carte blanche to violate copyright law merely because it incorporates AI."
The Market Nobody Had Sold Into Yet
The part of this opinion with the longest reach, I think, is its treatment of licensing. Thomson Reuters argued that ROSS had damaged a potential market for licensing headnotes as AI training data. ROSS answered that Thomson Reuters had never licensed headnotes to anyone for that purpose, so there was no market to damage.
The court disagreed. It found evidence that the market for licensing headnotes as training text is "rapidly developing," and that Thomson Reuters was already training its own AI search products on them. A copyright owner's choice not to sell licenses, the court said, doesn't prove the market is imaginary. By taking the headnotes without permission, ROSS "usurped" Thomson Reuters's opportunity to enter it. ROSS offered no evidence to the contrary, which is a large part of why this was decided on summary judgment.
Rights holders will quote this passage for years. Training rights can carry value of their own, whether the owner sells them, uses them internally, holds them for later, or refuses to sell them to a rival. No price list required.
This is also where I get uneasy. At oral argument in June, ROSS's lawyer called the theory circular: any copyright owner can claim it would have licensed whatever use is later defended as fair. If the bare possibility of a training license is enough to show market harm, fair use stops doing much work in training cases at all. The doctrine exists partly so that owners can't control every secondary use simply by proposing to charge for it.
The court didn't take that on directly. It rested on this record: a developing market, internal use, and a defendant building a direct substitute. Future courts will have to decide how much evidence it takes before a hypothetical training market counts, especially when the new product doesn't compete with the source at all.
There's one more fact worth setting next to that holding, and it isn't in the opinion. According to Thomson Reuters's 2020 complaint, as reported at the time, Thomson Reuters refused to license its content to ROSS, and only then did ROSS turn to LegalEase. Nothing in copyright law obliges a publisher to license its work to a competitor. But read alongside the market-harm holding, it describes a structure worth naming. The owner of the training data can decline to license a challenger, and then cite the training-data market it never opened as the harm when the challenger finds another way in.
Generative AI Gets Distance, Not Immunity
The court was explicit that this is not an LLM case, and it said so in a footnote with real news in it.
On September 1, the Department of Justice filed a statement of interest in In re OpenAI, the consolidated copyright litigation in the Southern District of New York. Relying on Bartz v. Anthropic, the DOJ argued that training a large language model capable of generating original responses is transformative, and that the training at issue didn't produce "substitutive competition." The Third Circuit said those concerns don't apply here. ROSS can't generate anything, and it was built precisely to be a substitute. Then the court added, pointedly, that the DOJ knows how to assert its interests in these cases and chose not to in this one.
So the federal government is now taking sides in AI training litigation, and an appeals court made a point of recording where it hadn't. For anyone tracking where AI copyright policy is actually being made, that footnote may matter more than the holding.
Generative capability does create real distance. A foundation model learns statistical patterns across an enormous corpus drawn from unrelated fields. Any single work is a sliver of it, and its outputs serve countless purposes that have nothing to do with the source. Those facts support a far stronger claim of different purpose than ROSS could make.
But this opinion suggests courts will look at the business wrapped around the model rather than accept "generative" as a label that ends the analysis. A general model trained across many domains is one set of facts. A specialized model fine-tuned on a publisher's work to deliver that publisher's core service is another. In between sits most of what companies are actually building right now: retrieval systems, domain copilots, research agents, enterprise tools with a proprietary collection behind them. Their exposure will depend less on what the technology is called than on whom it is built to replace.
Outputs still matter, especially where a model reproduces protected text. What ROSS shows is that a plaintiff doesn't always need to win on outputs. ROSS's outputs were public court opinions. It lost anyway.
If You Build on Someone Else's Judgment
For model builders, provenance is now where the analysis starts, not where it ends. The harder questions are what human judgment sits inside a dataset, why the model needs it, whether uncopyrighted material could do the same job, and how the finished product relates to the source's business. The closer that relationship, the harder it becomes to argue that training happened in a world apart.
I expect licensing agreements to get much more specific as a result. Whether data may train a general model or a direct competitor will become a priced distinction. So will what happens to the weights and derived datasets afterward.
Publishers, meanwhile, have every reason to document the editorial labor in their products. In this case, that labor is exactly where the line between protected expression and public law fell.
The uncomfortable side is concentration. Incumbents with deep archives can train their own products and decline to supply challengers. The company that owns the premium dataset may also be the dominant competitor downstream, and the one deciding who gets to train on it. Copyright becomes part of the competitive architecture of AI markets, next to compute and capital.
That might reward real investment in reliable information. It might also make incumbents very hard to dislodge. Fair use was never designed to settle that tension, and it won't. Competition policy, licensing norms, and public investment in open datasets will decide which way it breaks.
What the Court Actually Asked
ROSS lost because the court looked through the model to the business it served. It asked what was copied, whether the copying was necessary, what the copied material did inside the final product, and whose market that product was built to take.
None of those questions depends on whether a system is labeled AI. That's why I think this opinion will matter well beyond legal research, even though its holding is narrow.
Applied to foundation models, the answers may come out very differently, and the court deliberately left that fight open. What it closed is the shortcut: the idea that an intermediate training step turns copying into transformation on its own.
It's the same move I think governance has to make with every AI system. Don't ask what the paperwork says the system is. Ask what it does once it's running, and for whom.
ROSS has been gone for five years. The question the court asked about it is going to be put to companies that are very much alive.