@cmdicely LLM outputs generated solely by text prompts are exactly what the (c) office has said are not (c)able, and TOS alone does not a trade secret make when those secrets are readily ascertainable from the service itself.
EleutherAI's Stella Biderman highlights arguments challenging LLM IP theft claims
Arguments cite non-copyrightable outputs and reverse-engineering precedents.
Many users rejected challenges to IP theft claims in LLM training as pedantic or elitist, insisting models are built on stolen copyrighted material from the start.
No Digg Deeper questions have been answered for this story yet.
Most Activity
To summarize, distillation probably doesn't violate (C) because (1) reverse engineering is fair use, (2) outputs are not (according to the (C) office) copyrightable, (3) to the extent they are, the labs' terms of service typically assign that (C) to the user, and (4) to the extent their TOS' attempt to protect an intellectual property interest that doesn't exist they may be preempted by existing IP law (including trade secrecy law, where there is not likely any trade secret implicated because anyone can elicit these outputs from models released to the public). The labs objections are also, as widely pointed out, incredibly hypocritical, and as less widely pointed out, rather inconsistent with their fair use arguments in cases about their scraping data for AI training (which is why they haven't yet articulated a legal theory about how what China's doing is illegal, either.)
@KevinBankston Please don’t conflate things. One goes to public library and read books, wants to write better codes/poems by learning. The other(Kimi) don’t even want to learn and keep asking teachers how to write this/that and iterate without paying the teacher. Not talking IP here.
@dangaron Certainly a TOS violation but that’s not IP law. Nor is CFAA (and if this is a CFAA violation an enormous amount of other common online behavior, including by the labs and their scraping vendors, is too.)
@dangaron @mkratsios47 That particular tweet did not but it’s part of the usual set of talking points from the administration here, see eg
They’ve repeatedly called it “IP theft” of “proprietary information.” Arguable contract claim but also good argument that such claim would be preempted by IP law, see eg https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5049562. And if this violates CFAA then there are a whole lot of felons out there whose only crime was using a proxy or changing their browser user agent string. If they want to make an issue of it with China as a negotiable issue of industrial policy, fine. I’m just pushing back on vague intimations of illegality when they refuse to bring receipts.
@KevinBankston Presumably a TOS violation (a breach of contract) + perhaps a CFAA violation too
@adamconner http://e.gov/wp-content/uploads/2026/04/NSTM-4.pdf Repeats claims that the outputs being distilled are “proprietary”
@KevinBankston “According to your (C) office, LLM outputs aren’t copyrightable” That is not what the copyright office has said, which is more nuanced, and the argument here is more about trade secrets protected by contractual agreements around model access, not copyright on outputs anyway.
Yes. That’s a trade secret. But in this case the purported trade secret is freely available to anyone online to prompt, and unlike with an NDA there are no restrictions in the TOS on sharing the supposedly trade secreted outputs—in fact it assigns you copyright in them—just restrictions automated scraping. Hence, not a trade secret.
@KevinBankston Didn't anthropic this week just get a 1.5B judgement for using stolen copyrighted books (screwing over american authors and companies) and it was the LARGEST IP THEFT JUDGEMENT IN US HISTORY? How does nobody see the fucking irony lol
@dangaron @mkratsios47 And the Treasury Secretary yesterday: https://thehill.com/policy/technology/5980722-scott-bessent-china-sanctions-ai-theft/amp/ They keep saying the words but not articulating a legal theory
When a user signs up for Anthropic’s API (creating an account, obtaining API keys, and paying for usage), they enter into Anthropic’s Commercial Terms of Service. Those terms expressly form a binding agreement Even for the free version of http://Claude.ai (the free consumer tier), Anthropic’s Consumer Terms of Service form a binding civil contract. No one forced you to use their service. If and when you did, you agreed to the basic legal tenets of a voluntary agreement as the basis for binding obligations, which dates back 2000 years to the Roman Empire.
@tjl Then they should actually articulate that claim
@tjl As noted, if they want to make it a diplomatic issue more power to them, I’m questioning the continued intimation of illegality
@KevinBankston Is there an actual memo?
@KevinBankston If a US company did this, the claim would be tortious breach of contract for the purpose of misappropriating trade secrets
@KevinBankston Consider an analogy: I get on the phone with a customer. I’m under NDA. They make a series of purely factual statements which I use to make a copy of one of their secret inventions. I’m in breach and stole IP without the conveying information having (c)
@KevinBankston Doesn’t have to be an IP claim—I don’t think they specified, right? Torts, breach. I would expect potential gov support for a CFAA claim. And who knows what claims exist under the treaty system
If you're interested in this issue I wrote about it a few weeks ago in my newsletter, Converger: https://converger.kevinbankston.com/i/196771089/model-hypocrisy-unconsented-data-scraping-for-me-but-not-for-thee