Live data from Hacker News

Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

variety.com

201–210 of 484 posts

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#201

Funny how people are suddenly on Elsevier's side. It's clear to me that AI training is transformative fair use under existing law. Maybe this will be the case to prove it.

I also find it funny, I said this regarding the other thread and article[0]

'"They then copied those stolen fruits"

How are these fruits "stolen" if they still have what was allegedley stolen?

Dowling v. United States, 473 U.S. 207 (1985): The Supreme Court ruled that the unauthorized sale of phonorecords of copyrighted musical compositions does not constitute "stolen, converted or taken by fraud" goods under the National Stolen Property Act

And even if, arguendo, sure its stolen. The purpose of copyright is to "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries"

And you would be hard pressed to prove that LLM's haven't advanced the arts and sciences, so at bare minimum transformative, ie fair use.'

[0] https://news.ycombinator.com/item?id=48026207#48029072

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#202

Funny how people are suddenly on Elsevier's side. It's clear to me that AI training is transformative fair use under existing law. Maybe this will be the case to prove it.

Illegally obtaining copyrighted materials is usually the issue not the transformation part

Looking at the complaint ( https://publishers.org/wp-content/uploads/2026/05/2026-05-05... ), that seems like the part that's got the most solid foundation, especially given that while torrenting the books, they were also seeding to other peers.

The items they call out around training the models (and attempting to claim that each subsequent model generation should count as an additional instance of infringement) seem far less grounded in the current court interpretations of AI training.

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#203

Earlier quoted context omitted.

they'll litigate how meta acquired those materials to train. you can do whatever you want with a book after it's in your house. but how did it get there?

They’re already on record as hoovering up Library Genesis and Anna’s Archive. For their “fair use” copyright bonfire to train their LLM. So not are these publishers rightfully pissed, Meta didn’t even give them the $6.99 for each epub to begin with. They’ve stolen the whole thing as part of this “fair use” campaign to destroy human authorship free of even the most basic remuneration.

Fun fact, if you link AA on FB it gets removed

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#204

Earlier quoted context omitted.

IMO ASN-based blocking should be much more common, but unfortunately it is not supported as a first-class configuration option in many common tools.

Yeah, I dont know how anybody stays sane without it. I have a list of over a thousand ASNs I blackhole at this point... Mine is a daily bash cronjob that fetches a text-based database and uses grep to build an nftables-apply script with all the IPs for the blocked ASNs. I keep meaning to share it, but it's embarrassingly messy I haven't had time to clean it up...

It's been a real game of cat and mouse over the last few years. I used to do daily iptables updates to block repeat scrapers on my small niche stats site I run. About 5-6 ago it become more common to see broader ranges - so I started blocking ASNs which worked great (esp for the regulars like Alibaba, Tencent, compromised DigitalOcean/OVH, ...). In the last 2-3 years though the overall bot traffic has skyrocketed - it's easy to spot bot activity after the fact (no requests to the CDN for static assets, user agent changes from one request to the next, predictable ID enumeration, etc) but not in a real time. They're also often using residential-based proxies and Cloudflare bot detection has become pretty bad.

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#205
post #2

A lot of people would be very pleased if this leads to Zuckerberg getting even the statutory minimum damages ($750?) on each infringement. The previous infringement case with Anthropic said that while training an AI was transformative and not itself an infringement, pirating works for that purpose still was definitely infringement all by itself. The settlement was $1.5bn, so close to $3k for each of the 500k they pir…

What's frustrating is all those kids who got criminal charges for running MP3 sites back in the day [1], and this guy rips off every piece of media in existence and will walk away literally because he's too rich to be charged. [1] See, e.g. https://en.wikipedia.org/wiki/Oink%27s_Pink_Palace#Legal_pro...

What's frustrating is that I don't even consider infringement to be a crime. Why are you all so upset about this, rather than his real crimes?

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#206

Funny how people are suddenly on Elsevier's side. It's clear to me that AI training is transformative fair use under existing law. Maybe this will be the case to prove it.

> It's clear to me that AI training is transformative fair use under existing law. Maybe this will be the case to prove it.

That is not what this case is about. It is more about the illegal violation and piracy of copyrighted content done by Meta for commercial use and Zuck knew they were doing it.

Why did Anthropic settle [0] with a multi-billion dollar payout to authors after commercializing their LLMs that was trained off of copyrighted content that was illegally obtained and kept without the authors permission?

There's a reason why they (Anthropic) did not want it to go to trial. (Anthropic knew they would lose and it would completely bankrupt them in the hundreds of billions.)

AI boosters will do anything to justify the mass piracy and illegal obtainment of copyrighted material for commercial use (not research) which that is not fair use in the US. There is no debate on this. [0]

[0] https://images.assettype.com/theleaflet/2025-09-27/mnuaifvw/...

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#207
post #35

Earlier quoted context omitted.

Yes, I have no objection to that part. It's the arguments that training itself is the problem. Sarah Silverman as the most prominent example.

I mean the act of reproducing the copyrighted material is what is illegal. LLMs I've used for coding has outputted exact copyrights for code verbatim into my code before. When that happens it feels kind of fishy to be honest.

Yes. I agree. But many people argue that training itself is a copyright violation. That's the position I'm countering here.

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#208

Funny how people are suddenly on Elsevier's side. It's clear to me that AI training is transformative fair use under existing law. Maybe this will be the case to prove it.

If i could ask for a summary from an llm vs buy a book id go with the summary. That eats into commercial use and the supreme court case sided with Gerald Ford when a newspaper published a small gist of his autobiography because it ate into the sales

Re: Zuckerberg 'Personally Authorized and Encouraged' Meta's Copyright Infringement

#210
post #122

Earlier quoted context omitted.

But why should separate licensing be required at all? A search engine reads and indexes every word of every page it crawls. No one argues that requires licensing, only that the outputs must respect copyright. Why should training be different?

When google starting outputting summaries people asked the same questions. If you supplant the value of the original with the original as input then you probably have some legal questions to answer.

But that's about the output, not the training. We agree: outputs that supplant the original are the problem. A model constrained to produce only fair use outputs causes no such harm — regardless of what it was trained on.
Post reply on HN