Live data from Hacker News

Anthropic says Alibaba illicitly extracted Claude AI model capabilities

reuters.com

621–630 of 1001 posts

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#621
post #567

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

In a different world it is not fair use. The benefits of the crime should be always taken off. If you isolate the training and pirating, you may say that it was fair, but that completely misses the point. The sole purpose of pirating (aka crime) was to train the models.

Copyright infringement isn't usually a crime.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#622

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

> That's why Anthropic switched to scanning paper books. After they threw away all the tainted data from the pirated books, right?

No, because the judge ruled that the training was fair use and the model itself wasn't infringing.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#623
post #534

I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting Anthropic's "research purpose" defense: "You can't just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want." https://www.joneswalker.com/en/insights/blogs/ai-law-blog/wh...

Exactly. Couldn't happen to better people. I'm pretty against piracy personally but if we find reliable ways to pirate Anthropic/OpenAI products in the future I'm all for it.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#624

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

Isn't scanning also a form of copyright infringement? You are making a digital copy of a book, which is the same thing as downloading a book from the internet...

I think that we can run a perhaps silly thought experiment.

Suppose that I have a nearly perfect memory and I could remember all the books I read. Suppose also that I have a million year life span so I could read 7 million books. Then, what happens if at the end of all of those years, or at any earlier moment I answer questions from people and I exploit commercially the knowledge I gathered reading those books? Would my reading those books be study or copyright infringement? Remember the nearly perfect memory hypotheses.

Of course it's a bit silly because the time to train a LLM and the time I need to read all those books is different by orders of magnitude and that changes the perspective. Who would complain with me today if their heirs lose some money on 7 million AD? Who would even notice that I started that million years long endeavor. Who's going to be there to ask me questions by then? Humans? Birds? Lizards? And I can say that I am studying like everybody else before me, but does an LLM study? And I am sure there are many other nuances.

Anyway, I don't think that scanning is any different than photons hitting my retina. The difference is in what happens next: the faithfulness of memory, the amount of knowledge, the speed of accumulating it. After all a huge amount of quantity can become quality.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#625

Earlier quoted context omitted.

Im ok with this! Is there a site that list all these resellers, or better, a openrouter-like for these resellers?

They're called 中转站 (transfer stations/proxies). They can be a bit tricky to find on your own, so I'd suggest asking your preferred AI to search in Mandarin for you. I linked a larger operator in the parent comment, or have a look at https://hvoy.ai/ which lists a ton. You can also find many on Funpay, which may be easier to use. This is one seller I found, they're reselling "real Max 20x subscription accounts", at ~9…

How did you even find these? Even in discussions about cheap AI I've never seen anyone mention this. Great find!

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#626

Earlier quoted context omitted.

Using them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.

If using the books is fair use, then distilling the model, which is just a derived product of those books is also fair use. These companies are trying to have their cake and eat it too.

Hmm, training on a book’s text smears the content all over the weights, merging it with all other texts. The original text isn’t intentionally supposed to be reproducible in any larger part (although IIRC models were able to emit fairly large chunks verbatim).

Quite unlikely, training on behavior purportedly approximately replicates the behavior. It gets replicated intentionally as a whole.

IANAL, but I see significant differences with intent to copy a significant part as a whole into a competing product, surely shouldn’t fit under legal concept of fair use, no matter whether scanning books for LLM training fits or not.

Whether such things (behaviors) are copyrightable - and should they be so - is another interesting question. Those aren’t algorithms or databases (stuff clearly and explicitly covered in many copyright laws), those are human expectation models, something like how we train animals or teach our own.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#627
post #472

Earlier quoted context omitted.

Stupid question: I was under the impression that these models were trained on PB of data. Surely the amount of questions/response they can extract from querying a bigger model (Claude) is fairly modest. How is it not a drop vs the training dataset?

It's not about how big your dataset is - it's about how you use it. I jest, but I'm also completely serious. 1T tokens from Claude can teach a model something 1T tokens scraped from the open web can't. Things like "how an LLM can problem solve effectively", or "how an LLM should use tools", or "how to construct reasoning chains", or "when to double check", or "what innate capabilities an LLM can or can't rely on". Th…

Unremarkable base model will remain an unremarkable fine-tuned model that memorised a couple thousand of input-output pairings.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#628

Earlier quoted context omitted.

Reuters is probably the most rigorous news agency in the world. > it said was the largest known attack > Anthropic said in the letter it was supportive of the U.S. government's efforts to combat the attacks both times the word "attack" appears it's clearly stated that the word was used by the company, it's a direct company quote. actually putting it into quotes would be editorializing > Unfortunately, the Reuters pie…

Well, let’s say you put the picture of some political figure, and put in highly contrasted red, bold large catchy font, "TERRORIST THAT KILLED MILLION PEOPLE", then below that in barely visible contrast, in tiny discrete letters, "is what this person probably will claim to be against". This whole sentence technically will be correct, 100% guarantee, whatever this person actually even said or think. From a propaganda…

nice slippery slope you manufactured there - what if Reuters becomes Daily Mail

what framing are you talking about? they are literally quoting a company.

please explain what Reuters should have done here. Should they have added in parentheses: (editor note: we don't agree with Anthropic calling this an "attack")

Is that what you want? News outlets giving their opinion and moral judgement on company quotes? I mean, Fox News/CNN do have a large following, so there is clearly a market for that.

Re: Anthropic says Alibaba illicitly extracted Claude AI model capabilities

#630

Earlier quoted context omitted.

Apple gave Xerox the right to buy $1 million of pre-IPO stock before the meeting took place.

Glad you pointed this out. I believe the sequence was that Jobs himself got a shorter demo during his first visit with no prior arrangements. He then negotiated bringing back a group of his key people to get a more in depth demo and that included the stock deal. When Apple was accused of 'ripping off' PARC, Steve didn't seem keen to bring up this rather salient point. I suspect it may have been a combination of wanti…

> the million dollar stock deal could seem a bit like trading beads to Native Americans for Manhattan Island

But in both cases the value only existed because of the people offering the deal. XeroX doing nothing with a UI or native Americans doing nothing with some land would mean the UI and the land would continue to be worth nothing. It was the others coming with ideas and effort that made them valuable.

Post reply on HN