Earlier quoted context omitted.
I haven’t seen any accusations that they’ve done that, though. Usually people get pirated material from sources that intentionally share pirated material.
They're not just training on pirated content, they've also scraped literally the entire internet and used that too.
Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
301–310 of 354 posts
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#302Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#303Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#304Earlier quoted context omitted.
At the risk of stepping on a well-known land mine around here, how'd you do on the IMO problem set this year?
I didn't participate. I probably wouldn't have done well. I disagree with your framing.
Either both AI teams cheated, in which case there's nothing to worry about, or they didn't, in which case you've set a pretty high bar. Where is that bar, exactly? What exactly does it take to justify blowing off copyright law in the larger interest of progress? (I have my own answers to that question, including equitable access to the resulting models regardless of how impressive their performance might be, but am curious to hear yours.)
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#305Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#306Earlier quoted context omitted.
Who's to say why I downloaded and am now watching a movie? Is it for my enjoyment? Is it because I'm training my brain? How is me training my brain any different from companies training their LLMs? Same goes for recording: I'm just training my skills of recording. Or maybe I'm just recording it so I can rewatch it later, for training purposes, of course.
>Who's to say why I downloaded and am now watching a movie? Is it for my enjoyment? Is it because I'm training my brain? How is me training my brain any different from companies training their LLMs? None of this is relevant because Anthropic was only left off the hook for training, and not for pirating the books itself. So far as the court cases are playing out, there doesn't appear to be a special piracy exemption f…
Good example, because this is exactly what websites are doing with LLM companies, who are doing their damnest to evade the blocks. Which brings us back around to "trespassing" or the CFAA or whatever.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#307Earlier quoted context omitted.
> Besides, if it really bothered you, then we might not see this weird tone-switch from one sentence to the next, where you seem to think that piracy is shocking and "something should be done" and then "it's not good tht someone should spend time in jail for it". What gives? What a weirdly condescending way to interpret my post. My point boils down to: Either prosecute copyright infringement or don't. The current sta…
> Either prosecute copyright infringement or don't This is the absolute core of the issue. Technical people see law as code, where context can be disregarded and all that matters is specifying the outputs for a given set of inputs. But law doesn’t work that way, and it should not work that way. Context matters, and it needs to. If you go down the road of “the law is the law and billion dollar companies working on pro…
Copyright laws target everyone. SEC laws don't.
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#308Earlier quoted context omitted.
They're not just training on pirated content, they've also scraped literally the entire internet and used that too.
Scraping the public internet is also not a CFAA violation
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#309Earlier quoted context omitted.
>Who's to say why I downloaded and am now watching a movie? Is it for my enjoyment? Is it because I'm training my brain? How is me training my brain any different from companies training their LLMs? None of this is relevant because Anthropic was only left off the hook for training, and not for pirating the books itself. So far as the court cases are playing out, there doesn't appear to be a special piracy exemption f…
> If you're caught with AV gear in a movie theater once, you'd likely be ejected and banned from the establishment/chain, not have the FBI/MPAA go after you for piracy Good example, because this is exactly what websites are doing with LLM companies, who are doing their damnest to evade the blocks. Which brings us back around to "trespassing" or the CFAA or whatever.
That argument is pretty much dead after https://en.wikipedia.org/wiki/Van_Buren_v._United_States and https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
Re: Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
#310Earlier quoted context omitted.
Scraping the public internet is also not a CFAA violation
CFAA bans accessing a protected computer without authorization. Hitting URLs denied by robots.txt has been argued to be just that.
"Has been argued" -- sure, but never successfully; in fact, in HiQ v. LinkedIn, the 9th Circuit ruled (twice, both before and on remand again after and applying the Supreme Court ruling in Van Buren v. US) against a cease and desist on top of robots.txt to stop accessing data on a public website constituting "without authorization" under the CFAA.