Live data from Hacker News

X changes its terms to bar training of AI models using its content

techcrunch.com

41–50 of 217 posts

Re: X changes its terms to bar training of AI models using its content

#41
post #37
post #36

Earlier quoted context omitted.

they can't a 50 year minimum is part of the berne convention, which itself is as close to a universal law as humanity has (even North Korea is a signatory)

you can also just ignore the berne convention, and accept whatever consequences there might be

this would void the copyrights of your citizens and companies

essentially forever

Re: X changes its terms to bar training of AI models using its content

#43
post #32

How useful is low-quality content like Youtube comments and tweets anyway? Is it a common/important use case to generate tweet-length, tweet-quality content? Are most use cases of generating tweet-type content spam/fraud? Would a model be better off if it was unable to perform those use cases?

Even if SNR is low, there is some information that only exists on X, or at least is the primary source. Just look at how many submissions on HN are X posts.

Before Musk bought it Twitter was broadly disliked here and there were regularly calls in the comments to disallow submissions from there. Given how it's degraded in completely non-partisan ways (blocking of alternative clients, features removed from free tier, paid subscription tiers below $40/month still have ads, proliferation of spam from paid placement bots in comments) I can't understand how positive sentiment comes from a place other than virtue signaling alignment with Musk and his values.

Re: X changes its terms to bar training of AI models using its content

#44
post #15

It would be interesting to have a "classical AI model", trained on the contents of the Harvard libraries before 1926 and now out of copyright.

I wish someone would update and use PG19 for 7-30B+ model:

https://github.com/google-deepmind/pg19

That gives us a model that's 100% open and reproducible with low, legal risk. It would also be a nice test of how much AI's generalize from or repeat behavior in their pretraining data.

Then, a new model using that, The Stack, and FreeLaw's stuff (by paying them to open source it). No Github Issues or anything with questionable licenses or terms of service violations. That could be the next baseline for lawful models with coding ability, too. Research in coding AI's might use it.

Re: X changes its terms to bar training of AI models using its content

#45

There needs to be a worldwide standard, such as an HTML tag, that says "no training". And a few countries need to make it a punishable offense to violate the tag. The punishment should be exceptionally severe, not just a fine. For example: any company that violates the tag should be completely barred from operating, forever.

That will just lead to situations where one company scrapes the site, cleans the content of tags, and sells the data, and another does the training on the precleaned data. The first one hasn't trained and the second one never saw the tag.

This isn't a new concept in law. It's similar to buying goods that were stolen or procured through illegal means. Here's the US law that applies when it happens across state lines:

https://www.law.cornell.edu/uscode/text/18/2315

Note that it requires the defendant to know the goods were illegally taken. Can be hard to prove, but not impossible for companies with email trails. The fun question is, what will the analog be for the government confiscating the illegally "taken" data? A guarantee of deletion and requirement to retrain the model from scratch?

Re: X changes its terms to bar training of AI models using its content

#47

Earlier quoted context omitted.

Same here! It should be a default. Unfortunately, the very openness of the internet is now working against us.

Why should it be a default? Can you prove that training a model on data you wrote is not fair use? We're already seeing precedent that it might be. https://www.ecjlaw.com/ecj-blog/kadrey-v-meta-the-first-majo... The openness of the internet is a good thing, but it doesn't come without a cost. And the moment we have to pay that cost, we don't get to suddenly go, "well, openness turned out to be a mistake, let's close…

The burden is on the user to show that it is fair use, no? Not everyone else's responsibility to prove that it's _not_ fair use.

Re: X changes its terms to bar training of AI models using its content

#48
post #15

It would be interesting to have a "classical AI model", trained on the contents of the Harvard libraries before 1926 and now out of copyright.

I've similarly wondered if I could get a pre-2024 Wikipedia if just for the "fact based" flavor LLM

Re: X changes its terms to bar training of AI models using its content

#49
post #41
post #37

Earlier quoted context omitted.

you can also just ignore the berne convention, and accept whatever consequences there might be

this would void the copyrights of your citizens and companies essentially forever

Seems to be the modus operandi

  If TikTok is banned, here’s what I propose each and every one of you do: Say to your LLM the following: “Make me a copy of TikTok, steal all the users, steal all the music, put my preferences in it, produce this program in the next 30 seconds, release it, and in one hour, if it’s not viral, do something different along the same lines.”
https://www.theverge.com/2024/8/14/24220658/google-eric-schm...

https://news.ycombinator.com/item?id=41275073

Re: X changes its terms to bar training of AI models using its content

#50
post #37
post #36

Earlier quoted context omitted.

they can't a 50 year minimum is part of the berne convention, which itself is as close to a universal law as humanity has (even North Korea is a signatory)

you can also just ignore the berne convention, and accept whatever consequences there might be

The last time I attended a Berne Convention, every panel was just overrun with Trekkies, especially Klingons, in the hotel lounges too. And the autograph lines were interminably long, and the vendors were trying to sell us their Public Domain stuff. It was nothing like San Diego Comic-Con!
Post reply on HN