Live data from Hacker News

X changes its terms to bar training of AI models using its content

techcrunch.com

21–30 of 217 posts

Re: X changes its terms to bar training of AI models using its content

#21
post #18

There needs to be a worldwide standard, such as an HTML tag, that says "no training". And a few countries need to make it a punishable offense to violate the tag. The punishment should be exceptionally severe, not just a fine. For example: any company that violates the tag should be completely barred from operating, forever.

That will play out exactly like the "Do not track" bit did.

Perhaps we should try anyway, in case you are wrong.

Re: X changes its terms to bar training of AI models using its content

#22
post #5

Weird this just happened. I assumed all sites with any sort of content changed their terms soon after ChatGPT hit the scene.

Yep, from https://the-decoder.com/reddit-ends-its-role-as-a-free-ai-tr... :

You must not, and must not allow those acting on your behalf to:

...use the Data APIs to encourage or promote illegal activity or violation of third party rights (including using User Content to train a machine learning or AI model without the express permission of rightsholders in the applicable User Content);

Re: X changes its terms to bar training of AI models using its content

#23
post #22
post #5

Weird this just happened. I assumed all sites with any sort of content changed their terms soon after ChatGPT hit the scene.

Yep, from https://the-decoder.com/reddit-ends-its-role-as-a-free-ai-tr... : You must not, and must not allow those acting on your behalf to: ...use the Data APIs to encourage or promote illegal activity or violation of third party rights (including using User Content to train a machine learning or AI model without the express permission of rightsholders in the applicable User Content);

In my eyes that is considered fair use, and I think the courts will come to agree unless they are financially incentivized to look the other way and thus create a moat for existing players at the expense of newcomers.

Re: X changes its terms to bar training of AI models using its content

#24
post #9
post #2

If an artist or author can't do this, social media shouldn't be able to do it either. If Xai wants to train on public corpus, it shouldn't be allowed to prevent its own corpus from being used. We need regulations to limit the power grabs. Train all you like, but don't dare try to constrain to your walled gardens. We should also probably nip the "foundation model company / also a social media company" conglomeration i…

> If an artist or author can't do this, social media shouldn't be able to do it either. Even if this is done, the case of starving artist v. megacorp will probably go to whoever wields the most money and lawyers. To add insult to injury, the artist’s opponent is fueled by their ill-gotten gains.

This is dependent on country. USA, yes with their draconian methods. Countries like the UK, the looser of the suit pays all the cost. UK layers have no problem taking low wealth client cases they know will win. UK allows for David vs Goliath and David to win. US up lifts Goliath as a God.

Re: X changes its terms to bar training of AI models using its content

#25

There needs to be a worldwide standard, such as an HTML tag, that says "no training". And a few countries need to make it a punishable offense to violate the tag. The punishment should be exceptionally severe, not just a fine. For example: any company that violates the tag should be completely barred from operating, forever.

>There needs to be a worldwide standard, such as an HTML tag, that says "no training"

Any country that seriously implemented this would just end up being completely dominated by the autonomous robot soldiers of another country that didn't, because it effectively bans the development of embodied AGI (which can learn live from seeing/reading something, like a human can).

Re: X changes its terms to bar training of AI models using its content

#26
post #4
post #2

If an artist or author can't do this, social media shouldn't be able to do it either. If Xai wants to train on public corpus, it shouldn't be allowed to prevent its own corpus from being used. We need regulations to limit the power grabs. Train all you like, but don't dare try to constrain to your walled gardens. We should also probably nip the "foundation model company / also a social media company" conglomeration i…

Artists can do this, and they do

Yes, but do artists have the ability to actually monitor and enforce this? You have to have the capacity and the wherewithal and to test these models to even know that your data is being ingested into AI.

Big companies like the New York Times and Twitter/X have the funds to pay for this. Miscellaneous artists probably don't.

Re: X changes its terms to bar training of AI models using its content

#27
How useful is low-quality content like Youtube comments and tweets anyway? Is it a common/important use case to generate tweet-length, tweet-quality content? Are most use cases of generating tweet-type content spam/fraud? Would a model be better off if it was unable to perform those use cases?

Re: X changes its terms to bar training of AI models using its content

#28
post #10

wish I could change my terms to bar training of AI models on my content

Same here! It should be a default. Unfortunately, the very openness of the internet is now working against us.

Why should it be a default? Can you prove that training a model on data you wrote is not fair use?

We're already seeing precedent that it might be.

https://www.ecjlaw.com/ecj-blog/kadrey-v-meta-the-first-majo...

The openness of the internet is a good thing, but it doesn't come without a cost. And the moment we have to pay that cost, we don't get to suddenly go, "well, openness turned out to be a mistake, let's close it all up and create a regulatory, bureaucratic nightmare". This is the tradeoff. Freedom for me, and thee.

Re: X changes its terms to bar training of AI models using its content

#29
post #9

Earlier quoted context omitted.

> If an artist or author can't do this, social media shouldn't be able to do it either. Even if this is done, the case of starving artist v. megacorp will probably go to whoever wields the most money and lawyers. To add insult to injury, the artist’s opponent is fueled by their ill-gotten gains.

This is dependent on country. USA, yes with their draconian methods. Countries like the UK, the looser of the suit pays all the cost. UK layers have no problem taking low wealth client cases they know will win. UK allows for David vs Goliath and David to win. US up lifts Goliath as a God.

Also in many countries legal costs are just generally lower than in the US.

Re: X changes its terms to bar training of AI models using its content

#30

There needs to be a worldwide standard, such as an HTML tag, that says "no training". And a few countries need to make it a punishable offense to violate the tag. The punishment should be exceptionally severe, not just a fine. For example: any company that violates the tag should be completely barred from operating, forever.

That will just lead to situations where one company scrapes the site, cleans the content of tags, and sells the data, and another does the training on the precleaned data. The first one hasn't trained and the second one never saw the tag.

Companies who are found guilty of this should also be rendered bankrupt then.
Post reply on HN