Live data from Hacker News

X changes its terms to bar training of AI models using its content

techcrunch.com

91–100 of 217 posts

Re: X changes its terms to bar training of AI models using its content

#91
post #81

VAT for content should be a thing. Ultimately all users should be getting paid

We really need LLMs for music to become more advanced.

Then maybe the recording companies will start defending artist rights.

Because not sure what all the other industry bodies are doing.

Re: X changes its terms to bar training of AI models using its content

#92
post #74

Right, stealing training data from others is OK, having it stolen from you is not. What else is new?

Almost certainly the easter egg found in the Trump "Big Beautiful Bill" which prevents states from enacting AI regulations also came from Musk.

That way he can continue to steal from others and lock competitors out whilst being comfortable knowing that no laws will be enacted to prevent it.

Re: X changes its terms to bar training of AI models using its content

#93
post #74

Right, stealing training data from others is OK, having it stolen from you is not. What else is new?

Almost certainly the easter egg found in the Trump "Big Beautiful Bill" which prevents states from enacting AI regulations also came from Musk. That way he can continue to steal from others and lock competitors out whilst being comfortable knowing that no laws will be enacted to prevent it.

why do you think he is so evil but all others are benign ?

Re: X changes its terms to bar training of AI models using its content

#94

Who's training an AI on the "Tweet" button text? Or are they trying to forgo section 230 protection and claim ownership of content uploaded to the site?

Perhaps they want the prohibition on using the site content for AI training to be considered based on something other than their ownership of it, like bandwidth usage or users' rights

Re: X changes its terms to bar training of AI models using its content

#95

Earlier quoted context omitted.

So I get to use the platform for free, but I also get paid to post on the platform? I'm not sure that makes sense. Like I hate to take the side of big tech, but they can't literally be paying users to use their platform. Just use something else, there are a million social media sites

> I get to use the platform for free You actually get to generate content for the platform for free. Without you (all of the X users), the platform would be devoid of content, just botspeak and corporate promos. Plus, as the sibling mentioned, they monetize your visit through ads (and data use).

Most posts are ignored and are an absolute loss to the company. Which is why platforms like Twitter only allow you to make money from posting once you reach a certain threshold.

Re: X changes its terms to bar training of AI models using its content

#96
Copyright is not going well. The rights of millions of people are trampled by companies, both the content we post on social networks and our private AI chats. Our voice doesn't matter.

Copyright was supposed to protect expression and keep ideas freely circulating. But now it protects abstractions (see the Abstraction-Filtration-Comparison test). It is much more difficult to be sure you are not infringing.

Re: X changes its terms to bar training of AI models using its content

#99
post #81

VAT for content should be a thing. Ultimately all users should be getting paid

I wanted to do some quick math on this idea- supposed we trained a vanilla transformer model from scratch, as GPT2/GPT3 was done- the number of seen input tokens is known perfectly, as is the sources of those training tokens (since then, everyone has either kept quiet about the sources post-Books3-fiasco, or have been finetuning on top of previous models making this more difficult of a calculation)

GPT-3 was trained on approximately 300 billion tokens. An small sized technical textbook might contain something like... 130,000 tokens? (1 token ~= 0.75 words, ~100k words in the book).

Thus, say you wrote a textbook on quantum mechanics that was included in the training corpus. A naive computation of the fraction of your textbook's contribution to the total number of training tokens would be 300B/130K = 0.0000004333333333, or 0.000043%.

If our hypothetical AI company here reported, say $500M in yearly profit, if all of that was distributed 100% based on our naive training token ratio (notice I say naive because it isn't as simple to say that every training token contributes equally to the final weights of a model. That is part of the magic.) then $500M * 0.000043% = $215.

You could imagine a simpler world where it was required by law that any such profitable company redistribute, say, %20 (taking the 'anti-VAT' idea) back to the copyright holders / originators of the training tokens. So, our fictitious QM textbook author would receive a check in the mail for $43 for that year of $500M in revenue. Not great, but not zero.

Since then, training corpuses are much, much larger, and most people's contributions would be much smaller. Someone who writes witty tweets? Maybe 1/100th the length of our above example in am model with now 100x the training corpus.

So fractions of a penny for your tweets. Maybe that is fitting after all...

Re: X changes its terms to bar training of AI models using its content

#100

Earlier quoted context omitted.

Same here! It should be a default. Unfortunately, the very openness of the internet is now working against us.

Why should it be a default? Can you prove that training a model on data you wrote is not fair use? We're already seeing precedent that it might be. https://www.ecjlaw.com/ecj-blog/kadrey-v-meta-the-first-majo... The openness of the internet is a good thing, but it doesn't come without a cost. And the moment we have to pay that cost, we don't get to suddenly go, "well, openness turned out to be a mistake, let's close…

Yeah, I don't think downloading my paid-for books, from an illegal sharing site, to scrape and make use of, is in any way fair use.

From the decision in 1841, in the US (Folsom vs Marsh):

> reviewer may fairly cite largely from the original work, if his design be really and truly to use the passages for the purposes of fair and reasonable criticism. On the other hand, it is as clear, that if he thus cites the most important parts of the work, with a view, not to criticize, but to supersede the use of the original work, and substitute the review for it, such a use will be deemed in law a piracy

Further, to be "transformative", it is required that the new work is for a new purpose. It has to be done in such a way that it basically is not competing with the original at all.

Using my creative works, to create creative works, is rather clearly an act of piracy. And the methods engaged, to enable to do so, are also clearly piracy.

Where would training a model here, possibly be fair use?

Post reply on HN