Live data from Hacker News

X changes its terms to bar training of AI models using its content

techcrunch.com

181–190 of 217 posts

Re: X changes its terms to bar training of AI models using its content

#181

Earlier quoted context omitted.

You've got it backwards. It's on the defendant to prove that their use is fair. The plaintiff has to prove that they actually own the copyright, and that it covers the work they're claiming was infringed, and may try to refute any fair-use arguments the defense raises, but if the defense doesn't raise any then the use won't be found fair.

It's true that the process is copyright strike/lawsuit -> appeal, but like I said, it's in their best interests to just prove that it's fair use because otherwise the judge might not properly consider all facts, only hear one side of the story and thus make a bad judgement about whether or not it is fair use. If anything, I'm just being pedantic, but we do ultimately agree here I think.

Well, lawsuits have multiple stages. First the plaintiff files the suit, and serves notice to the defendant(s) that the suit has been filed. Then there's a period where both sides gather evidence (discovery), then there's a trial where they present their evidence & arguments to the court. Each side gets time to respond to the arguments made by the opposing party. Then a verdict is chosen, and any penalties are decided by the court. So there's not really any chance the judge only hears one side of the story.

That said, I think we do agree. The plaintiff should be prepared to refute a fair-use argument raised by the defendant. I'm just noting that the refutation doesn't need to be part of the initial filing, it gets presented at trial, after discovery, and only if the defendant presents a fair-use defense. So they don't have to prove it's not fair use to win in every case. I'm probably also being excessively pedantic!

Re: X changes its terms to bar training of AI models using its content

#183

Earlier quoted context omitted.

if that is any consolation, no one gives a shit about xitter's ToS either. it will continue to be scrapped by every major player.

How exactly is it being scraped? My understanding is Twitter and LinkedIn are both huge pains in the ass to scrape right now.

There's a number of companies out there, like "brightdata", which pay a small amount to app developers to install a native "sdk". That SDK mimics a browser, and makes requests as if the user's device is doing it.

Since it's using a large number of real user's devices, and closely mimicing real web browsers, it ends up looking incredibly similar to real user traffic.

Since twitter allows some amount of anonymous browsing, that's enough to get some amount of data out. You can also pay brightdata for one large aggregated dataset.

https://bright-sdk.com/

This is part of the AI revolution, user's devices being commandeered to DDoS small blogs and twitter alike to feed data to the beast.

Re: X changes its terms to bar training of AI models using its content

#184
post #142
post #80

Earlier quoted context omitted.

Linux is generally a functional tool, and struggles with overall coherence. There are far fewer success stories of artworks being made in this style. (E.g. there are successful multiplayer open-source games or clones of existing games, but very few original single-player games, and those that there are are largely the work of a single individual)

Linux is both a kernel (which is under GPL), and an operating system, whose other components are under a variety of licenses (and you can pick and match which components you want). That's why some people like to call it 'Gnu/Linux', but thanks to recent advances we can make Gnu-free Linuxes today, too. > There are far fewer success stories of artworks being made in this style. (E.g. there are successful multiplayer o…

> Linux is both a kernel (which is under GPL), and an operating system

I was talking about the kernel, though what I said applies to both.

> Humans have made art since forever.

Perhaps, but not the kind of long-form narrative experiences that we're talking about here. (Sagas and epics predate copyright, but those are a quite different form, and indeed have much the same downsides - struggles with coherence and consistency when there are multiple authors, inability to put everything together in a sensible arc).

Re: X changes its terms to bar training of AI models using its content

#185

Earlier quoted context omitted.

Can you expand on that?as written I can't make much sense of your comment. FTR I don't think he thought he could save us, I think he thought he could do cool stuff (space and EVs) and now says climate change isn't as bad as he used to think (despite mountains of evidence to the contrary).

If you repeat your lies enough, you'll end up believing them yourself. Especially if you surround yourself with yes-men. It's entirely possible Musk genuinely believed he was a savior of humanity.

I think there's a long way between what you're saying (which is true particularly as there seem to be a lot of thin skinned leaders of the tech industry) and ending up being responsible for teenage mental health crises, assisting genocide and destroying foreign aid.

You don't hear this sort of stuff about ebays founder, for instance.

Re: X changes its terms to bar training of AI models using its content

#186
post #18

There needs to be a worldwide standard, such as an HTML tag, that says "no training". And a few countries need to make it a punishable offense to violate the tag. The punishment should be exceptionally severe, not just a fine. For example: any company that violates the tag should be completely barred from operating, forever.

That will play out exactly like the "Do not track" bit did.

how did that play out?

Re: X changes its terms to bar training of AI models using its content

#187

Earlier quoted context omitted.

> Elon is not a generous guy Why would he be? Why shouldn't he be? He has 10x more of everything in the world than he could ever possibly use in his lifetime. Greed is not a virtue.

My uncle has 10x more of everything in the world than he could ever possibly use in his lifetime. A lake house, a main house, a few boats and cars. Elon is somewhere around 10,000x.

The median American net worth is $192,700. Elon’s net worth is $393.4 billion, so if I’m doing math right he’s about 204,000,000x more

Re: X changes its terms to bar training of AI models using its content

#189

Earlier quoted context omitted.

My uncle has 10x more of everything in the world than he could ever possibly use in his lifetime. A lake house, a main house, a few boats and cars. Elon is somewhere around 10,000x.

The median American net worth is $192,700. Elon’s net worth is $393.4 billion, so if I’m doing math right he’s about 204,000,000x more

[deleted]

Re: X changes its terms to bar training of AI models using its content

#190

Earlier quoted context omitted.

I've similarly wondered if I could get a pre-2024 Wikipedia if just for the "fact based" flavor LLM

Do you think Wikipedia starting in '24 was polluted by AI slop? This is certainly possible, I'm just not aware of it happening. Wikipedia periodically publishes database dumps and the Internet Archive stores old versions: https://archive.org/search?query=subject%3A%22enwiki%22%20AN... Plus you could also grab the latest and just read the 12/31/23 revisions.

It was already slop, let's not pretend it is significantly different today.
Post reply on HN