Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

181–190 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#181

That’s nothing. AOL has just finished training on 29 years of emails and messages. it’s hoped that with more H100s the AI will finally be able to calculate the full amount due by BillG for the emails mom has been forwarding.

TIL AOL still exists.

Re: Facebook scraped every Australian adult user's public posts to train AI

#182

Earlier quoted context omitted.

Yes, if one over-narrowly construes any analogy, it can be quickly dismissed. I suppose that's my fault for putting an analogy on the internet. We've had copying technologies since people invented the pen. It was such an important activity that there were people who spent their whole lives copying texts. With the rise of the printing press, copying became a significant societal concern, one so big that America's foun…

I don't think that facebook should be allowed to violate copyright law, but clearly they have the same rights as you do to copy works made publicly avilable on the internet.

We are talking about more than the current law here. We're talking about what the law should be, based on what people see as right. And I'd add that Facebook is doing a lot more here than just quietly having a copy of something.

Re: Facebook scraped every Australian adult user's public posts to train AI

#183

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. Or we could be ethical and encourage others to be ethical.

Okay, but that doesn't change how the Internet works.

Encouraging people to be ethical isn't actually a real way to prevent people copying photos you put up online.

Re: Facebook scraped every Australian adult user's public posts to train AI

#184

Earlier quoted context omitted.

I don't think that facebook should be allowed to violate copyright law, but clearly they have the same rights as you do to copy works made publicly avilable on the internet.

We are talking about more than the current law here. We're talking about what the law should be, based on what people see as right. And I'd add that Facebook is doing a lot more here than just quietly having a copy of something.

If the concern isn't copyright infringement what would the new law be about? What's the harm people want solved? Is it just some philosophical objection, like people not liking the idea of other people doing something they don't like? Is it fear of future potential harms that haven't been seen in real life yet?

Is it just "someone else may be using what I published publicly to make money somehow in a way that is legal and doesn't infringe on my copyrights but I still don't like it because they aren't giving me a cut?" What would the new law look like?

Re: Facebook scraped every Australian adult user's public posts to train AI

#185
post #121

Earlier quoted context omitted.

> Question would be whether machine learning, unlike human learning, should be treated as copyright infringement. No, the question is whether those genAI we have around are mass copyrights violation machines or whether they "learn" and build non-violating work. And honestly, I have seen evidence pointing both ways. But the "copyrights protection" institutions are all quickly to decide the point dismissing any evidenc…

> No, the question is whether those genAI we have around are mass copyrights violation machines or whether they "learn" and build non-violating work. I refer to the training process in question, which may or may not be be violating copyright, as "machine learning" since that's the common terminology. Question is whether that process is covered by fair use. Whether or not it actually "learn"s is not irrelevant, but I'…

> I refer to the training process in question

Yeah, you go for the red herring.

All of the worthwhile debate is about the real violations. But the public discourse is surely inundated with that exact red herring.

Re: Facebook scraped every Australian adult user's public posts to train AI

#186
post #121

Earlier quoted context omitted.

> No, the question is whether those genAI we have around are mass copyrights violation machines or whether they "learn" and build non-violating work. I refer to the training process in question, which may or may not be be violating copyright, as "machine learning" since that's the common terminology. Question is whether that process is covered by fair use. Whether or not it actually "learn"s is not irrelevant, but I'…

> I refer to the training process in question Yeah, you go for the red herring. All of the worthwhile debate is about the real violations. But the public discourse is surely inundated with that exact red herring.

I addressed model output (infringes copyright if substantially similar, as with manually-created works) and the process of training the model (requires collating/processing ephemeral copies, possibly fair use). What do you think the "real violations" are, if not those?

Re: Facebook scraped every Australian adult user's public posts to train AI

#187

Earlier quoted context omitted.

OK, but I think there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". I don't think – I hope , at least – no one is going to start a shooting war over their framework of choice. You can't say the same about much of the content circulating around social media. (Edit: You added more to…

> there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". Is there? Perhaps the trouble here is that your examples are too far apart to recognize how they compare? What if Facebook, instead, said "Fat workers are too bloated, you should hire skinny workers"? Or if HN said "so and so pro…

> Is there?

The difference is that talking about a framework with passion doesn't generally end up with severe real world consequences for other people.

The difference between tech and people is that people are... people. They have lives, feelings, all those fun human things. Passionately talking about how X group of people should die (or ) is quite different than passionately talking about how Y framework is the worst thing programmers have ever concocted. Generally the latter won't end up with someone getting stabbed. (Generally...)

> What if Facebook, instead, said "Fat workers are too bloated, you should hire skinny workers"? Or if HN said "so and so projects are all ____ and they should all be ____, and oh, your vacuum doesn't suck".

I think this is a terribly unfair switcharoo to make. Let's fill in the blanks for that second quote.

> "so and so group of people are all inhuman and they should all be killed, and oh, your mom sucks"

> "so and so projects are all terrible and they should all be deleted, and oh, your vacuum doesn't suck"

Re: Facebook scraped every Australian adult user's public posts to train AI

#188

Earlier quoted context omitted.

What AI does is much more like the Old Masters approach of going to a museum and painting a copy of a painting by some master whose technique they wish to learn. This has always been both legal, and encouraged. Or borrowing a thick stack of books from the library, reading them, and using that knowledge as the basis for fiction. That's a transformative work, and those are fine as well. My take is that training AI mode…

> This has always been both legal, and encouraged. Not always. The copy must be easily identifiable as copy. An exact reproduction can't have the same dimensions as the original for example. Drawing just a person or a detail of the picture, or redoing the picture in a different context or style, is encouraged. Selling a full scale photo of the picture is forbidden. The copyright of famous art belongs to the museum.

The second example is better than the first, yes. I was thinking about the process more than the fact that painting a study produces a work, and a derived one at that, so more normal copyright considerations apply to the work itself.

> An exact reproduction can't have the same dimensions as the original

This is a rule, not a law, and a traditional and widespread one. Museums don't want to be involved in someone selling a forgery, so that rule is a way of making it unlikely. But the difference between "if you do this a museum will kick you out" and "this is illegal" is fairly sharp.

> The copyright of famous art belongs to the museum.

Not in a great number of cases it doesn't, most famous art is long out of copyright and belongs to the public domain. Museums will have copyright on photos of those works, and have been known to fraudulently claim that photos taken by others owe a license fee to the museum, but in the US at least this isn't true. https://www.huffpost.com/entry/museum-paintings-copyright_b_...

Re: Facebook scraped every Australian adult user's public posts to train AI

#189

Earlier quoted context omitted.

"That's just how the internet works" is nonsensical when AI is changing how the internet works. Just because the tradeoffs of sharing on the internet used to work before AI, doesn't mean those tradeoffs continue to be workable after AI. It's like having drones follow everyone around and publish realtime telephoto video of them because they have "no expectation of privacy" in public places. Maybe before surveillance t…

> Maybe before surveillance tech existed, there was no expectation of privacy in public places, but now that surveillance tech exists, people naturally expect that high-res video of their every move won't be collected, archived and published even if they are in public. currently, that'd be an unrealistic expectation. I'd agree that it would be nice if that wasn't the case but laws need to catch up with technology. Ri…

The topic of expectations reminds me of this article

https://spectrum.ieee.org/online-privacy

Re: Facebook scraped every Australian adult user's public posts to train AI

#190
post #135

Earlier quoted context omitted.

> If that isn't what you want, don't publish your works there. "Women are oppressed in Iran. Well, that's just how Iran is. Just leave it if you don't want to be oppressed" Oh my. Yea, and whatever is some way, is that way – "it is how it is, deal with it". It's an empty statement. The topic is an ethical and political discussion in light of current technologies. It's a question of whether it should work this way. Th…

> "Women are oppressed in Iran. Well, that's just how Iran is. Just leave it if you don't want to be oppressed" Yea, and whatever is some way, is that way – "it is how it is, deal with it". It's an empty statement. No, because Iran can stop oppressing women and still exist as a functional country. oppressing women today is "how it is". The internet on the other hand is designed to be a system for the distribution of…

all received messages are copies of the original. a broadcast is a copy. language itself is an incessant copying. so it’s a truism to say that of the internet. and a generality that doesn’t apply to specifics. downloading cracked software is also copying, but this genus is irrelevant to its discussion. its morality is beside its being a copy, even though it is essential that it be a copy. likewise with other data.

we don’t have rules set yet, that’s why this discussion is active, ie, not just a niggle from after a couple of beers. it’s a question of respect for the author.

Post reply on HN