Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

131–140 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#131
post #59

Earlier quoted context omitted.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied

Until now, this has been an acceptable tradeoff because there's some friction to theft. Directly cloning the work is easy, but that also means an artist can sue or DMCA. It also means the original artist's work can go more viral, which, despite the short-term downsides, can help their popularity long term.

The important difference is that imitating an artist's style with new work used to take significant time (hours or days). With an LLM, it takes milliseconds, and that model will be able to churn out the likes of your work millions of times per day, forever. That's the difference, and why the dilemma is new.

> Ultimately "AI did it" should never be allowed to be used an as an excuse

With the exception of an LLM directly plagiarizing, the only way to prove it didn't is by not allowing it to train on something. LLMs are the sum of everything. We could say the same about humans, sure, we are a model trained on everything we've ever seen too. But humans aren't machines who can recreate stuff in the blink of an eye, with nearly perfect recall, at millions of qps.

Re: Facebook scraped every Australian adult user's public posts to train AI

#132
post #46
post #4

Earlier quoted context omitted.

Is this an important distinction?

Yep. For reasons of propriety, for one thing. But also because the data protection laws get especially opinionated about what you do with kids’ speech, and one line they draw is at age 13. The American variant, COPPA, dates back to 2000, and requires verifiable parental consent to process the data of under-13s. No idea if that matters retrospectively in legal terms—it’s seems to me that the main problem was providing…

Surely I am not a fan of FB, but what else should they do other than relying on self-reported age?

Re: Facebook scraped every Australian adult user's public posts to train AI

#133

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

AFAIK, AI models have no way of differentiating high quality input from garbage. If it's fed peer-reviewed, academic papers as well as a paranoid, violent person's Facebook manifesto it treats them with equal weight as long as the sentences are coherent.

On some level it needs to be fed some amount of garbage because it takes in all sorts of garbage inputs like we do.

AI that needs painstakingly curated training data isn't interesting in the same way that early lightbulbs that used precious metals and cost too much to be commercially viable aren't interesting.

Re: Facebook scraped every Australian adult user's public posts to train AI

#134

Earlier quoted context omitted.

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" Heck, you could make the same argument about leaving the house. "That's just how guns work! Don't go out in pu…

> You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" No, the argument would be about photocopies, not paper. "That's just how photocopiers work! Don't put something into a photocopier if you don't want photocopies of it." It isn't possible for anyone to access anything on the internet without m…

Yes, if one over-narrowly construes any analogy, it can be quickly dismissed. I suppose that's my fault for putting an analogy on the internet.

We've had copying technologies since people invented the pen. It was such an important activity that there were people who spent their whole lives copying texts.

With the rise of the printing press, copying became a significant societal concern, one so big that America's founders put copyright into the constitution. [1] The internet did add some new wrinkles, but if anything the surprise is is that most of the legal and moral thinking that predates it translated just fine to the internet age. That internet transmission happens to make temporary copies of things changed very little, and certainly not the broad principles.

I understand why Facebook and other people lining their pockets would like to claim that they are entitled to take what they want. But we don't have to believe them.

[1] https://constitution.congress.gov/browse/essay/artI-S8-C8-1/...

Re: Facebook scraped every Australian adult user's public posts to train AI

#135

Earlier quoted context omitted.

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" Heck, you could make the same argument about leaving the house. "That's just how guns work! Don't go out in pu…

> You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" No, the argument would be about photocopies, not paper. "That's just how photocopiers work! Don't put something into a photocopier if you don't want photocopies of it." It isn't possible for anyone to access anything on the internet without m…

> If that isn't what you want, don't publish your works there.

"Women are oppressed in Iran. Well, that's just how Iran is. Just leave it if you don't want to be oppressed"

Oh my. Yea, and whatever is some way, is that way – "it is how it is, deal with it". It's an empty statement. The topic is an ethical and political discussion in light of current technologies. It's a question of whether it should work this way. That's how all moral questions come about – by asking if something should be the way it is. And the current state of technology brings a dilemma that hasn't existed before.

And no, the internet was not designed for that. Quite obviously. Sounds like you haven't heard of private messages.

I'm very surprised this has to be stated.

Re: Facebook scraped every Australian adult user's public posts to train AI

#136

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. And if someone takes a picture of your artwork, or takes a picture of your person, and posts that to the internet without your consent? Have you given up your rights then? My answer: Absolutely not.

What AI does is much more like the Old Masters approach of going to a museum and painting a copy of a painting by some master whose technique they wish to learn. This has always been both legal, and encouraged.

Or borrowing a thick stack of books from the library, reading them, and using that knowledge as the basis for fiction. That's a transformative work, and those are fine as well.

My take is that training AI models is a bespoke copyright situation which our laws were never designed to handle, and finding an equitable balance will take new law. But as it stands, it's both legal and encouraged for a human to access a Web site (thereby making a copy) and learn from the contents of that website.

That is, fundamentally, what happens when an LLM is trained on corpus data. The difference in scale becomes a difference in kind, but as I said, our laws at present don't really account for that, because they weren't designed to.

LLMs sometimes plagiarize, which is not ok, but most people, myself included, wouldn't consider the dilemma satisfactorily resolved if improvements in the technology meant that never happened. Outside of that, we're talking about a new kind of transformative work, and those are legal.

Re: Facebook scraped every Australian adult user's public posts to train AI

#137

Earlier quoted context omitted.

Well, just as another perspective... I'm not convinced that the philosophy of copyright is a net positive for society. From a certain perspective, all art is theft, and all creativity builds upon preexisting social influences. That's how genres develop, periods, styles... and yes, blatant ripoffs and copycats too. If the underlying goal is to be able to feed creators, maybe society needs better funding models...? The…

> I'm not convinced that the philosophy of copyright is a net positive for society. I'm ok with that. But the philosophy of copyrights is not under debate here. All that is being debated is if it should protect small people from big corporations too.

It's not? I thought we were talking about "AI SHOULD be trained on everything that is in the public sphere" and "[your work] will be sampled by a computer and instantly recreated at scale. [...] Commercial art producers have always ripped off minor artists". Isn't that all about copyright and the ability to make money off your creative works?

When I put something on Wikipedia or any other commons, I don't worry about which other person, algorithm, corporation, or AI ends up reusing it.

But if my ability to eat tomorrow depended on that, then I would very much care. Hence, copyright seems an integral part of people's ability to contribute creatively.

My argument is that by detaching their income from the reusability of their work, we would be able to free more creators from that constraint. Under such a system, the little guy would never get rich off their work, but they wouldn't starve when a big corporation (or anyone else) rips them off either.

Re: Facebook scraped every Australian adult user's public posts to train AI

#138

Earlier quoted context omitted.

> It optimizes against controversy and vitriol rather than encouraging it [...] a bunch of nerds mostly talking shop Something doesn't add up here. Nerds talking shop and controversy and vitriol about the their technical preferences is the same thing. Perhaps what you're saying is that controversy and vitriol is only apparent when you're an "innocent bystander" who doesn't have a passion for the subject? Which HN avo…

OK, but I think there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". I don't think – I hope , at least – no one is going to start a shooting war over their framework of choice. You can't say the same about much of the content circulating around social media. (Edit: You added more to…

> there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks".

Is there? Perhaps the trouble here is that your examples are too far apart to recognize how they compare?

What if Facebook, instead, said "Fat workers are too bloated, you should hire skinny workers"? Or if HN said "so and so projects are all ____ and they should all be ____, and oh, your vacuum doesn't suck".

I see no practical difference. It seems the only difference is that Facebook tends to talk about people, HN about tech. But that's not a significant distinction – aside from where your interests lie. Certainly tech-minded folks often find people to be uninteresting.

Re: Facebook scraped every Australian adult user's public posts to train AI

#139

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. This is true for average people. Is it true for the wealthy? Is it true for Disney? Does our law acknowledge this truth and ensure equal justice for all?

It's 100% true for everyone. You can't access anything at disney.com without making a copy of that thing. Disney can't access anything at yourdomain.whatever without making a copy of that thing.

Whatever crimes either of you can get away with using your copies is another matter entirely. Any rights you had under the legal system you had before AI haven't gone away, neither have the disadvantages you have against the wealthy.

Re: Facebook scraped every Australian adult user's public posts to train AI

#140

Earlier quoted context omitted.

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" Heck, you could make the same argument about leaving the house. "That's just how guns work! Don't go out in pu…

> You could make the same argument about paper. Most paper doesn't come with Terms and Conditions that everything you write on it belongs to the paper company. I hate Facebook (with a fiery passion) but people gave them their data in exchange for the groundbreaking and unprecedented ability to make friends with another person (which has never been done before). It sucks, but don't use these "free" systems without und…

I think you're confusing a legal point (whether a T&C really gives Facebook any particular legal right in court) with the moral question of whether or not people should just roll over for large companies because of language we all, Facebook included, know that nobody ever reads.

Even if FB's T&C made it clear they could do this (something I haven't seen proven), that at best means people would have a hard time suing as individuals. They can still get upset. They can still protest to the regulators and legislators whose job it is to keep these companies in line, and who create the legal context that gives a T&C document practical meaning.

Post reply on HN