Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

201–210 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#201

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

Are you concerned that this approach will lead to the abandonment of an open web?

If AI companies, and whatever comes next, are expected to take advantage of everything shared online, regardless of copyrights, it seems reasonable that people will stop sharing most things of value.

Re: Facebook scraped every Australian adult user's public posts to train AI

#202
post #99
post #59

Earlier quoted context omitted.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

> Why would we embrace this now that a computer can do it and there's a level of deniability? Generally I don't think people are arguing that copyright law should be more lenient to AI than it is to humans. If your work gets ripped off (a substantially similar copy not covered by fair use) you can sue regardless of tools used in its creation. Question would be whether machine learning, unlike human learning, should b…

> Generally I don't think people are arguing that copyright law should be more lenient to AI than it is to humans. If your work gets ripped off (a substantially similar copy not covered by fair use) you can sue regardless of tools used in its creation.

With humans, copyright law deals with knowing and intentional infringement more severely than accidental and unintentional infringement.

With an AI, any infringement on the part of the AI end-user is very likely going to be accidental and unintentional rather than knowing and intentional, so the legal system is going to deal with it more leniently, even if actual infringement is proven. The exception would be if you deliberately prompted it to create a modified version of a pre-existing copyrighted work.

With humans, whether infringement is knowing or not, intentional or not, can turn into a massive legal stoush. Whereas, if you say it is AI output, and it appears to actually be AI output, it is going to be much harder for the plaintiff (or prosecution) to convince the court that infringement was knowing and intentional.

Re: Facebook scraped every Australian adult user's public posts to train AI

#203
post #131

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied Until now, this has been an acceptable tradeoff because there's some friction to theft. Directly cloning the work is easy, but that also means an artist can sue or DMCA. It also means the original artist's work can go more viral, which, despite the short-term downsides, can help their p…

Outside of competing profit motives, there is no dilemma. It's that underlying motive, and it's root, that will have to undergo a drastic change. The Pandora that AI is is already out of the box and there's no putting it back in; only dealing with the consequences.

Re: Facebook scraped every Australian adult user's public posts to train AI

#205

Presumably “scraped” isnt the right term here. They already have the raw data, they Won’t be “scraping “ it from the website they’ll just be investing it from where they store it

Agreed, scraping is more appropriate for when one gathers data from a 3rd party site.

Re: Facebook scraped every Australian adult user's public posts to train AI

#207
If the service is free, you're the product. Here's a radical idea: if you don't want Facebook to use your information and content, don't post it to Facebook.

... or does everyone around here think anything different is happening to their posts?

Re: Facebook scraped every Australian adult user's public posts to train AI

#209
post #181

That’s nothing. AOL has just finished training on 29 years of emails and messages. it’s hoped that with more H100s the AI will finally be able to calculate the full amount due by BillG for the emails mom has been forwarding.

TIL AOL still exists.

Verizon.

Re: Facebook scraped every Australian adult user's public posts to train AI

#210
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

It doesn’t recreate anything outside of edge cases you really have to go looking for. It will ingest and spit out the style though and I see nothing wrong with that. It’s basically what people do right now.
Post reply on HN