Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

211–220 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#211

Earlier quoted context omitted.

This benefits actually everyone. If our combined creative work until this point is what turns out to be necessary to kick-start a great shot at abundance (and if you do not believe that, if it's all for nothing, why care at all about the money wasted on models?) it might simply be our societal moral obligation to endorse it -- just as is will be the model creators moral obligation to uphold their end of this deal. In…

I want to see any indication that abundance form AI would benefit man kind first. While I would love Star Trek society has been going very much towards Cyberpunk aesthetic aka "the rich hold all the power". To be precise AI models fundamentally need content to survive but they need so much content there is no price that makes sense. Allowing AI to monetize without enriching the people who allowed it to exist isn't a…

There are glimpses. Getting a high score on an Olympiad means there is the possibility of being able to autonomously solve very difficult problems in the future.

Re: Facebook scraped every Australian adult user's public posts to train AI

#212

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

Are you concerned that this approach will lead to the abandonment of an open web? If AI companies, and whatever comes next, are expected to take advantage of everything shared online, regardless of copyrights, it seems reasonable that people will stop sharing most things of value.

If you never display your work online, you’ll probably never gain any traction as an artist.

Re: Facebook scraped every Australian adult user's public posts to train AI

#214

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

This is just copyright infringement reworded to pretend it's not. I own the things I write, and publishing it on the internet doesn't negate that. OpenAI doesn't have the right to claim it, no matter what they think, and neither does anyone else.

Re: Facebook scraped every Australian adult user's public posts to train AI

#216

Earlier quoted context omitted.

Are you concerned that this approach will lead to the abandonment of an open web? If AI companies, and whatever comes next, are expected to take advantage of everything shared online, regardless of copyrights, it seems reasonable that people will stop sharing most things of value.

If you never display your work online, you’ll probably never gain any traction as an artist.

I'd be concerned with getting traction if my art is online and anyone can feasibly copy my style.

Maybe that's unimportant and no different than being able to make physical copies, though good forgeries haven't always been so easily done and the forgery is meant to be an identical copy of the original work. It could just be me, but the idea of an attempted identical copy of a well known work feels different than a new creation being passed off as the work of a well known artist. For example, you can claim to have a really god copy of the Mona Lisa but that wouldn't be as valuable as claiming you have a previously unknown, unique work from the artist.

Re: Facebook scraped every Australian adult user's public posts to train AI

#217
post #214

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

This is just copyright infringement reworded to pretend it's not. I own the things I write, and publishing it on the internet doesn't negate that. OpenAI doesn't have the right to claim it, no matter what they think, and neither does anyone else.

Firstly publishing something on Facebook explicitly gives them the right to "copy" it. It certainly gives them the right to exploit it (it's literally their business model.)

Secondly, Facebook is behind a login, so it's not "public" in the way HN comments are public. You'd have gained more kudos had you argued that point.

Thirdly this article I about MetaAI not OpenAI. So, no, OpenAI isn't claiming anything about your Facebook post.

I'll assume however that you digressed from the main topic, and were complaining about OpenAI scraping the web.

Here's the thing. When you publish something publically (on the internet or on paper) you can't control who reads it. You can't control what they learn from it, or how they'll use that knowledge in their own life or work.

You can of course control republishing of the original work, but that's a very narrow use case.

In school we read setwork books. We wrote essays, summaries, objections, theme analysis and so on. Some of my class went on to be writers, influenced by those works and that study.

In the same way OpenAI is reading voraciously. It is using that to assign mathematical probabilities to certain word pairings. It is studying published material in the same way I did at school, albeit with more enthusiasm, diligence and success.

In truth you don't "own the things you write" not in the conceptual sense. You cannot own a concept, argument or position. Ultimately there is nothing new under the sun (see what I did there?) and your blog post is already a rehash of that which came before.

Yes, you "own" the text, to the degree to each any text can be "owned" (which is not much.)

Re: Facebook scraped every Australian adult user's public posts to train AI

#218
post #214

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

This is just copyright infringement reworded to pretend it's not. I own the things I write, and publishing it on the internet doesn't negate that. OpenAI doesn't have the right to claim it, no matter what they think, and neither does anyone else.

> publishing it on the internet doesn't negate that

The terms of use of most sites (including this one) include giving the site owners a license to use what you post, often in any way they see fit.

Re: Facebook scraped every Australian adult user's public posts to train AI

#219
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

Information wants to be free, man.
Post reply on HN