Earlier quoted context omitted.
This is just copyright infringement reworded to pretend it's not. I own the things I write, and publishing it on the internet doesn't negate that. OpenAI doesn't have the right to claim it, no matter what they think, and neither does anyone else.
Firstly publishing something on Facebook explicitly gives them the right to "copy" it. It certainly gives them the right to exploit it (it's literally their business model.) Secondly, Facebook is behind a login, so it's not "public" in the way HN comments are public. You'd have gained more kudos had you argued that point. Thirdly this article I about MetaAI not OpenAI. So, no, OpenAI isn't claiming anything about you…
Facebook scraped every Australian adult user's public posts to train AI
231–240 of 267 posts
Re: Facebook scraped every Australian adult user's public posts to train AI
#232Presumably “scraped” isnt the right term here. They already have the raw data, they Won’t be “scraping “ it from the website they’ll just be investing it from where they store it
It’s interesting to think about what the right verb is. I’d probably say Meta trained their models using all self-hosted, public AU citizen’s data. But it doesn’t really sound as scary as “scraped” to non-technical users.
Re: Facebook scraped every Australian adult user's public posts to train AI
#233Presumably “scraped” isnt the right term here. They already have the raw data, they Won’t be “scraping “ it from the website they’ll just be investing it from where they store it
In some legislations there are rules about scraping. And for many less technical people it sounds scary.
Re: Facebook scraped every Australian adult user's public posts to train AI
#234Earlier quoted context omitted.
I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…
Where the puck is about to be is very different from where it is. Generative AI hasn't cracked the creativity problem yet. It can generate new art but it can't develop its own style like a human can (from first principles, humans basically caricature high quality video feed). There is pretty good reason to believe that this will be a solved problem inside a decade. We're moving towards processing video and the power…
Re: Facebook scraped every Australian adult user's public posts to train AI
#235Presumably “scraped” isnt the right term here. They already have the raw data, they Won’t be “scraping “ it from the website they’ll just be investing it from where they store it
Re: Facebook scraped every Australian adult user's public posts to train AI
#236That’s nothing. AOL has just finished training on 29 years of emails and messages. it’s hoped that with more H100s the AI will finally be able to calculate the full amount due by BillG for the emails mom has been forwarding.
TIL AOL still exists.
Re: Facebook scraped every Australian adult user's public posts to train AI
#237People just submitted it. I don't know why. They 'trust me'. Dumb fucks. -Mark Zuckerberg Things change, but this never stop being a concise summary of Meta's ethos as a company.
Re: Facebook scraped every Australian adult user's public posts to train AI
#238Presumably “scraped” isnt the right term here. They already have the raw data, they Won’t be “scraping “ it from the website they’ll just be investing it from where they store it
The authors of the title most likely wanted to suggest a similarity between metas use of the data and scraping. In some legislations there are rules about scraping. And for many less technical people it sounds scary.
Re: Facebook scraped every Australian adult user's public posts to train AI
#239Earlier quoted context omitted.
We are talking about more than the current law here. We're talking about what the law should be, based on what people see as right. And I'd add that Facebook is doing a lot more here than just quietly having a copy of something.
If the concern isn't copyright infringement what would the new law be about? What's the harm people want solved? Is it just some philosophical objection, like people not liking the idea of other people doing something they don't like? Is it fear of future potential harms that haven't been seen in real life yet? Is it just "someone else may be using what I published publicly to make money somehow in a way that is lega…
Copyright is a thing we made up because of "people not liking the idea of other people doing something they don't like". The current specific boundaries are a careful compromise about exactly when we protect "using what I published publicly to make money somehow".
Those boundaries in large part exist because of responses to specific technologies. Before the technologies, people didn't care. After, people got upset and then we changed the laws, or the court rulings that provide detailed meaning to the laws shifted.
As an example, you could look at moral rights, which are another intellectual property right the laws for which came much later than copyright. Or you could look at how intellectual property law around music has shifted in response to recording, and again in response to sampling. Or you could look at copyright for text, which was invented in response to the printing press. And more specifically in response to some people using pioneering technology to profit from other people's creative work.
And we might not need any changes in the law here. The people at OpenAI and elsewhere know they're doing something that could well be illegal. They've been repeatedly sued, and they've chosen to make deals with some data-holders. They wouldn't be paying for published data at all if they knew they were in the clear, but they've chosen to cut many deals with publishers powerful enough to sue. They're hoovering up data anyhow because, like too many startups, they've decided they'll do what they want and see if they can get away with it.
Re: Facebook scraped every Australian adult user's public posts to train AI
#240I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.