Live data from Hacker News

Microsoft, OpenAI sued for ChatGPT 'privacy violations'

theregister.com

121–130 of 231 posts

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#121
post #96

Earlier quoted context omitted.

It doesn't require humans to work for free — while that's been a common default MO since everyone looked at Google making a search index and thinking to themselves "if they're doing it surely do can we", there are data sets made by paying people.

There are such datasets, and AI companies absolutely pay to have data curated. But I suspect it would be just unimaginably expensive to create a dataset from scratch with enough tokens to feed a model with hundreds of billions of parameters, all the while paying every participant fairly.

"fair" is somewhat undefined, as the fair-looking number for being paid for effort can be very different to the fair-looking number for being paid for the resale value of the end product on an open market.

I wonder what would an LLM trained on Google code and internal documents look like?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#122

Earlier quoted context omitted.

Anyone can read your blog and then post their own blog post using knowledge they learned while reading yours. ChatGPT "learned" from your blog that same way

Anthropomorphizing that it "learned" is disingenuous and I expect better from the HN crowd. If ChatGPT regurgitates verbatim or nearly verbatim, something it slurped up from OP's blog, is that not plagiarism? Where do you draw the line? Where would a reasonable person draw the line?

A human is both capable of reciting things from memory in an infringing manner, and learning from experiences to create something new. Maybe we should tape people's mouth shut if they dare to violate copyright by reciting a copyrighted book word for word or put them in a straight jacket if they recreate a copyrighted painting from memory.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#123

What I noticed is that the privacy setting which should prevent OpenAI to use my data for training purposes, was already deleted twice and I had to set it again. No idea what that means and if the data that I entered before I noticed that setting was gone is now being owned by OpenAI. Anyway, it is obvious that privacy is no priority to them. Also, it's known that YC companies are informally being told they should no…

As I understood it that setting is an opt-out cookie. So must be set on all new browser sessions.

Seems to be a blatant violation of GDPR. So I assume they’ll be fined for it sooner or later and forced to cleanup the training data anyway.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#124
post #102

Earlier quoted context omitted.

> Plaintiff ... is concerned that Defendants have taken her skills and expertise, as reflected in [their] online contributions, and incorporated it into Products that could someday result in [their] professional obsolescence ... It's been a bit surreal seeing modern day Luddites come out of the wood works basically coming up with any ethical/legal argument they can that is a thinly veiled way of saying "I don't want…

As far as I remember Luddites were smart and not against all technology, they were just protecting their jobs. And they were ultimately right. Why? Except for the longshoremen in the US getting compensation and an early retirement due to the introduction of containers, I know of exactly 0 (ZERO!) mass professional reconversions after a technological revolution. Look at deindustrialization in the US, UK, Western Europ…

Stables became gas stations. Nintendo used to be a toymaker.

Businesses change and adapt. Workers too — but people often don’t like change, so many choose to stay behind. Should we cater to them?

I used to do a lot of work which is now mostly automated. Things like sysadmin work, spinning up instances and configuring them manually, maintaining them. I reconverted and learned terraform, aws etc when it became popular.

Should I have gotten help from the government to instead stick to old style sysadmin work?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#125
post #65

Earlier quoted context omitted.

LLMs were invented at least five years ago (BERT) though you could make the case for a few years earlier. My guess is the majority of Reddit users are new since then, not 0.1%?

Your guess is that the majority of Reddit users have joined since 2018? 1) I do not think that is correct, 2) the mere existence of LLMs isn't public awareness about how LLMs are trained, and 3) you know exactly what I'm saying and that 99.9% might be slight hyperbole.

1: Reddit has ~1.6B monthly active users, compared to 0.3B in 2018. [1] So 2x user growth seems more likely to me than not.

2: You're the one who went with "invented" ;)

3: I know you're exaggerating, but I think you think you're exaggerating much less than you actually are.

[1] https://www.bankmycell.com/blog/number-of-reddit-users/

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#126
post #65

Earlier quoted context omitted.

LLMs were invented at least five years ago (BERT) though you could make the case for a few years earlier. My guess is the majority of Reddit users are new since then, not 0.1%?

My account(s) are 17 years old on reddit.

Yes? Mine is nearly that old. But we are very clearly the minority!

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#127

Earlier quoted context omitted.

> including personal information obtained without consent Obtained from (check notes) public internet forums > For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model. You've got to be incredibly naive if you think public Reddit data isn't used to train ML models, not…

Or maybe when you started posting on reddit, LLMs hadn't been invented yet. This is true for 99.9% of the people who post on Reddit.

People have been training ML models on data scraped from Reddit since at least 2015 [1], back when there were less than a million users

[1] https://www.kaggle.com/datasets/ehallmar/reddit-comment-scor...

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#128

What I noticed is that the privacy setting which should prevent OpenAI to use my data for training purposes, was already deleted twice and I had to set it again. No idea what that means and if the data that I entered before I noticed that setting was gone is now being owned by OpenAI. Anyway, it is obvious that privacy is no priority to them. Also, it's known that YC companies are informally being told they should no…

As I understood it that setting is an opt-out cookie. So must be set on all new browser sessions. Seems to be a blatant violation of GDPR. So I assume they’ll be fined for it sooner or later and forced to cleanup the training data anyway.

How is that a GDPR violation?

GDPR doesn’t prevent opt outs of this kind of thing.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#129
post #65

Earlier quoted context omitted.

LLMs were invented at least five years ago (BERT) though you could make the case for a few years earlier. My guess is the majority of Reddit users are new since then, not 0.1%?

Your guess is that the majority of Reddit users have joined since 2018? 1) I do not think that is correct, 2) the mere existence of LLMs isn't public awareness about how LLMs are trained, and 3) you know exactly what I'm saying and that 99.9% might be slight hyperbole.

> Your guess is that the majority of Reddit users have joined since 2018?

It's not really important to the debate around unlicensed use of copyrighted works to train AI models, but it wouldn't surprise me at all if the majority of Reddit users have joined since 2018. It's tough to get reliable active user counts, but they seem to have risen substantially over the past five years.

It also wouldn't surprise me if the majority of Reddit users were indeed from prior to 2018, but at the very least > 2018 would be a very substantial minority.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#130
post #88

Earlier quoted context omitted.

> Plaintiff ... is concerned that Defendants have taken her skills and expertise, as reflected in [their] online contributions, and incorporated it into Products that could someday result in [their] professional obsolescence ... It's been a bit surreal seeing modern day Luddites come out of the wood works basically coming up with any ethical/legal argument they can that is a thinly veiled way of saying "I don't want…

In this case I think it's a little different. People are saying that they don't want to have their own productive or creative output used to undermine their own standard of living. That's not the same as simply not wanting to have your job automated away by someone else's business innovation.

To make chatGPT analogous to coal mining automation it would have to be able to automate the thing it is doing without learning from sources online.

To make coal mining automation analogous to chatGPT the machinery would have had to use something the coal miner did to learn how to automate their work? I'm imagining a camera looking at all the coal miner's work and then the machine can immediately do it, but better.

I agree it is a tad different, but like with someone's coal mining which is in the public domain for anyone in the tunnel to see, likewise anything you write unprotected online is in the public domain and fair game I think?

Post reply on HN