Live data from Hacker News

Microsoft, OpenAI sued for ChatGPT 'privacy violations'

theregister.com

201–210 of 231 posts

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#201

Earlier quoted context omitted.

Regardless of access rights to the data, I've yet to read a compelling argument why LLMs are even derivative works. You can't identify your Reddit comment in a ChatGPT conversation. How is it any different than a human learning English by reading Reddit? That human wouldn't be violating copyright every time they said a phrase that was repeated by hundreds of Redditors. My favorite LLM analogy so far is the "lossy jpe…

I've been thinking of the output as fanfiction/fan art. It shares many of the same complications regarding the ownership of ideas, commerical intent of writing, competition, and copyright. Fanfiction is generally a protected form of expression, but requires the work to be "transformative". Unlike with parodies and critisisms, fanfiction can be much harder to distinguish from original work. From that perspective, a la…

Fanfiction isn't as protected as many people think it is.

https://en.wikipedia.org/wiki/Legal_issues_with_fan_fiction

Fanfiction and fan art also tend to run afoul of the infrequently (but occasionally) litigated part of copyright - copyright of fictional characters.

https://en.wikipedia.org/wiki/Copyright_protection_for_ficti...

I came across this with the Eleanor lawsuits - https://www.caranddriver.com/news/a42233053/shelby-estate-wi... - and while I believe that that instance Eleanor falls on the "this shouldn't have been copyrightable" (took a bit to get there), the question is "what protects the representation of Darth Vader?"

In general it tends to be ignored and tacitly encouraged... but it isn't protected.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#202
post #198

Earlier quoted context omitted.

No, the work has not been. The impression that the work leaves on a neural network has been though. AIs are not massive repositories of harvested data. The models are relatively small (<20GB).

A resized, smaller, or encoded version of an image is still subject to copyright. Calling an encoding an 'impression' is deceitful.

Not always.

https://www.pinsentmasons.com/out-law/news/google-thumbnails...

> A US court ruled this week that Google's creation and display of thumbnail images does not infringe copyright. It also said that Google was not responsible for the copyright violations of other sites which it frames and links to.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#203

Earlier quoted context omitted.

I don’t get your point. Whether you use copyrighted material in commercial context or not always matters. That’s one of the most important aspects of different open source licenses.

This is not true for copyright law (the 4-factor test[0]) or for OSI licenses (they almost universally place no restrictions on commercial use). The only exception that comes to mind right now is the Creative Commons NC, which is generally recognized as being unsuitable for software[1]. [0]: https://fairuse.stanford.edu/overview/fair-use/four-factors/ [1]: https://creativecommons.org/faq/#can-i-apply-a-creative-comm.…

And CC-NC isn't considered an open source license by the FSF or OSI anyway. And IMO the NC clause is pretty much impossible to define for non-trivial use and Creative Commons basically came up. Not sure non-derivatives is a lot better especially given remixing was one of the original drivers behind CC but it's at least less controversial.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#204
post #141

Earlier quoted context omitted.

This is a fantastic point. I can legally go pick up any strictly copyrighted book at a store and read parts of it for free which I will then have learnt and have in my brain to share with to anyone else. If I happen to have a superintelligent brain I can potentially gain a lot more and make a lot more inferences from this one outing and consequently add a lot of value to others I share my info to. But telling me it i…

If you go read a book, memorize it, write it down later in a substantively similar form, and share it freely or sell it — yes, you might get into copyright trouble. It has happened before and it is at best a tricky gray area. If you pick up a book and learn a fact, then yeah, you’re allowed to share that fact. It’s weird that this topic keeps devolving into a form of “so what, it’s illegal for me to learn things?” Be…

> You have a different set of rights than ChatGPT.

Gods, no. Where did you get that from?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#205

Earlier quoted context omitted.

As I understood it that setting is an opt-out cookie. So must be set on all new browser sessions. Seems to be a blatant violation of GDPR. So I assume they’ll be fined for it sooner or later and forced to cleanup the training data anyway.

How is that a GDPR violation? GDPR doesn’t prevent opt outs of this kind of thing.

In the sense that consent requires active opt-in. The passive “opt-in” by failing to set the cookie doesn’t count as consent.

So if they’re claiming they have the right to process data on the legal basis of consent, and they claim the absence of that cookie constitutes that consent, then they have no legal basis, and are thus in violation of the law.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#206

Earlier quoted context omitted.

It seems people prefer power distributed by capital, rather than military might or factionalism/leaders/politics. Not that all capital is distributed by merit, plenty of people used military might or factionalism/leaders/politics to obtain disproportionate amount of capital. But if you are against the last 2 happening, I don't see what you expect a reorganization of society to accomplish since you are going to get a…

Capital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might. I don't necessarily disagree with your later points. I do, however, disagree with giving up.

>Capital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might.

At least its equitable (based on value of output), ofc there are legacy issues as well.

Some demagogue can swoon the masses and take it all if not for capital. That demagogue could be Trump or Stalin.

Know the consequences of what you are advocating for.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#207
post #193

Earlier quoted context omitted.

Suppose you're a weaver. It's hard, fiddly work, and you have to get your timing and your tension just right to make quality material. Now, there are mechanised looms that can do the job faster (though the quality's not great : they could still do with some improvement, in your opinion). From this efficiency gain, who should reap the profits? Suppose you're a farmer. You've been working on your tractors for decades,…

I've answered this before. The container revolution split some of the resulting profits with those whose livelihoods were destroyed, the longshoremen. "You build a dam that destroys 10000 homes, who should reap the profits?"

It's a good answer, but it raises further questions:

• Should we be destroying people's homes to build dams without their consent?

• In general, are people being compensated when these things happen to them? i.e., while it might be nice, does this actually happen?

The Luddites (the real ones, not the mythological bastardisation of them) continue to be sympathetic characters.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#208

Earlier quoted context omitted.

Capital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might. I don't necessarily disagree with your later points. I do, however, disagree with giving up.

>Capital primacy is maintained by the capitalist state, i.e. the monopoly on violence. This is literally military might. At least its equitable (based on value of output), ofc there are legacy issues as well. Some demagogue can swoon the masses and take it all if not for capital. That demagogue could be Trump or Stalin. Know the consequences of what you are advocating for.

> At least its equitable (based on value of output), ofc there are legacy issues as well.

It's not. By definition, it's based on control of capital. That's why it's called capitalism. In other words, those aren't "legacy" issues; they are literally the system as designed.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#209
post #86
post #77

Earlier quoted context omitted.

>You don’t get to make information publicly available. But not publicly available. But we do? Open sourcing something with caveats is common. This code is public BUT not for commercial use. This code is public BUT you must display attribution etc. Sure, blogposts are unlicensed (that I know) but the idea of something publicly available being held to restrictions is nothing new.

Do you allow commercial employees to read the code and incorporate knowledge obtained from the code into their brains?

Yes, it's completely unfeasible to make a license to control that.

On the other hand, it's completely feasible to make a license that stops someone from training their model with some piece of info, is it not?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#210

Earlier quoted context omitted.

Another day, another person on HN showing us how they don't understand the difference between Public Domain and Open Source or Copyleft etc. And regardless -- the problem now is that expectations of how content can be consumed are now fundamentally violated by automation of content ingestion. People put stuff up on the Internet with the expectation of its consumption by human minds, which have inherent limitations on…

> People put stuff up on the Internet with the expectation of its consumption by human minds Then people obviously aren’t aware that bots have been indexing web pages and showing summarized information without going to the web page for three decades.

I think it's a bit intellectually dishonest to claim an equivalence between content indexing for search engines and machine learning for LLMs. They might share an underlying harvesting technique, but their uses -- indexing for information accessibility vs automatic content production are qualitatively different.

Further, almost every site has had an e.g. robots.txt which has permitted content harvesting only for certain accepted purposes for a couple decades now. So clearly people already had a sense of how they wanted their content harvested and for what purposes.

Post reply on HN