Live data from Hacker News

Microsoft, OpenAI sued for ChatGPT 'privacy violations'

theregister.com

171–180 of 231 posts

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#172
post #44

I mean, it ingested all of the content from my blog. Without my permission. It's not a major part of their corpus of data, but still -- I wasn't asked and I don't really care to donate work to large corporations like that. So the technology is cool, but I'm firmly of the stance that they cut corners and trampled peoples' rights to get a product out the door. I wouldn't be entirely unhappy if this iteration of these p…

Anyone who reads your blog is "ingesting content" from it. That is presumably the purpose of your blog in the first place. Whether that content is used to train a human mind or an artificial one is probably not up to you as the author.

This type of comments can be seen every single time a thread about LLM, or OpenAI or some such comes up.

And it adds nothing. I'm sorry but saying "Whether that content is used to train a human mind or an artificial one is probably not up to you" may be worse than saying nothing at all.

First because it shows enough doubt on whether it's up to the authors of content (IP laws, fair use, intent of the use, and many things I ignore), while giving no laws as an example or frame of reference.

And second because it's comparing a human mind that we know exist, to an artificial one, which implies:

1. An LLM is an artificial mind, or close to one, whatever that is (again, not defined).

2. If they were to exist, they would be both equivalent and treated the same as a human one.

The amount of jumps in a couple sentences, added to the uncertainty of how copyright would/will work, multiplied by the numer of times I/we read that type of comment every single time, it's getting tiresome. And it's adding noise to the noise-signal ratio.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#173
post #147

Earlier quoted context omitted.

> I can personally memorize and recite copyrighted works all I want, Whoever told you that is lying to you. You are not legally allowed to personally memorize and recite copyrighted works all you want, any more than you're allowed to personally memorize, write down copyrighted works, and distribute them as much as you want. All piracy is a process of computer-assisted remembering and reciting.

Last I checked I can legally enter any bookstore with copyrighted books, pick up a book, and read it. And then tell anyone what I read. I can't go write and commercialize what I learnt directly, but I'm not breaking the law by quickly seeing how some book I didn't buy ends so I can talk about it at a party - and then everyone knows how it ends which might affect whether they want to buy said book and upset the author…

Maybe it would be useful what "tell anyone what I read" means. Because if you mean 1 to some in a room, then most likely. If you use any type of broadcasting then most definitely no. Try reading outloud a script from a recent movie on twitch/youtube/radio/tv and whether it gets DMCA'd or not. Same for books, songs I guess... not? But not sure.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#174
post #146

Earlier quoted context omitted.

My point is that you have to separate the method for collecting the data versus the usage of the data as separate legal questions. Scraping is legal. What you do with the data that you scrap though is a whole other question. To put it another way, it's legal for me to go to the library and borrow a DVD or a book or poems. That doesn't give me the right to publish the poems again under my own name. Whether I find the…

What you describe misrepresents how LLMs/neural networks and the math works, your analogy does not apply. There's no static data in the networks. The output of LLMs are much closer to parodies and fanfiction. In that case, you very clearly own the copyright to the new work you make.

That's weird, since my comment literally said nothing about LLMs. I was simply pointing out that making scraping legal doesn't invalidate any of the other data laws that were out there, and gave one example.

You keep making the claim that because it was scraped people can do whatever they want, as scraping is legal. That is the only thing I'm arguing against, because that is a gross misinterpretation of how the case that made scraping legal was decided. LLMs aren't relevant to that point (which is exactly what I keep saying- the method of collection doesn't magically change the legality of it).

That being said, you're still wrong. The USPO has said that the output of LLMs are the outputs of algorithms and are not creative works. Therefore you can't "own the copyright to the new work you make" because the work itself can't be copyrighted at all. No one can own the output of an LLM.

Also, just because it seems you want to be wrong on every level, it is absolutely possible that a neural network would be able to repeat data from its training set. This is an incredibly known problem in the field.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#175
post #102

Earlier quoted context omitted.

As far as I remember Luddites were smart and not against all technology, they were just protecting their jobs. And they were ultimately right. Why? Except for the longshoremen in the US getting compensation and an early retirement due to the introduction of containers, I know of exactly 0 (ZERO!) mass professional reconversions after a technological revolution. Look at deindustrialization in the US, UK, Western Europ…

But as the corollary to that, I know of zero successfully stopped technological revolutions. You can't put the genie back in the bottle, and there is no way to stop progress, aside from a one-world authoritarian government that forcibly stops as much of it as they can. But even that would only be marginally effective. Progress would eventually resume.

Yes, you do know of revolutions stopped and it worked for centuries.

Tokugawa Japan, Qing China, many other places including in Europe for centuries.

That's too extreme.

My point is that we're reaching a point where people need to be compensated. We can't just destroy their lives, collect all the money in 2 bank accounts and call it a day.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#176
post #141
post #86

Earlier quoted context omitted.

Do you allow commercial employees to read the code and incorporate knowledge obtained from the code into their brains?

This is a fantastic point. I can legally go pick up any strictly copyrighted book at a store and read parts of it for free which I will then have learnt and have in my brain to share with to anyone else. If I happen to have a superintelligent brain I can potentially gain a lot more and make a lot more inferences from this one outing and consequently add a lot of value to others I share my info to. But telling me it i…

If you go read a book, memorize it, write it down later in a substantively similar form, and share it freely or sell it — yes, you might get into copyright trouble. It has happened before and it is at best a tricky gray area.

If you pick up a book and learn a fact, then yeah, you’re allowed to share that fact.

It’s weird that this topic keeps devolving into a form of “so what, it’s illegal for me to learn things?” Because: no, it’s not. And: You and a piece of software are treated differently under the law. You have a different set of rights than ChatGPT.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#177
post #148

Earlier quoted context omitted.

Can you point to where that "right" is codified in law?

Common law of contracts dictates that you can commit to performing certain services in exchange for the counter-party performing certain services. For example, you provide both money, viewing data, and permission to run DRM and proprietary code on your property (e.g. set-top boxes or smart TVs) to Netflix in exchange for obtaining access to their library of TV shows and movies. It's codified in the fact that saying y…

You still haven't said where it's legal that all rights can be signed away. I know for a fact that you can't waive tenant rights when signing a lease, for example. We also don't allow people to sign over so many rights that they're considered slaves, as slavery has been made illegal. I also can't sign away my right to not be sexually harassed- if a company makes me sign something saying that they can sexually harass me they will still end up losing in court. The US has also limited the ability for NDAs to cover discussions about labor practices, so there's another right we can't sign away.

It seems to me there are a to of counter examples to this "right" you speak of. So many that it doesn't seem like it really exists.

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#178

Earlier quoted context omitted.

1. How would this not make tools like Github Copilot exorbitantly expensive? Why should I have to pay a tax to everyone else in the United States to use something that was disproportionately trained on my own data? 2. Given that the internet is global, is every country supposed to make their own versions of this? Will I have to pay the EU tax to use models that might have been trained on data that Europeans posted on…

To your first question, it would incentivize training of models on one's own data exclusively -- companies could train something like Copilot on their own code, for instance. To your second question, there's no way to have an international policy like this so yes each jurisdiction would do it independently -- just as they do with thousands of other similar things.

I don't think a model trained on a single company's data would be nearly as helpful as a model trained on all publicly licensed code on the internet. But suppose it were...

What if I'm not a massive corporation with millions of lines of code to train on and I want to pay for an AI coding assistant? Doesn't this make it effectively illegal for me to purchase such a product for a reasonable price when big companies will presumably be able to use it without paying the tax?

Another situation - let's say you're a company that contributes heavily to open source, but also accepts external contributions. Could Facebook train a model on the React codebase, for example, without having to pay the AI tax?

Another situation - suppose I start an LLM coding assistant and sell it to my friend. Presumably I don't have to pay the tax as a "low revenue" company. Then I get acquired or get some huge seed round and suddenly my customers have to pay the AI tax. Doesn't this just nuke all my customers?

Anyway, as a software engineer, I personally want people to use my code for whatever they want to use it for, without having to pay me for it. I indicate that by using an MIT license. Why throw that precedent out the window?

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#179
post #84

Earlier quoted context omitted.

I agree that there is additional nuance, but so far public data scraping has very clearly been ruled as legal. It's possible that at the time of scraping, copyrighted data was incorporated into the training data because it hadn't been taken down by the host platform yet. But in my opinion, the core idea proposed by the suit that private data was used intentionally, is not true. The GPT4 browsing plugin is equivalent…

Even if they were exposing static data, how would that be different than a search engine? Google has been scraping the web for two decades, indexing even explicitly copyrighted content, and then making money by selling ads next to snippets from that content. If you're going to make the case that an LLM is violating copyright, then surely you must also assert that Google is too, because it's the same concept, but Goog…

By putting something on a public-facing website, it's generally agreed that (absent a robots.txt to the contrary), you intend it to appear in web search results, and you're granting a public limited semi-transferable revocable license to request, download and view your site to your visitors.

That doesn't mean you grant a license to produce derivative works other than search indexes. Legally, it's different. (Germany codifies these as separate "moral rights": Urheberpersönlichkeitsrecht.)

Re: Microsoft, OpenAI sued for ChatGPT 'privacy violations'

#180
post #44

I mean, it ingested all of the content from my blog. Without my permission. It's not a major part of their corpus of data, but still -- I wasn't asked and I don't really care to donate work to large corporations like that. So the technology is cool, but I'm firmly of the stance that they cut corners and trampled peoples' rights to get a product out the door. I wouldn't be entirely unhappy if this iteration of these p…

You sent your content to them in response to their HTTP requests. That sure looks like affirmative consent to me.

You’re right! Just like Disney+ did when I watched Star Wars the other day. I’m excited to know Disney has consented to me posting Star Wars in its entirety free online.
Post reply on HN