Live data from Hacker News

ChatGPT is not ‘artificial intelligence.’ It’s theft

americamagazine.org

11–20 of 168 posts

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#11
post #5

Using this logic, every writer who has learned to read and write by reading books, or every artist who improved their craft by studying works, or every musician who learned the piano by practicing pieces, is also "stealing" in whatever they "originally" create due to learning via pattern recognition "tiny pieces of every work" in their data set. It's ridiculous to compare agents that generalize well to "stealing" pie…

AI is not writers artists or developers, it’s software that ingests data and generates an output. Anthropomorphism is going out of fashion fast.

"Stealing" is an action that does not depend on who or what the perpetrator is.

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#12

It doesn't chop text content into bits and reproduce it, that's absurd. It's a text prediction engine, it predicts what tokens come next based on some context

This guy did his homework. I was looking at some interactive ed material and it is precisely what you are saying. Interesting stuff.

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#15
I think this boils down to an argument against the copyrightability of LLMs trained on public data. If all the data in the LLM was public then how can you assert that your copy is proprietary just because the copy is in a very novel type of lossy compression encoding? If I take someone’s art and JPEG compress it or chop it up and store it in some kind of database do I own it now?

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#18
I think one cool benefit of AI advances will be, that many more ordinary people will start thinking about philosophical problems such as the one posed in this article.

What is the meaningful difference between something truly being learned, and mere "copying" of "tiny pieces of material" that have been observed in the past?

Could learning exist in a vacuum universe, with nothing to copy?

Is there any LLM output that could convince the author of originality, or will it never be convincing? Can the author always tell whether he's interacting with one?

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#19
post #14

No one owns basic shapes. You can claim ownership to the complex geometry of your work - and that’s it. The exact complexity of said geometry is impossible to define - you’ll know it when you see it.

Shapes (and colors) can be trademarked, so yes, they can.

Re: ChatGPT is not ‘artificial intelligence.’ It’s theft

#20
post #5

Using this logic, every writer who has learned to read and write by reading books, or every artist who improved their craft by studying works, or every musician who learned the piano by practicing pieces, is also "stealing" in whatever they "originally" create due to learning via pattern recognition "tiny pieces of every work" in their data set. It's ridiculous to compare agents that generalize well to "stealing" pie…

I think this is a fallacy that I see a lot in recent AI discussions. An LLM is not the same as a human brain. You might see some superficial similarities in both being able to produce a block of text, but the method by which the text is produced is entirely different. For example, we can't download entire libraries of books instantly to our brains and then reproduce those books word for word in memory. Things that operate at different scales and by different methods should have different regulations, in the same way a bike or a car is regulated differently from a truck or another piece of heavy machinery.

Also, humans can be, and often are, found liable for copyright infringement or for piracy depending on how they conduct themselves. If a human was to reproduce a copyrighted book word for word, that would consist of copyright infringement regardless of whether it was done by rote memory, by copy and paste, or assisted by a black box LLM. Even if a human paraphrases another work they can still be found guilty of plagiarism if the paraphrase is still overly similar to the original source material. A human can also be guilty of copyright infringement if they use a copyright work as source material in certain ways. If I steal a stock image without paying for a license and add it in my Photoshop collage, I might be found to have pirated or infringed on the original image creator's property.

LLMs are trained on copyright data and can often reproduce that copyright data. It's an open question how we regulate this.

I personally think it would be fair for an artist or author to say their work was not licensed to be used in training a neural net or otherwise request to opt out.

Post reply on HN