Live data from Hacker News

Things are about to get worse for generative AI

garymarcus.substack.com

491–500 of 769 posts

Re: Things are about to get worse for generative AI

#491

Earlier quoted context omitted.

> And OpenAI quite literally sells access to their models, and if those models are pushing out verbatim copyrighted works as has been alleged by the NYT, then they are by definition reselling copyrighted works without permission. This style of argument has been previously made regarding things like torrenting during the heyday of piracy ("why would you need except for illegal purposes!") In my opinion, it's the exact…

> They aren't encouraging people to misuse them and it is solely on the user's shoulders for their choice to use them in a way that would cause infringement if the result is used commercially. I agree in principle, but that they can in the first place, especially when it accidentally happens, and at such massive scales more importantly, is the issue methinks. And no one's talking about abolishing the AIs here, we're…

You won't get what you want with those sorts of deals.

OK, say every artist gets $100, one time (exact amount varies but would not be much). Everything's properly licensed according to you and the artists are essentially no better off, and the models are now good enough to create new training data for the future and artists never see any more money.

You've won, I guess?

Re: Things are about to get worse for generative AI

#492

Earlier quoted context omitted.

Copyright should be the problem of the person using the works and not the problem of the AI generating it. Unless Nintendo plans on busting down the doors of every person who tries to draw Mario or preventing little Timmy from making a parody of Coca-Cola, making it where AI cannot generated copyrighted works is insane imo. Those brands should be proud to be such a big part of the cultural fabric that it is difficult…

When I cover generative AI in my Ethics in AI lecture, one of few soapbox opinions I give is that GenAI is doing essentially what people do - copy others. Picasso has a quote about "Good Artists copy, Great Artists steal", which doesn't mean try to pass Lario and Muigi off as your own, but rather that great artists are able to take aspects from other works (also called 'inspiration') without being caught. My personal…

Question: That is a point that would protect GPT models in the abstract, but that doesn't hold for OpenAI and Microsoft that provide "Image generation as a service"? The actual implementation is irrelevant, if must not be able to provide images that are infringing copyrights? (Just like a designer in an agency cannot use Mario for a print).

So using a model running on my laptop to generate a "Mario like" image would be fine, but it would make monetizing this difficult?

Re: Things are about to get worse for generative AI

#493
post #470

Everybody just buying into the corporate narrative that anyone can actually own these sorts of things. Who truly owns the tales of Snow White and Cinderella? These stories didn't originate with Disney; they are part of a rich tapestry of folklore passed down through generations. Disney's success was partly built on adapting these existing narratives, which were once shared and reshaped by communities over centuries.…

Copyright has never been based on a moral stance. It has always been determined by the lobbying power of various groups.

The idea that we should dispense with it to let generative AI companies make even more money seems totally bizarre.

Re: Things are about to get worse for generative AI

#494

I am constantly suprised by the amount of apologizing for generative AI infringement here. The fact that it's already being done and is a technical breakthrough is irrelevant to existing copyright law. "We are big and innovative" may hold weight with legislators, but it won't with the courts. Remember when everyone and their dog discovered sampling in the late 80's and they all thought they could get away with it bec…

[deleted]

Re: Things are about to get worse for generative AI

#495
post #467

I did an interesting thing and looked at how well the Llama2 models could compress text. For example, I took the first chapter of the first Harry Potter book and recorded the index of the 'correct' predicted token. The original text, compressed with 7zip (LZMA?) to about 14kB. The Llama2 encoded indexes compressed to less than 1kB. Then, of course, I can send that 1kB file around and decode the original text. (Unless…

This is a little confusing. You turned the text into indices? So numbers? Then compressed that? Or the text as numbers without any extra compression is only 1kb?

The tokenizer the models use,(sentence piece) is more or less based on one way to do compression.(bpe). It's not really clear what your testing.

Re: Things are about to get worse for generative AI

#497

Earlier quoted context omitted.

Wouldn’t this be well handled by suing the person that prompted and distributed the results?

No, whoever operates the LLM service is liable for unauthorized modification, reproduction and distribution of copyrighted work to users.

Ok, so is it ok if I run the whole thing on my own hardware, and never distribute?

If not, how does that differ from me making an unauthorized pencil drawing of Mario?

Re: Things are about to get worse for generative AI

#498

Earlier quoted context omitted.

It's going to be hard to remove every single "shorthand descriptions of well-known entities" or other prompts that can be used to generate copyrighted or trademarked content. Sure, if you're not deliberately trying to generate infringing content, you can probably remove or discard those results, the trouble is the people who will try to trick the AI to generate this content, blocking those people is going to be impos…

I don't understand why some peoole thinks any infringing content can be singled out and removed. Aren't LLMs giant coefficient matrices like, a punched out croissant dough, made of all training data plyed over? How can you say you can remove one specific ply out of dough and declare that ever potential effect that the offending ply had created is now completely removed?

The "reasonable" singular removal is more about coming up with ways to block prompts that can produce infringing content, and having filters on the other end to catch infringements before they are published to the user. It's an endless whack-a-mole that never actually addresses the problem but might look good enough to the legal system or to keen supporters.

Barring some major breakthrough, the actual answer is to train a new model without the infringing data.

I think some of the people saying "remove it from your model" are aware of this and are simply being glib and needling; "you've created this infringement monstrosity, so surely you made sure to include a way to deal with this problem without throwing away all of your work, right?"

Re: Things are about to get worse for generative AI

#499
The NYTimes case is a clear one because they are delivering nearly the same content as an end product to users. The others seem like dead ends. The infringer would be the prompter, not the AI which operates more like a search engine. This is Napster all over again, what a phenomenal waste of time and money, where the artist will definitely come out with 0 at the end of it and a few corporations control everything - not to mention, there's nothing stopping anyone from releasing a tool that will crawl all spongebobs, generate your model for you and allow you to produce locally copyright infringing material it to your hearts content locally. You could drown yourself in local spongebobs.

Re: Things are about to get worse for generative AI

#500

Earlier quoted context omitted.

When I cover generative AI in my Ethics in AI lecture, one of few soapbox opinions I give is that GenAI is doing essentially what people do - copy others. Picasso has a quote about "Good Artists copy, Great Artists steal", which doesn't mean try to pass Lario and Muigi off as your own, but rather that great artists are able to take aspects from other works (also called 'inspiration') without being caught. My personal…

Irrespective of the current legal situation, there is no reason to regulate machines the same way as humans (and it is generally not what happens).

I hope no self-respecting instructor in ethics could with a straight face teach how an LLM is like a human being when it comes to copyright while glossing over the blinding implication that if it truly were so we would then be subjecting that being to unthinkable abuse.

That hypocritical, self-contradictory take is transparently geared to benefit commercial LLM operators (at the expense of individuals who stand to suffer material harm and/or authored the very creative works thanks to which the tool even exists).

Post reply on HN