Live data from Hacker News

Things are about to get worse for generative AI

garymarcus.substack.com

531–540 of 769 posts

Re: Things are about to get worse for generative AI

#531

These don't seem all that difficult to fix to me. Most of the examples are not really generic, but are shorthand descriptions of well-known entities. "Video game plumber" is practically synonymous with "Mario" and anyone that has the slightest familiarity with the character knows this. Likewise, how difficult is it to just use descriptive tools to describe Mario-like images [1] and then remove these results from anyo…

It's going to be hard to remove every single "shorthand descriptions of well-known entities" or other prompts that can be used to generate copyrighted or trademarked content. Sure, if you're not deliberately trying to generate infringing content, you can probably remove or discard those results, the trouble is the people who will try to trick the AI to generate this content, blocking those people is going to be impos…

  What happens when an AI is used to make decisions where the user/victim is entitled to know exactly why the AI did what it did? From a business and legal perspective I think the current AI solutions are dangerous and should be used very sparsely, exactly because even the creators can't point to the exact pieces of information that caused the AI to make the choices it did.
I totally agree that we need explainability (probably through symbolic systems, not fancier models), but I think you’re overestimating how much more satisfying explanations from more traditional AI are. “My rules told me to do X” is a bit more helpful for a troubleshooting engineer than “my training data trained me to do X”, but from a ‘business and legal perspective’ the difference is much less pronounced IMO.

Both answers mean you did something wrong in creating the machine. The fault will always lie with the creator.

Re: Things are about to get worse for generative AI

#532

Earlier quoted context omitted.

And that tech was not destroyed by regulation. It was replaced by the superior tech of torrents.

The company, however, was destroyed. Along with any possibility for a similar company to exist (for very long).

https://www.napster.com/us/

Re: Things are about to get worse for generative AI

#533

Earlier quoted context omitted.

> Would you ever actually use generative AI to pirate something when you could just torrent it? While there may be an argument that generative AI is infringing copyright, it is not really a very good tool for it. How do you square this with literally the first image in the OP showing side by side GPT reproing copyrighted work? imo a good modern art project would be someone making a website that “archives” NYT article…

Here is a picture of Darth Vader: https://lumiere-a.akamaihd.net/v1/images/darth-vader-main_45... Please show me a prompt that reproduces it. Also to pass this test, it has to be just as easy as right clicking "download image" The images in the article are done in reverse. They find a prompt that shows a copyrighted character and then search for the matching image. That's not how piracy is done.

Woah this is really moving the goalposts and is pretty disingenuous. When I responded to your prompt about GPT being bad at reproducing copyrighted material with a counter example where it appears to in fact be good at it, you tell me that I must reproduce a specific image as easily as “clicking download image.”

Not what I was arguing and you’re not going to win many arguments with anyone who is paying attending by coming out of left field with only tangentially related demands.

Re: Things are about to get worse for generative AI

#534

No they are not. This is a negotiation tactic by the NYT to drive up the licensing price. Period. The Napster/Music Industry analogy has no resemblance to this situation. The only meaningful question that might be answered as a result of this is, what permission and access rights do crawlers have to content that is publicly and legally available.

Surely there's a meaningful question about copying and distributing content verbatim, which GPT has been shown to do.

Re: Things are about to get worse for generative AI

#535

Earlier quoted context omitted.

That's a really eloquent way of saying "It's already happening, so give up on it." I'm sure it works out great for taking action and solving problems.

Isn't this what most of the world is saying to environmental activists who argue that we should go back to pre-industrial levels of production to "save the Earth"? I for one think that indeed there are many cases like this where the only feasible way out is forward. The film GATTACA expressed this very human sentiment well: > You want to know how I did it? This is how I did it, Anton: I never saved anything for the s…

Which environmental activists are saying that? That's a pretty specific claim.

Re: Things are about to get worse for generative AI

#536
post #420

Earlier quoted context omitted.

It's already happening and most people like having AI more than the DMCA. Selling people on the idea that ML training is piracy to people who on average pirate content with no moral quandary will go nowhere.

People liked having Napster, but it didn't stop file sharing going from a big mainstream app to underground sites run out of Russia (or other places that ignore copyright law). Sure, you can download music/movies still, but it's not like the Napster days. "Generative AI" is obviously copyright infringement, so owners of the copyright will win in court. Either Microsoft will have to fight a mass of legal cases, some w…

> "Generative AI" is obviously copyright infringement

You're saying this as a matter of fact when it's not clear at all. We'll see what happens with the NYT case because it touches on all the major points.

It's gonna call into question all web scraping and indexing because they're also distillations of copyrighted content in the same manner.

Re: Things are about to get worse for generative AI

#537

No they are not. This is a negotiation tactic by the NYT to drive up the licensing price. Period. The Napster/Music Industry analogy has no resemblance to this situation. The only meaningful question that might be answered as a result of this is, what permission and access rights do crawlers have to content that is publicly and legally available.

The article does not mention napster, where did this reference come from?

Re: Things are about to get worse for generative AI

#538
Related ongoing thread:

NY times is asking that all LLMs trained on Times data be destroyed - https://news.ycombinator.com/item?id=38816944 - Dec 2023 (93 comments)

Also:

NY Times copyright suit wants OpenAI to delete all GPT instances - https://news.ycombinator.com/item?id=38790255 - Dec 2023 (870 comments)

NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 - Dec 2023 (84 comments)

The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 - Dec 2023 (861 comments)

The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/item?id=38781863 - Dec 2023 (11 comments)

Re: Things are about to get worse for generative AI

#539

Earlier quoted context omitted.

You won't get what you want with those sorts of deals. OK, say every artist gets $100, one time (exact amount varies but would not be much). Everything's properly licensed according to you and the artists are essentially no better off, and the models are now good enough to create new training data for the future and artists never see any more money. You've won, I guess?

Training AI on AI generated data doesn't add anything. The AI already has all the weights to generate the image, so you are at best just reinforcing the existing weights by weighing them more than others. The closest thing you could do is e.g. have a second model that does something novel like create a 3D model from a 2D image and then you try to animate the model and a third model verifies the quality of the output.…

My point is that say every artist gets some small token payment once, and then what? That's not enough to live on, so we're right back to square one and we've solved nothing.

Incidentally yes, training AI on AI output will work fine, as long as you have a signal of quality. For example, upvotes in a subreddit would work fine. But that's not crucial to my point, which is that what OP is asking for will accomplish exactly nothing.

Re: Things are about to get worse for generative AI

#540

Earlier quoted context omitted.

Your argument is nonsense. The junior artist in your hypothetical would have as much liability, if not more.

but would they have liability if they submitted their "output" to a senior artist, who immediately shot it down as obviously infringing? Surely not. It's not illegal to draw Mario - just illegal to make money off your drawing. I think the real question is whether OpenAI should be allowed to charge for generating infringing content. Even though the unit cost of the Mario drawing is negligible, the sum total of their i…

you don't have to make money off it, you just can't publish it, except as a parody or commentary or possibly a tutorial on how to draw mario if the judge is having a good day

but "making money = infringement" is folk wisdom. you could certainly say making money attracts attention and increases likelihood of legal action

Post reply on HN