Live data from Hacker News

Things are about to get worse for generative AI

garymarcus.substack.com

701–710 of 769 posts

Re: Things are about to get worse for generative AI

#701
post #529
post #55

To me that’s the wrong question. Everyone knew it was trained on copyrighted material and capable of eerily similar outputs. But it’s already done. At scale. Large corps committing fully. There is no chance of that toothpaste going back in the tube. It’s a bit like when big tech built on aggressive user data harvesting. Whether it’s right, ethical or even legal is academic at this stage. They just did it - effectivel…

This comment is ignorant of history It happened with Napster, then Apple Music, now streaming services There is no widespread file sharing in the general public, instead we have devices that we don’t own, and streaming subscriptions Apple didn’t just copy all the music onto iPods and sell it — it took them a decade of deal making and lots of money to acquire the rights to the content I’m not saying what’s right or wr…

I've was never willing to riot over Napster - this is different.

This is one of the substantial jumps, I refuse to be cut out of this innovation.

Seriously, use Bing, try your free Photos built in generation system, they are rolling out GPT built into Word. Microsoft is easily the advanced tech company right now, as far what services can be provided to a consumer at scale. This is still like the alpha phase of all this. Apparently I talk to Copilot soon - that levels that up so much and it's already the best assistant I've ever had.

This is equivalent to trying to keep us all off smartphones and stuck on dumb phones I guess - I think you get what mean.

The NYT decided for all of us, the new smartphone equivalent thing is bad and we can't have it... that is something I'll riot over.

Re: Things are about to get worse for generative AI

#702
post #529

Earlier quoted context omitted.

This comment is ignorant of history It happened with Napster, then Apple Music, now streaming services There is no widespread file sharing in the general public, instead we have devices that we don’t own, and streaming subscriptions Apple didn’t just copy all the music onto iPods and sell it — it took them a decade of deal making and lots of money to acquire the rights to the content I’m not saying what’s right or wr…

I've was never willing to riot over Napster - this is different. This is one of the substantial jumps, I refuse to be cut out of this innovation. Seriously, use Bing, try your free Photos built in generation system, they are rolling out GPT built into Word. Microsoft is easily the advanced tech company right now, as far what services can be provided to a consumer at scale. This is still like the alpha phase of all th…

Just now I asked Copilot why my keyboard RGB lights were turned off every time I opened a game, that's almost verbatim - it told me exactly where to go and exactly what to turn off, took about 10 seconds to entirely search and correct the problem.

Re: Things are about to get worse for generative AI

#703
There are an alarming number of responses seemingly completely unaware of the core thrust of the article (and NYT lawsuit). ChatGPT was able to reproduce and publish significant portions of NYT articles, completely verbatim for hundred-to-thousand word stretches.

It’s not derivative work. We’re way past that. NYT has an exceptionally strong case here and anyone arguing about the merits of copyright is way off the mark. This court case is not going single-handedly to undo copyright. OpenAI has very little going for them other than “this is new, how were we to know it could do this”. So knowing that, the currently trained models are in a very sticky situation.

Further, I don’t see NYT settling. The implications are too large, and if they settle with OpenAI, they will have a similar case pop up with every other model. And every other publisher of digital content with have a similarly merited case. This is an inflection point for generative AI, and it’s looking like it will be either much more expensive or much more limited than we originally thought.

A side effect of this: I am predicting that we will start to see a rise in “pirate” models. Models who eschew all legality, who are trained in a distributed fashion, and whose weights are published not by corporations but by collectives (e.g. torrent models). There is a good chance we see these surpass the official “well behaved” models in effectiveness. It will be an interesting next few years to see this play out.

Re: Things are about to get worse for generative AI

#704
post #696

Earlier quoted context omitted.

The only problem with this view, is: > "the rest of us don’t care as long as the work is of good quality" Copyright protects Disney. But it also protects every creative author, no matter how disadvantaged, from mass shareholder driven behemoths. Today "Disney likes your work" is ear music. Without copyright it would be a death nell.

So how can we modify copyright so that it protects the little guy more than it protects Disney?

That's not how laws work.

Re: Things are about to get worse for generative AI

#705
post #184

Just make LLMs be like your average human and forget details. I know that it's easier to say than to do, but so are many things worth doing. I can't plagiarize - my language and visual memory doesn't work that way. Such an LLM will have to "create" and answer from more fuzzy memory.

The class of models that Yann Lecun is bullish on (look up I-JEPA) do exactly this.

But I-JEPA is non-generative. It does semantic image interpretation.

Okay, I guess it is related as my brain only does semantic image interpretation. (edit: my brain can create images, but only when I'm unconscious)

So with such a model, if you ask it to create an image, it would first create a semantic grammatical model of what you had asked for, and then perhaps draw it with colored pencils. I sort of like that. It's all that I could do. And it would be unlikely to violate any copyrights.

Re: Things are about to get worse for generative AI

#706

There are an alarming number of responses seemingly completely unaware of the core thrust of the article (and NYT lawsuit). ChatGPT was able to reproduce and publish significant portions of NYT articles, completely verbatim for hundred-to-thousand word stretches. It’s not derivative work. We’re way past that. NYT has an exceptionally strong case here and anyone arguing about the merits of copyright is way off the mar…

Such a thing happened with DALLE, Midjourney, and Stable Diffusion.

Stable Diffusion, when used to its fullest with thing like Control Net and LoRAs, blows the pants off of other proprietary models.

Re: Things are about to get worse for generative AI

#707

There are an alarming number of responses seemingly completely unaware of the core thrust of the article (and NYT lawsuit). ChatGPT was able to reproduce and publish significant portions of NYT articles, completely verbatim for hundred-to-thousand word stretches. It’s not derivative work. We’re way past that. NYT has an exceptionally strong case here and anyone arguing about the merits of copyright is way off the mar…

My guess is that OpenAI will be able to basically copy Google/YouYube on this and offer a system like content-ID. Specifically, ChatGPT doesn't reproduce copyrighted works by default; only by request/action of a third party user much like YouTube serving whatever videos people upload. It wasn't the intent of OpenAI to infringe copyright and in fact a lot of or most researchers believed the models were not overfitted enough to reproduce significant portions of arbitrary works.

Re: Things are about to get worse for generative AI

#708

Earlier quoted context omitted.

On the latter example, my question is whether anyone is actually using ChatGPT to read NYT articles. My understanding is that to produce the examples of word-for-word text in their lawsuit, they had to feed GPT the first half dozen paragraphs of the article and ask it to reproduce the rest. If you can produce the first half dozen paragraphs, you already have access to the article's text. Given that, is this theoretic…

I think it would be quite enough to prompt OpenAI with article title and author name. This is how LLMs are working.

I tried that a few different ways and couldn't get it to work. I don't think just the title and author are enough. I'd be interested to see if anyone else can find a prompt that does it.

Two of my attempts:

https://chat.openai.com/share/5cd17ff3-e142-4a7d-91c2-0b2479...

https://chat.openai.com/share/04fd722b-8b3c-469b-a1a2-d58e64...

Re: Things are about to get worse for generative AI

#709
post #589

Earlier quoted context omitted.

I think the world would be completely fine without a copyrighted C3PO or Robocop. George Lucas didn’t have billions of merchandising revenue in mind when working on his wild and thought unlikely to be successful science fiction movie in the 70s. Robocop was also a labor of love. We don’t really need Nth Star Wars sequel powered by those extra profits. The art form could be healthier overall.

Fine if they weren't copyrighted today, or ever? Because if copyright was eliminated the day Star Wars was released, other people would have copied the film reels and charged for entry, and Lucas would have hardly made a cent. Or if copyright was eliminated the day he went looking for funding, it wouldn't have ever been made. Personally, I think the world's a little richer for star wars's existence.

He may have made less money and heay have made more, with different monetization schemes.

Copyright is a monetization scheme, but it's not the only one.

In this imagined world, cinemas would have no movies to show, so they'd have to pay people like Lucas to create the films such that there'd be something to put on the screen. If many cinemas got together, and maybe got loans, they could pay for bigger budget films, too

Re: Things are about to get worse for generative AI

#710
post #470

Everybody just buying into the corporate narrative that anyone can actually own these sorts of things. Who truly owns the tales of Snow White and Cinderella? These stories didn't originate with Disney; they are part of a rich tapestry of folklore passed down through generations. Disney's success was partly built on adapting these existing narratives, which were once shared and reshaped by communities over centuries.…

> culture is a communal property Public domain / communal property is also part of copyright, so it's not as if this is some forgotten concept that needs to be restored to the discourse. Georgism is underconsidered, though. > By focusing solely on the legal implications and ignoring the historical context of cultural storytelling The legal implications are human implications and as much a part of culture as anything…

The incentives remain poorly aligned though. Otherwise the people who actually author the copyrighted works (actors, special effects artists, etc) wouldn't have had to go on strike for so long to get proper compensation.

The value still remains with the people who own the reproduction capabilities, and only scraps go to the artists. Artists can get scraps without selling copyright too, just look at patreon

Post reply on HN