Live data from Hacker News

Things are about to get worse for generative AI

garymarcus.substack.com

181–190 of 769 posts

Re: Things are about to get worse for generative AI

#181
post #161

Earlier quoted context omitted.

If a child is instructed to read a copyrighted work at school, which later becomes a factor in his own derivative works, he won't be in breach of copyright. Why should other intelligent entities be prevented from reading copyrighted works and gaining whatever there is to gain from those works the way any human might?

This argument might hold more water when generative models are more than fancy compression algorithms/text completion engines. A more practical way of looking at this is: who is making money off of these models? How did they get their training data? I’m not a fan of copyright in general, but we have serious outstanding issues with companies and organizations stealing or plastering work without compensating the origin…

> This argument might hold more water when generative models are more than fancy compression algorithms/text completion engines.

I doubt that part of the argument would change even if we perfected brain uploads.

Now, if you gave the current LLMs a robot body with a cute face, that'll probably change minds faster, regardless of the underlying architecture.

> who is making money off of these models?

When the models are open source, or at least may be downloaded and used locally for no cost, that would be the users of the models.

And back to the biological comparison: I learned to read (and also to code) in part from the Commodore 64 user manual, should I owe the shareholders anything for my lifetime earnings? As I got to the end of that sentence, a thought struck me: taxes do that. And in the UK the question of if university should be funded by taxes or by the students themselves followed the same lines.

Re: Things are about to get worse for generative AI

#182
post #137
post #12

Or... things are about to get worse for copyright holders. I don't see any developped country pressing the brake on AGI in the near future to protect a few copyright holders from getting "stolen" in hypothetic scenarios.

> I don't see any developped country pressing the brake on AGI in the near future It’s already happening with EU AI Act https://www.europarl.europa.eu/news/en/headlines/society/202...

[deleted]

Re: Things are about to get worse for generative AI

#183
post #141

Earlier quoted context omitted.

Apple could buy most of the NYT, RIAA and MPAA companies combined with petty cash. The big ones are Disney and Sony with a combined market cap about 250b. Microsoft alone is worth over 10 times that.

Honestly I've always wondered what would happen (and how much the entertainment world would change) if a company like Apple, Google, Microsoft, etc did just that. Or heck, if it turns out you need the rights to train LLMs and its easier to do that with public domain stuff, they just flat out bought half the entertainment industry and assigned everything to the public domain. Every Disney work every for example.

No-one is going to buy a major media company and then throw the rights into the public domain. What they would do is buy the rights and then sue all competitors in the GenAI space.

Re: Things are about to get worse for generative AI

#184

Just make LLMs be like your average human and forget details. I know that it's easier to say than to do, but so are many things worth doing. I can't plagiarize - my language and visual memory doesn't work that way. Such an LLM will have to "create" and answer from more fuzzy memory.

The class of models that Yann Lecun is bullish on (look up I-JEPA) do exactly this.

Re: Things are about to get worse for generative AI

#185
post #135

Earlier quoted context omitted.

That's where I'm at. Dall+E spitting out C3PO should be entirely ok, unless I'm making money with the output, Disney should pound sand.

Put that c3p0 on a website that gets revenues from views and someone is getting paid.

Ok, sure, but that's not a GenAI thing, that's a plain old boring copyright thing. If I draw a bunch of C3POs and slap them on my Adwords website then I can expect a C&D letter post haste, who cares if the material in question came out of my pen, Photoshop or a GenAI model?

Re: Things are about to get worse for generative AI

#186

Earlier quoted context omitted.

thats irrelevant since an LLM is not an intelligent entity. Whatever you're arguing about is fiction.

ChatGPT is not an intelligent entity? What’s been comprehending and rewriting all my crappy code for several months? An auto-complete? There’s obviously emergent behavior there that is actually defined by the maker and most users as “intelligence.” Edit: typo

I would paraphrase one of Clarke's laws and say that "Any sufficiently advanced text generator is indistinguishable from an intelligent entity."

Just because a computer program's output is remarkably good does not mean there is any emergent intelligence, any more than a technology we don't understand means there is magic.

Re: Things are about to get worse for generative AI

#187

Earlier quoted context omitted.

> a few copyright holders By which you mean every copyright holder. > AGI in the near future Something that is purely speculative, undefined, and has been promised in the near future for 50+ years. I don't see copyright holders lying down for someone else's benefit and I don't see governments gutting copyright, contract law, and several other avenues of protection that copyright holders can deploy in the name of some…

If a child is instructed to read a copyrighted work at school, which later becomes a factor in his own derivative works, he won't be in breach of copyright. Why should other intelligent entities be prevented from reading copyrighted works and gaining whatever there is to gain from those works the way any human might?

If llms are intelligent entities legally equivalent to a human child, then they incur an even more serious legal problem, as we are all in violation of the 13th amendment.

Re: Things are about to get worse for generative AI

#188

Earlier quoted context omitted.

disclaimer: I work on GenAI at google, but views are my own The question is, how did the model create Mario&Luigi or Scrooge McDuck without training on copyrighted data? It can't just crawl Wikipedia because Fair Use in Wikipedia doesn't constitute Fair use for a commercial AI model. One possible outcome is more transparency on what datasets were used to train the models.

Disclaimer: ibid > It can't just crawl Wikipedia because Fair Use in Wikipedia doesn't constitute Fair use for a commercial AI model. Why not? The lawyers I've discussed this with socially think that questions like this are unresolved. There are certainly competing legal theories, but we're in uncharted territory. No one knows what the outcome will be until rulings come down or Congress acts. I find the NYT's argumen…

That's a good point. I agree it's not clear cut one way or another and we gotta let it play out.

Re: Things are about to get worse for generative AI

#189

Earlier quoted context omitted.

Two questions: (1) Do you think "developing AGI" a realistic, achievable goal? If so, what evidence do you see that we're making progress on the problem of "general" intelligence? Specifically, what does any of that have to do with Large Language Models? (2) Are there any "national security" applications of Large Language Models that you're aware of? It seems to me that it would be a very difficult case to make that…

Regarding (2), automating surveillance at scale. If you manage to put a bunch of listening devices at a place you're moderately interested in, a cafeteria at an enemy base for example, you might end up with literally hundreds of hours of conversations, most of them completely uninteresting, but a few that might possibly contain nuggets of information of the utmost importance. Listening to all these conversations requ…

I think we'd need to see these things get a lot more reliable for them to be viable in this use case. This seems like a "leaky net" as opposed to some more deterministic strategy (e.g. grepping large lists of keywords, or parallelizing the task over thousands of human analysts). When you're looking for a needle in a haystack you need to inspect every leaf and stalk.

So should we put copyright through the shredder on the wager that somehow generative techniques will find applications for mass surveillance?

Post reply on HN