Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

281–290 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#281

Earlier quoted context omitted.

Computers aren't humans and LLMs aren't human brains. We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts. Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay,…

So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question? It doesn’t seem clear to me that it does. Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in…

I don't think it matters how fancy you make your LLM, to be honest.

If you turn an LLM into a person, you may have a ethical and legal basis for treating that LLM like a person. There's no law about artificial intelligence being sentient or not, but law applies only to people, so that'd be the supreme court case of the century. I remember the Star Trek TNG episode about this topic and while the answer was perhaps more obvious with mister Data, the best arguments for and against synthetic consciousnesses have all been made in that episode.

The law doesn't care for how human-like programmers may think their program is, and neither should it in my opinion. What matters to the law is that the output is a result of an automated process, which comes with a completely separate set of rules and conditions compared to fair use.

The complexity of the program isn't a very good legal defence in my opinion because there's no clear line when the complexity would be enough to be considered human like. You could, for example, also claim that a computer is just very good at doing imitations, just like a person can be good at doing imitations on stage, and that an mp3 file is just an elaborate imitation act.

Even with a digital system identical to a physical system I don't think you can state human-ness as an argument. A tape recorder is just a sophisticated, automated way of sending an electric field through a magnet, similar to what a human can do with a dynamo and a spool of tape; a sort of delayed-action theremin, which would turn it into a musical signal. In turn, a neural network can be solved by human brain power if you pay enough people to work on a single iteration for an entire year. Almost everything a computer can do is just a sophisticated way of doing what humans are already doing, so I don't see why this is different when it comes to this topic.

There's no obvious "this is human like now" threshold and I doubt there will be until we know exactly how the human brain works.

When it comes to AI generated works, we don't currently know where the line between copyright violation and copyright exemption lies. If this goes through, it's the third major lawsuit of its type, the other two being actions against Stable Diffusion by artists to prevent it from producing derived works from their art.

IANA but I think you can assume that the "but computers are just like digital humans" approach won't fly in court. That's not really important, though; both sides of the coin are already having clever lawyers write up legal defences for their points of view, and it's more than likely that some other factors ("is a model a derived work" (probably) and "is the output of a model a derived work" (who knows!)) will decide the future of AI and copyright. The interesting thing is that academic research is essentially exempt from copyright law, so nobody can demand takedown of their content from research data sets, but whether the commercial branch of AI companies can use their academically generated models to serve their customers?

As an upside, I don't think the death of ChatGPT is the end of AI. This whole scenario could've easily been avoided if AI companies paid for their data set or had restricted themselves to works they had the license to (public domain, CC0, etc.), and OpenAI in particular has been pretty brazen in their "we'll see about it if it ever comes up" approach. Companies like Github are probably in the best place legally, where their users are already signing off on "we can take your content and do whatever the fuck we want" terms and conditions.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#282

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

Search engines have a big difference which is that they usually direct you to the original website.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#283

Earlier quoted context omitted.

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills. And I don't think it…

Seems like a rerun of the argument over snippets, or the "answer onebox" as Google used to call it where the info you need is directly inlined in the SERP rather than being behind a link.

Yes I would say that is very similar. Many sites saw their revenue crater, based on data they themselves had created.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#284
post #252

Earlier quoted context omitted.

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

Seems unfair as we converge on AGI. If I memorize the lyrics to a song, is that a copyright violation? The lyrics are encoded in the arrangement of my neurons, after all.

That's not what GP is saying. The (potentially) copyright infringing part is the reproduction of copyrighted material, not the encoding itself. In the same way that learning the lyrics to a song isn't copyright infringement, but performing that song live without permission would be.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#285

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

As previously said, search engines index and provide links. I’ll add that it constitutes fair use because a search engine isn’t itself a replacement for the articles that it indexes. But ChatGPT is actually providing an alternative that obviates the original articles themselves.

Google started moving away from just providing links along time ago. They routinely scrap data and show it, keeping people from visiting the links. I don't see how this behavior could be allowed while also crushing LLMs.

Personally, I like the flexibility of an LLM being able to describe a process at different skill levels. This is of tremendous educational value to the world.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#286

Earlier quoted context omitted.

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

Search engines have a big difference which is that they usually direct you to the original website.

Bing's LLM does that too.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#287

Earlier quoted context omitted.

Copywrite law doesn't ban people from reading the source altogether. That is what is being proposed, ban AI from being allowed to 'read' the content. The real argument is how does a human brain aggregate knowledge and then profit from it, and is it really that different from an AI model aggregating knowledge. They both read in data, perform calculations on the data, and spit out something.

Why should it be okay for humans to make a profit off of reading copyrighted content, though?

https://news.ycombinator.com/item?id=37159753

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#288

Earlier quoted context omitted.

Copywrite law doesn't ban people from reading the source altogether. That is what is being proposed, ban AI from being allowed to 'read' the content. The real argument is how does a human brain aggregate knowledge and then profit from it, and is it really that different from an AI model aggregating knowledge. They both read in data, perform calculations on the data, and spit out something.

This is a religious belief and not scientific knowledge or strong legal argument.

Are you projecting a religious argument onto LLM's?

It seems you are assuming that humans contains some metaphysical 'essence' like a soul that is doing the thinking?

Please connect the dots here.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#289

Earlier quoted context omitted.

ATMO you shouldn't have to maintain knowledge of what kind of crawler bot exist and having to maintain deny list. It should be the opposite, only expressedly allowed content should be crawled by mainaining allow lists.

You can do the opposite since the inception of robots.txt: User-agent: * Disallow: / and then whitelist google bot and whatnot. Most of the web is already configured this way. Just check robots.txt of any major website, e.g. https://twitter.com/robots.txt

The Allow: directive was an extension to robots.txt added later.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#290

Earlier quoted context omitted.

>>"If you do allow that, the many many affected industries have catastrophic problems." That is the problem. Technically, AI should be allowed to 'read' content, it isn't hidden, and it gets mixed with other content in a 'brain' like thing. AI and Humans can both spit out a new product that is 'similar' and thus be sued on that similarity. But it can also produce endless similar variations at low cost and fast. It is…

There’s an unbelievably vast difference between a human’s creative process and the mechanical reproduction of reweighed training data. Machines don’t create, people do.

Is there? I think you are giving too much credit to human creativity. Falling prey to the 'human exceptionalism' argument, or to the 'mysterious'. It just assumes humans are unique, or ineffable, it doesn't provide any of the 'how' it is done.

Ask GPT something like "write a screen play for Othello using dialog like Tarantino, but with bit of style like Baz Luhrmann". It produces something pretty creative. And if you say, well it is just combining what came before, that also applies to humans, there are no new ideas.

Humans being complicated doesn't mean AI wont catch up. Its just time/money/engineering at this point.

In Nature, why would Carbon be more holy than Silicon.

Post reply on HN