Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

561–570 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#561

Earlier quoted context omitted.

If the bar for copyright is as low as ordering of words, then I don't even know what to say.

So stop saying anything. Go learn how copyright works in the real world

What makes you think my opinions will change based on how legacy systems work in real world? Just because a stupid system exists doesn't mean it's correct.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#562

Earlier quoted context omitted.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

The legal argument, which I'm sure you are very well aware of, is that training a model on data, reorganizing, and then presenting that data as your own is copyright infringement.

Can you elaborate a bit more? That’s actually just a claim, not a legal argument.

Copyright law allows for transformative uses that add something new, with a further purpose or different character, and do not substitute for the original use of the work. Are LLM’s not transformative?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#563
post #429
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

Can't have your cake and eat it too.

1. If you run different software (LLM), install different hardware (GPU/TPU), and use it differently (natural language), to the point that in many ways it's a different kind of machine; does it actually surprise you that it works differently? There's definitely computer components in there somewhere, but they're combined in a somewhat different way. Just like you can use the same lego bricks to make either a house or a space-ship, even though it's the same bricks. For one: GPT-4 is not quite going to display a windows desktop for you (right-this-minute at least)

2. Comparing to humans is fine. Else by similar logic a robot arm is not a human arm, and thus should not be capable of gripping things and picking them up. Obviously that logic has a flaw somewhere. A more useful logic might be to compare eg. Human arm, Gorilla arm, Robot arm, they're all arms!

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#564

Earlier quoted context omitted.

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. In what sense are they claiming their generated contents as their own IP? https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-... > OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assig…

They can’t transfer rights to the output of it isn’t theirs to begin with. Saying they don’t claim the rights over their output while outputting large chunks verbatim is the old YouTube scheme of upload movie and say “no copyright intended”.

Exactly. And while one can easily just take down such a movie if an infringement claim is filed it’s unclear how one “removes” content from a trained model given how these models work. Thats messy.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#565
post #472
post #429

Earlier quoted context omitted.

a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.

When all the legal precedents we have are about humans, human analogies are incredibly relevant.

There is a hundred years of legal precedents in the realm of technology upsetting the assumptions of copyright law. Humans use tools - radios, xerox machines, home video tape. AI is another tool that just makes making copies way easier. The law will be updated, hopefully without comparing an LLM to a man.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#566
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

A human can't credit the source of each element of everything they've learnt. AI's can't either, and for the same reason. The knowledge gets distorted, blended, and reinterpreted a million ways by the time it's given as output. And the metadata (metaknowledge?) would be larger than the knowledge itself. The AI learnt every single concept it knows by reading online; including the structure of grammar, rules of logic,…

> And the metadata (metaknowledge?) would be larger than the knowledge itself.

Because URLs are usually as long as the writing they point at?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#567
post #531
post #480

Earlier quoted context omitted.

LLMs are not databases. There is no "citation" associated with a specific query, any more than you can cite the source of the comment you just made.

That's fine. Solve it a different way. OpenAI doesn't just get to steal work and then say "sorry, not possible" and shrug it off. The NYTimes should be suing.

And god willing if there is any justice in the courts NYTimes will lose this frivolous lawsuit.

Copyright law is a prehistoric and corrupt system that has been about protecting the profit margins of Disney and Warner Bros rather than protecting real art and science for living memory. Unless copy/paste superhero movies are your definition of art I suppose.

Unfortunately it seems like judges and the general public are so clueless as to how this technology works it might get regulated into the ground by uneducated people before it ever has a chance to take off. All so we can protect endless listicle factories. What a shame.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#568
post #518
post #485

Earlier quoted context omitted.

Ok, now please cite the source of this comment you just made. It's okay if the citation list is large, just list your citations from most probably to the least probable.

"Now displaying 3 citations out of ~150,000,000.." [1] http://web.archive.org/web/20120608192927/http://www.google.... [2] https://steemit.com/online/@jaroli/how-google-search-result-... [3] https://www.smashingmagazine.com/2009/09/search-results-desi... [4] Next page :)

This is not answering the GP question and does not count as a satisfactory ranked citation list. The first one is particularly dubious. Also you didn’t clarify which statement was based on which citation. I didn’t see “dog” in your text.

To help understand the complexity of an LLM consider that these models typically hold about 10,000 less parameters than the total characters in the training data. If one wants to instruct the LLM to search the web and find relevant citations it might obey this command but it will not be the source of how it formed the opinions it has in order to produce its output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#569

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

SciHub was an early warning, IMHO, that there's a strong risk of the first world fumbling the ball so badly with IP that tech ecosystems start growing in the third world instead. The dominant platform for distributing scientific journal papers is no longer Western. Maybe SciHub is economically inconsequential, but LLM's certainly are not! Imagine if California had banned Google spidering websites without consent, in…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#570
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.)

Hacker News consistently have upvoted posts to let users circumvent paywalls. And even when it doesn't, conversations here (and on Twitter, Reddit, etc.) that summarize the articles and quote the relevant bits as soon as the articles are published are much more of a threat to The New York Times than ChatGPT training on articles from months/years ago.

Post reply on HN