Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

601–610 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#601
I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things).

Overall, current LLMs remind me of those bottom-feeder websites that do no original research--those sites that just find an article they like, lazily rewrite it, introduce a few errors, then maybe paste some baloney "sources" (which always seems to disinclude the actual original source). That mode of operation tends to be technically legal, but it's parasitic and lazy and doesn't add much value to the world.

All that aside, I tend to agree with the hypothesis that LLMs are a fad that will mostly pass. For professionals, it is really hard to get past hallucinations and the lack of citations. Imagine being a perpetual fact-checker for a very unreliable author. And laymen will probably mostly use LLMs to generate low-effort content for SEO, which will inevitably degrade the quality of the same LLMs as they breed with their own offspring. "Regression to mediocrity," as Galton put it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#602

Earlier quoted context omitted.

We all remember when Aaron Swartz got hit with a wire tapping and intent to distribute federal crime for downloading JSTR stuff right? It's really disgusting, IMO, that corporations that go above and beyond that sort of behavior are seeing NO federal investigations for this sort of behavior. Yet a private citizen does it and it's threats of life in prison. This isn't new, but it speaks to a major hole in our legal sy…

What happened to Aaron Swartz was terrible. I find that what he was doing was outright good. IMO the right reading isn't to make sure anyone doing something similar faces the same way, but to make the information far more free, whether it's a corporation using it or not. I don't want them to steamroll everyone equally here, but to not steamroll anyone.

I don't want them to steamroll everyone equally here, but to not steamroll anyone.

I think you're nissing the point, and putting cart before horse. If you ensure that corporations are treated as stringently as people are sometimes, the reverse is true. And that means your goal will presumably be obtained, as the corporate might, becomes the little guy's win.

All with no unjust treatment.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#603
post #148

When the US entered WWI, they couldn't build a plane despite inventing them. They had to buy planes from the French. Why? The Wright Brothers patent war [1]. This led to Congress creating a patent pool for avionics that exists to this day. Honestly, I get this feeling about these lawsuits about using content to train LLMs. Think of it this way: in growing up and learning to read and getting an education you read any…

Not silly at all when you actually read the article - "Photographing the Eiffel Tower at night is not illegal at all. Any individual can take photos and share them on social networks. But the situation is different for professionals. The Eiffel Tower's lighting and sparkling lights are protected by copyright, so professional use of images of the Eiffel Tower at night requires prior authorization and may be subject to…

All you're doing is repeating the fact. OP's point was, bluntly, "the situation is dumb".

I happen to agree on that one. What is the benefit of copyrighting the Eiffel Tower? The purpose of copyright is not to say you can always make money off of what you created. It is to incentivize the creation of new things by allowing you to exclusively make money off of it for a while before its benefits can go to broader society.

So what is the purpose of copyrighting the Eiffel tower? Would it not have been made if copyright wasn't in place? (obviously it would have because it was and the law wasn't in place yet). Second the claim is that the copyright is on the "lighting design" visible at night. Is the lighting design of the tower so unique that no-one else could come up with it? or is the lighting design necessitated by the structure of the tower itself?

I'd say given the structure of the tower which restricts the lights, there is nothing sufficiently remotely unique or different to warrant copyright of the lighting design. Almost any design on that tower would look about the same.

So how is society benefiting from copyrighting that lighting design?

Exclusivity deals are almost always a net loss for society. Which is why whenever you see one you should be questioning if it should be in place. Exclusive contracts are anti free-market. Now there are absolutely valid places where they are justified and should be in place - but they should be questioned by default.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#604
post #499

Earlier quoted context omitted.

It’s likely fair use.

I think we need a lot of clarity here. I think it's perfectly sensible to look at gigantic corpuses of high quality literature as being something society would want to be fair use for training an LLM to better understand and produce more correct writing... but the actual information contained in NYT articles should probably be controlled primarily by NYT. If the value a business delivers (in this case the information…

The case for copyright is exactly the opposite: the form of content (the precise way the NYT writers presented it) is protected. The ideas therein, the actual news story, is very much not protected at all. You can freely and legally read an NYT article hot off the press and go on air on Fox News and recount it, as long as you're not copying their exact words. Even if the news turns out to be entirely fake and invented by the NYT to catch you leaking their stuff, you still have every right to present the information therein.

This isn't even "fair use". The ideas in a work are simply not protected by copyright, only the form is.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#605
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

> the first being at the birth of modern search engines.

Why do you say that? Search engines would at least direct the viewer to the source. NYT gets 35%+ of its traffic from Google: https://www.similarweb.com/website/nytimes.com/#traffic-sour...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#607

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

I think the exact opposite is true: as long as AI depends critically on scrupulous news media to be able to generate info about current events, it is far more important to protect the news media than the AI training models. OpenAI could survive even if it had to pay the NYT for redistributing their works. But OpenAI can't survive if no one is actually reporting news fairly accurately. And if the NYT were to go bankrupt, all smaller players would have gone under looooong before.

In some far flung future where an AI can send agents to record and interpret events, and process news feeds and others to extract and corroborate information, this would greatly change. But probably in that world the OpenAI of those times wouldn't really bother training on NYT data at all.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#608
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

Anyone in Open Source or with common sense would agree that this is the absolute minimum that the models should be doing. Good comment.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#609
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

for what it's worth, i asked altman directly and he denied using libgen or books2, but also deferred to murati and her team on specifics. but the Q&A wasn't recorded and they haven't answered my follow-ups.

Why would he know the answer in the first place?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#610
post #583

Earlier quoted context omitted.

> And the metadata (metaknowledge?) would be larger than the knowledge itself. Because URLs are usually as long as the writing they point at?

I’m not an expert in AI training, but I don’t think it’s as simple as storing writing. It does seem to be possible to get the system to regurgitate training material verbatim in some cases, but my understanding is that the text is generated probabilistically. It seems like a very difficult engineering challenge to provide attribution for content generated by LLMs, while preserving the traits that make them more usefu…

Conceptually, it wouldn't be very hard to take the candidate output and run it through a text matching phase to see if there are ~exact matches in the training corpus, and generate other output if there are (probably limited to the parts of the training corpus where rights couldn't be obtained normally). Of course, it would be quite compute heavy, so it would add significantly to the cost per query.
Post reply on HN