Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

801–810 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#801

Earlier quoted context omitted.

> rent seeking media companies Rent seeking? Media companies that actually create content are rent seeking? Versus the garbage hallucinations AI creates?

The New York Times is dying company that is rent seeking here. Along time ago, their content was valuable, yet now you can't even give it away to researchers. I know because they tried to make a deal with my company, we passed because social media data is infinitely more valuable.

To me, your comment only reinforces the point that NYT's content is actually valuable, rather than valuable to rent seekers. But maybe you can give a bit more detail.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#802
post #674

Earlier quoted context omitted.

How about things that aren’t quite facts? Reviews, opinions, etc.

Illegal to have the same opinion as someone?

I was inarticulate. Imagine a business that goes to some trouble to review businesses or products. Can we lift those and serve them ourselves? Non facts…

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#803
post #95
post #76

Earlier quoted context omitted.

"Teaching" by copying source books word for word, would be copyright infringement; see, for example, the well-known issues around photocopying books or even excerpts. Also lying on source materials (e.g. telling students that some respected historian denies the Holocaust happened, when it's obviously not the case) is not "teaching" - it's defamation, and the NYT is absolutely right to pursue that angle too. Using LLM…

> Teaching" by copying source books word for word, would be copyright infringement; see, for example, the well-known issues around photocopying books or even excerpts. Incorrect. Educational use helps satisfy one of tests for fair use. Teachers can, in many cases, photocopy copyrighted work without infringing on that copyright.

Teachers can in some very limited cases photocopy very small chunks of copyrighted work. This also varies significantly from country to country; the starting position is that they cannot reproduce works in their entirety.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#804

Earlier quoted context omitted.

> suddenly everyone's a copyright lawyer Roll back 20+ years ago on Slashdot and you'll see the exact same thing. Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.

Legality aside I think copyright of digital things in the digital age is a net negative to humanity.

Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers.

This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing everything they possibly can to hang on for dear life. Society needs to move on already.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#805
post #772

Earlier quoted context omitted.

First, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it. Second, you ignored the "purely derivative" bit. You have to l…

> people who buy the CliffNotes version typically already have the original work as well (for example from coursework) Is there data that supports this? I’d be interested to know what % of people who buy a Cliffs Notes have already _bought_ the original.

I'm sure that somewhere, someone, has data confirming or denying this.

But, anecdotally, it's what I've seen to be the case.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#806

Earlier quoted context omitted.

I see the exact opposite - any open source model is going to become prohibitively expensive to train if quality data costs billions of dollars. We’re going to be left with the OpenAI’s and Google’s of the world as the only players in the space until someone solves synthetic data.

Exactly this. I work at a small web scraping company (so I might be a bit bias) and any small business can collect a fair, capable datasets of public data for model training, sentiment analysis or whatever today. If public data is stopped by copyright as this lawsuit implies that would just mean only giant corporations and pirates would be able to afford this. This would be a huge blow to open-source and research dev…

You may remember the Google Books lawsuit where Google was digitally copying the entirety of books and making them available online.

Google won that suit under fair-use as a massive searchable database was found to be transformative as well as the non-commercial nature.

So; if your web scraping companies goal is to allow people to bypass a paywall I suspect you'll have trouble in the future. If your web scraping company instead say allows people to do market analysis on how many people need a piano tuner in NYC and it doesn't do that by copying a NYT article doing original research I think you'll be fine.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#807

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I think we're suffering from an excluded middle when it comes to this kind of intellectual property. Naturally, most readers want to pay zero. Naturally, owners of the publication think it is probably worth a couple hundred dollars a year to be this well-informed.

The current arms race got us scrapers, and then paywalls, and then ad-blocking archivers ...

But in reality, I might drop a penny to read a NYT article. Maybe a nickel. There's no reasonable way of performing microtransactions right now. Everything is still in hefty increments, so nobody can work out what the market would bear.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#808

Earlier quoted context omitted.

A court in Japan will have no impact on the outcome of a copyright lawsuit in USA. Not to mention that it doesn't really matter how a Japanese court ruled since it's all governed by treaties anyway. They will change their laws if required to.

its not about applying laws across different countries its about a precedent. If you don't keep up with international competition, you lose.

Japan has the right idea about this matter.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#809

Earlier quoted context omitted.

This seems very false to me. Spotify is the prime example. They offer a good product that covers a 100% of my needs at a reasonable price. If that was an option for say UFC or engineering books, you bet I’d be subscribed. But being forced to read through some crappy reader software when I need the book source to take annotations in another software doesn’t work, so here we are. Same with the absurd pay per view busin…

This is also along the lines of how I think about things. If you make it convenient enough (compared to the alternative of paywall bypass or piracy) and provide enough overall/general value then I'm happy to subscribe. At the point where the experience degrades, or seems beyond the point of what one person could reasonably subscribe to, I basically just give up. Spotify hits this sweet spot where one subscription del…

I suppose part of the challenge here is that music and video content holds value much longer. Studios can invest in music and video content and see a return from the catalog over a long period of time as more enduring hits are produced and the duds fall away. But with news, they have to make the money on it now because yesterday’s news isn’t worth much no matter how expertly crafted.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#810
post #625

Earlier quoted context omitted.

Any publisher can opt out of google. Publisher also have substantial control over titles and snippets shown in google, whether an article appears in google news, etc Paraphrasing is also known as cloning and is often a copyright violation

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

That’s not the only reason. Google search is also transformative and non competitive with the underlying publications. And that is why the opt out is important. If you feel google competes with your site you don’t have to sue Google: just tell them to to away
Post reply on HN