Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

911–920 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#911
post #728

Earlier quoted context omitted.

Exactly this. I work at a small web scraping company (so I might be a bit bias) and any small business can collect a fair, capable datasets of public data for model training, sentiment analysis or whatever today. If public data is stopped by copyright as this lawsuit implies that would just mean only giant corporations and pirates would be able to afford this. This would be a huge blow to open-source and research dev…

research is fair use, also providing something amazing like Wikipedia is arguably educational (again fair use), reselling NYT articles on-demand via an API is by itself neither, so likely not free use

Fair use is irrelevant here as no small business would ever risk court dragging even though they are in the right. Especially since breaking ToS and "business damage" are easiest attachments to any lawsuit related to digital space.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#912

Earlier quoted context omitted.

People used to leave newspapers in the trash, on the train, all over the place. Anyone could pick them up and read for free. I think it's reasonable for folks to carry this attitude into the digital age. People feel like news is something to share, it's not the source of creative expression, it's facts and as such we feel entitled to know the facts about our world and what is happening that might affect us.

That newspaper was likely paid for by someone, and could only be read by one person at a time.

The news on a website is paid for by someone, else they would not be in business. And I only have one screen, I don't share it. The difference is physical vs. digital copies. A physical copy costs $.10 a digital copy costs $.0000001 (made up numbers), the business can take a loss of numerous digital copies before hitting the cost of one copy of the physical paper.

The problem is really that their business model sucks. They are working with fewer and fewer advertisers and much more competition and expecting business like they had before. And so we have a business that is attempting to fix itself with paywalls which don't work 100% of the time, but good enough to get the found newspaper analogy.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#913

Earlier quoted context omitted.

People used to leave newspapers in the trash, on the train, all over the place. Anyone could pick them up and read for free. I think it's reasonable for folks to carry this attitude into the digital age. People feel like news is something to share, it's not the source of creative expression, it's facts and as such we feel entitled to know the facts about our world and what is happening that might affect us.

No it isn’t reasonable and people not paying for that newspaper read anymore is the reason all news is sensationalist opinion pieces today.

BS, the reason we have sensational opinion pieces goes back before the internet. People are bored with mundane lives and love drama. Drama sells, 90s newsrooms found this out and cable news out competed established real journalism. The rest is a race to the bottom.

The internet simply exacerbated this as anyone could publish news on an equal platform to the big boys. Then we get paid-per-click, and that drives click-bait.

Stealing information absolutely is not responsible for that. People pay for junk, and that's the reason. We don't eat junk food because it's given away.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#914

Earlier quoted context omitted.

"Transformative" seems to fit a lot more that "Derivative". On the other hand, it's understandable why NYT is worried. OpenAI itself says that occupations like: Writers and Authors, Web and Digital Interface Designers, News Analysts, Reporters, and Journalists, Proofreaders and Copy Markers are "90-100% exposed" to what OpenAI is building.

I don't buy into all these "dangers". The advent of cars did not decrease the amount of drivers and introduced various new jobs, that were not available for a lot of people. And the rise of computers, did not make the workforce smaller but instead opened many more opportunities for a lot of people.

I think focusing on lost jobs is the wrong angle to take this in. It's in how the content being used is being compensated. OpenAi isn't paying writers to train their engine.

I don't care that the car replaced the horse carriage because it didn't need to compensate horses nor handlers to do so. AI being the newest iteration of scraping data from artists, writers, etc. to profit millions off of is directly using the "horse handler's" work. If these LLM's threw NYT a royalty to use their articles as training material, there wouldn't be a lawsuit.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#915
post #824

Earlier quoted context omitted.

That seems like a great way to destroy what is left of art as we know it. Anna Karenina is just numbers. In The Mood For Love is just numbers. Right. What do you propose is the business model for artists in the absence of copyright?

The market at large will determine that. If people value cheap AI generated images more than talented human curated art then that's what it will be. If a market exists to buy unique pieces where an artist put brush to canvas and priced their work at $1000 instead of the cheap $10 poster that can be mass produced then thats what it will be. If no one wants to pay $1000 for your unique piece, then the market has spoken…

None of that is what copyright protects. And it lessens the argument when you can argue that LLMs are essentially stealing a human art's work to be used to generate cheap images. Similar to how if you took commissioned art, printed out 1000 copies, and sold them for $1 a piece.

Copyright means that you need to at least pay that artist you stole from in some way, which the government enforces so artists don't stop creating.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#916
post #852

Earlier quoted context omitted.

How do you know what the value of the art will be before it's created? Guns N' Roses is a top 40 artist on Spotify nearly 35 years after producing an album. Should they not have been paid after 1991? If you argue that they were a popular band and therefore should have been paid accordingly up front, well what about their debut record, which sold 30 million copies? How would you predict that value before its creation…

You should be paid the accurate value of the labor . The pay should not scale more when no additional labor takes place. This is how art worked for millenia; someone commissions a chapel roof painting, someone commissions a concerto, someone commissions a statue, someone buys a chair, etc. Artists still do this today, and there is no issue determining value beforehand. Artists list their commission prices, or their h…

>You should be paid the accurate value of the labor.

that's gone out the door in the digital age. Compaies at this point have spent centuries trying to enfoce this model while witholding stuff like stock and royalties to take a part of what the company enjoys by protifting for decades off of a single (underpaid) piece of labor.

I don't exactly sympathize with a robot now trying to do the same. Pay your labor.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#917
post #887

Earlier quoted context omitted.

Alright, so then you're NOT going to pay artists for their labor? In the current system, artists might work for many years on a single work, or work many years perfecting their craft before anyone wants to pay for their work. Copyright gives them a way to earn money in the future that compensates them for the work they did in the past. It incentivizes creativity. Don't get me wrong, I don't think copyright is perfect…

Unfortunately, it’s hard to explain these things to techies who only see the world in their one-sided startupy way. The fact that there’re starving creatives who have already been massively marginalized by the likes of Spotify of this world, means nothing to these tech workers who only see everything as numbers, or a “business model” to “validate”. (full disclosure, I’m a techie who’s gradually woken up to the idea t…

I'm in games, where art and tech crossroad. I 1000% empathize for the fact that art exploits, abuses, and underpays even if they at times may be doing more work than a junior web dev.

It's a bit ironic, because a lot of tech offers partial compensation in stock. Something else that really doesn't happen in games unless you work for like, the 3-4 largest studios. So they should at least understand that your compensation is not all based on labor for time worked.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#918

Earlier quoted context omitted.

You should be paid the accurate value of the labor . The pay should not scale more when no additional labor takes place. This is how art worked for millenia; someone commissions a chapel roof painting, someone commissions a concerto, someone commissions a statue, someone buys a chair, etc. Artists still do this today, and there is no issue determining value beforehand. Artists list their commission prices, or their h…

Exactly. Somehow this idea that you keep getting paid for literally the same thing over and over again for work you did once is the absurdity. And ridiculously greedy. It seems to have been invented by laywers, for lawyers. Nobody else really benefits as much as they do. The whole entirety of society vs. a single profession of dubious morality.

>Somehow this idea that you keep getting paid for literally the same thing over and over again for work you did once is the absurdity

meanwhile, most tech is moving towards subscriptions?

Art is getting paid "non-greedily". People buy a song or art piece, and then people 10 years later buy a song or art piece. That's not one person paying twice for the same song, it's two people buying the same thing.

If people still value that art for that price later, I don't see how this is a "greedy" thing. is art magically supposed to turn open source CC0 after 5 years? Tech sure doesnt work like that.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#919
post #852

Earlier quoted context omitted.

How do you know what the value of the art will be before it's created? Guns N' Roses is a top 40 artist on Spotify nearly 35 years after producing an album. Should they not have been paid after 1991? If you argue that they were a popular band and therefore should have been paid accordingly up front, well what about their debut record, which sold 30 million copies? How would you predict that value before its creation…

> How do you know what the value of the art will be before it's created? I don't know. Anyone funding the work is accepting a risk. > Should they not have been paid after 1991? They definitely should get paid for their shows and live performances. The band itself can't be copied. Artists are extremely scarce. Their art, however, is not. Once created, the scarcity of their recordings is artificial and fundamentally ti…

>In other words, even if we accept copyright as legitimate, they sure as hell shouldn't still be getting paid for some late 80s album.

Why not? The fact is that even if the album is free, there will be people paying spotify $10/month to listen to it on demand. How is it fair that Spotify can profit from it for decades to come because they offer convenience, over the artist who made the music 10 years earlier and now relinquishes their art not even a quarter into a typical career?

Copyright is abusrd now, but it's not a bad concept. I think the original copyright law of 14 + 14 worked well enough. Life expectancy increased so I'd increase it to 14 + 14 + 14 (or 10 years after the death of the original author, whatever comes first). You fund an artist for their typical career length (if they choose to extend twice) and once they are (near) retired the song is free to work off of. In the meantime you simply negotiate if you want to use their work.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#920
post #702

Earlier quoted context omitted.

He's talking about citing and quoting NYTimes articles, not republishing them verbatim. That said, it's very different if you're a publication that sometimes cites reporting from other publications vs. a website exclusively dedicated to indexing and summarizing NYTimes articles.

I couldn't get gpt to quote an actual nyt article no matter how hard I tried...it just hallucinated in the general style of a news article. Presumably, if it can remember at least a paragraph or two of each article, then surely the same would be true of any text it ingested and the model size would approach the dataset size (probably actually much larger). I don't believe this is the case at all, even searching aroun…

I don't think it's necessarily true that model size would need to be larger than dataset size. It's theoretically possible that the model encoding achieves significantly better compression than DEFLATE or GZIP or whatever compression algorithm is used to store the dataset itself.
Post reply on HN