Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

851–860 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#851

Earlier quoted context omitted.

Copying is not theft. Stealing a thing leaves one less left Copying it makes one thing more; that’s what copying’s for.

Quite the hill to die on. But hear me on this: why doesnt openai allow training using its data? Or why doesnt microsoft train against windows’ source code (not that it would be of quality)? Or why can’t we just copy whatever movie and music we want through whatever protocol we want? Is it because your bosses know that copying without approval is theft?

It's not quite the hill to die on. There are (at least) two definitions of "theft" and "stealing" in common use:

1) I take something away from you. You have less of it as a result. Copying is not theft.

2) I deprive you of something, such as exclusive use of your land (e.g. by trespassing) or failing to follow through on a contract. Copying is theft.

Both of those are used by different communities, who both become angry at the other.

This is a semantic argument. Most members of both groups believe that there are times when copying is wrong, and are split on when that is.

However, to group #1, "stealing" and "theft" is a highly offensive term. It's much like saying "You raped me up the ___ when you didn't pay my contractor bill on time" or other hyperboles. Not paying my bill was wrong, but it also wasn't rape. It devalues rape, insults you, and is imprecise. You should use the precise "copyright violation" which describes exactly what happened.

To group #2, NOT calling it theft is offensive, since it devalues the costs to businesses and creators of copyright violations. Whether you agree with them or not, they have certain rights under the law, and picking-and-choosing which laws to follow is wrong (especially when it's self-serving).

Because the two groups mean different things by the same words, they can never hold a rational conversation with each other, and become offended when they hear the other group speak. It's how we polarize. It's unfortunate, since there's an important discussion to be had about the limits and enforcement of copyright and patents, which really should start with the copyright clause in the constitution, and when it helps versus impedes progress and economic growth. That's a discussion possible to have analytically and rationally.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#852

Earlier quoted context omitted.

Why don’t they train their AI on non-copyrighted material? It’s only fair for the copyright owners to want a share of the pie. I’d want one as well for my work.

>It’s only fair for the copyright owners to want a share of the pie. No it's not, it's pure greed. Everyone'd think it absurd if copyright holders dared to demand that any human who reads their publicly available text has to pay them a fee, but just because OpenAI are training a brain made of silicon instead of a brain made of carbon all the rent-seekers come out to try to take advantage.

> No it's not, it's pure greed.

And Altman (Mr. Worldcoin) and fucking Microsoft are what, some gracious angels building chatbots for the betterment of humanity? How is them stealing as much content as they can get away with not greedy, exactly?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#853
post #41
post #25

Earlier quoted context omitted.

NYTimes is an ad-supported business, so you visiting their website to read the content those ads pay for is important.

NYT is an anomaly. The majority of their revenue is actually from subscriptions.

And?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#854
post #159
post #25

Earlier quoted context omitted.

NYTimes is an ad-supported business, so you visiting their website to read the content those ads pay for is important.

My browser doesn’t display ads and shows little regard for most paywalls. Do I owe NYTimes something?

How is that germane to the point I'm making?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#855

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

We use LLMs for classification. When you have limited data, LLMs work better than standard classification models like random forests. In some cases, we found LLM generated labels to be more accurate than humans.

Labeling few samples, LoRA optimizing an LLM, generating labels on millions of samples and then training a standard classifier is an easy way to get a good classifier in matter of hours/days.

Basically any task where you can handle some inaccuracy, LLMs can be a great tool. So I don't think LLMs are a fad as such.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#856

Earlier quoted context omitted.

AI certainly isn’t a replacement for journalism, but that doesn’t mean journalism will continue to exist if no one pays for it. If everyone gets their news from chatGPT or the like there will be no investigative reporting. We’re already beginning to see this with most people reading the google/Facebook blurbs instead of clicking the link and giving ad money let alone paying.

> If everyone gets their news from chatGPT But I've just explained that ChatGPT can't actually produce news articles. I can't ask ChatGPT what happened today, and if I could it would be because a journalist went out and told ChatGPT what happened.

So at first ChatGPT will copy journalists. Then journalists will stop working because nobody pays them. Then there will be no news. Some people may look at that situation and decide to start a new news business but that business will fail because ChatGPT will immediately rip it off. The end game is just no news other than volunteers.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#857
post #436
post #229

Earlier quoted context omitted.

No. It is a technology that can massively accelerate human progress—I don’t buy the theory that OpenAI used NYTimes content that was not freely available to them. If you read all of the public internet you probably have lots of snippets of NYTimes articles. Regarding the reading of the wirecutter by a browser tool, I don’t know how much of it is available online without subscription (because I subscribe to the NYTime…

It may have been freely available to them, but that doesn’t mean that they are free to reproduce its contents or otherwise make use of it at scale.

The chatGPT model is not reproducing the contents nor making it available at scale. I, the user, ask the model to read the website and other websites and come back with a useful summary to me. This is not training data thus not violating copyright any more than a browser showing the full text would.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#858
post #840
post #709

Earlier quoted context omitted.

Eliminating the right to patient privacy does not serve the greater good. People have enough distrust of the medical system already. I’m ambivalent to training on properly anonymized health data but, i reject out of hand the idea that OpenAI et al should have unfettered access to identifiable private conversations between me and my doctor for the nebulous goal of some future improvement on llm models.

> unfettered access to identifiable private conversations You misread the post I was responding to. They were suggesting health data with PII removed. Second, LLMs have proved that AI which gets unlimited training data can provide breakthroughs in AI capabilities. But they are not the whole universe of AIs. Some other AI tool, distinct from LLMs, which ingests en masse as much health data as it can could provide heal…

> could outweigh an individual's right to privacy.

If that’s the case, let’s put it on the ballet and vote for it.

I’m tired of big tech making policy decisions by “asking for permission later” and getting away with everything.

If there truly is some breakthrough and all we need is everyone’s data, tell the population and sell it to the people and let’s vote on it!

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#859

Earlier quoted context omitted.

> If everyone gets their news from chatGPT But I've just explained that ChatGPT can't actually produce news articles. I can't ask ChatGPT what happened today, and if I could it would be because a journalist went out and told ChatGPT what happened.

So at first ChatGPT will copy journalists. Then journalists will stop working because nobody pays them. Then there will be no news. Some people may look at that situation and decide to start a new news business but that business will fail because ChatGPT will immediately rip it off. The end game is just no news other than volunteers.

> So at first ChatGPT will copy journalists.

You still literally have not explained how this works. ChatGPT could write a news article, but it's not going to actively discover new social phenomena or interview people on the street. Niche journalism will continue having demand for the sole reason that AI can't reliably surface new and interesting content.

So... again, how does a pre-trained transformer model scoop a journalist's investigation?

> Then journalists will stop working because nobody pays them.

How is that any different than the status-quo on the internet? The cost of quality information has been declining long before AI existed. Thousands of news publications have gone out of business or been bought out since the dawn of the internet, before ChatGPT was even a household name. Since you haven't really identified what makes AI unique in this situation, it feels like you're conflating the general declining demand for journalism with AI FOMO.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#860
A ChatGPT that was English language literate to the level of Victorian England, scientifically literate from library books and fed news from the Lincoln Journal Star (a Nebraska newpaper) would be more than sufficient for most of my needs.

Cut NYT out of the loop, fu*'em! Let them sell their own damned GPT and then charge them like crazy for the license.

Post reply on HN