Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

521–530 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#521
Someone should train some AI on decompiled code of Windows (not encouraging, but it would be interesting). Copyright is important for corpos when it protects their interests. Producing exact text as in NYT articles is pretty much a copyright violation. At least the last time the companies were trying to blame each other that their Java API implementations look pretty similar.

Even for open source code you cannot just remove the authors and license, replace some functions and say "oh, it is my code now". Only public domain code would allow these. But with copilot you could.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#522
post #331
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

for what it's worth, i asked altman directly and he denied using libgen or books2, but also deferred to murati and her team on specifics. but the Q&A wasn't recorded and they haven't answered my follow-ups.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#523

Earlier quoted context omitted.

Doesn't this harm open source ML by adding yet another costly barrier to training models?

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

setting legality as a cornerstone of ethics is a very slippery slope :)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#524
post #367

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures. I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a cha…

It does exist, and you'd be glad to know that it's going in the pro-AI/training direction: https://www.reedsmith.com/en/perspectives/ai-in-entertainmen...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#525
post #345

Earlier quoted context omitted.

It matters what ends up being best for humanity, and I think there are cases to be made both ways on this

People often get buried in the weeds about the purpose of copyright. Let us not forget that the only reason copyright laws exist is > To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries If copyright is starting to impede rather than promote progress, then it needs to change to remain constitutional.

Do other countries all use the same reasoning?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#526

Earlier quoted context omitted.

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

>Why isn't robots.txt enough to enforce copyright You actually need a lot more than that. Most significantly, you need to have registered the work with the Copyright Office. “No civil action for infringement of the copyright in any United States work shall be instituted until ... registration of the copyright claim has been made in accordance with this title.” 17 USC §411(a).

But the thing is, you can only bring the civil action forward after registering your claim but you need not register the claim before the infringement occurs.

Copyright is granted to the creator upon creation.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#527

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

If the NYT goes under, why would its replacement fare any better?

News media like NYT, Fox etc are tools for high scale brainwashing public by the elite. This is why you see all the News papers have some political ideology. If they were reporting on truth and not opinions they won't have the need for leaning. Also you never see the journalists reporting against their own publication.

Humanity is better off without these mass brainwashing systems.

Millions of independent journalists will be better outcome for humanity.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#528

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

They probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued. These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.

I’m ambivalent.

On the one hand, they should realize they are one of today’s horse carriage manufacturers. They’ll only survive in very narrow realms (someone has to build the Central Park horse carriages still), but they will be miniscule in size and importance.

On the other hand, LLMs should observe copyright and not be immune to copyright.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#529

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

The verbatim responses come as part of "Browse with Bing" not the model actually verbatim repeating articles from training data. This seems pretty different and something actually addressable.

> the suit showed that Browse With Bing, a Microsoft search feature powered by ChatGPT, reproduced almost verbatim results from Wirecutter, The Times’s product review site. The text results from Bing, however, did not link to the Wirecutter article, and they stripped away the referral links in the text that Wirecutter uses to generate commissions from sales based on its recommendations.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#530
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

A human can't credit the source of each element of everything they've learnt. AI's can't either, and for the same reason.

The knowledge gets distorted, blended, and reinterpreted a million ways by the time it's given as output.

And the metadata (metaknowledge?) would be larger than the knowledge itself. The AI learnt every single concept it knows by reading online; including the structure of grammar, rules of logic, the meaning of words, how they relate to one another. You simply couldn't cite it all.

Post reply on HN