Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

511–520 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#511
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

We all remember when Aaron Swartz got hit with a wire tapping and intent to distribute federal crime for downloading JSTR stuff right? It's really disgusting, IMO, that corporations that go above and beyond that sort of behavior are seeing NO federal investigations for this sort of behavior. Yet a private citizen does it and it's threats of life in prison. This isn't new, but it speaks to a major hole in our legal sy…

What happened to Aaron Swartz was terrible. I find that what he was doing was outright good. IMO the right reading isn't to make sure anyone doing something similar faces the same way, but to make the information far more free, whether it's a corporation using it or not. I don't want them to steamroll everyone equally here, but to not steamroll anyone.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#512

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

oh, sure. NYT could go, we could replace it with AI generated garbage with non-verifiable information without sources. AI changed the landscape. Google search would be working less reliable because real publishers would be hiding info behind the login (twitter/reddit). I.e. sites would be harder to index. There would be a lot of AI generated garbage which would be hard to filter out. AI generated review articles, AI generated news promoting someones agenda. Only to have a chatgpt which could randomly increase their price 100 times anytime in the future.

There was outrage about Amazon removing DPReview site recently. But, it would be a common practice not to publish code/info, which could be used to train the model of another company. So, expect less open source projects, that companies just released because they were feeling like it could be good for everyone.

Actually, there is the use case that NYT would become more influential and important, because if 99% of all info is generated by AI and search is not working anymore, we would have to rely on the trusted sources to get our info. In the world of garbage, we would have to have some sources of verifiable human-generated info.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#513
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

So you're advocating giving open AI and incumbents a massive advantage by now delegitimizing the process? It's kinda like why Netflix was all for "fast lanes"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#514

Earlier quoted context omitted.

The fact that copyright protection is far too long is entirely separate from the need for some kind of copyright protection to exist at all. All evidence suggests that it's completely impossible to live off your work unless you copyright it for some reasonable period, with the possible exception of performance art (music, theater, ballet). A writer or journalist just can't make money if any huge company can package t…

> All evidence suggests that it's completely impossible to live off your work unless you copyright it for some reasonable period Which evidence?

The fact that it has never been done successfully outside performance arts.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#515

Earlier quoted context omitted.

A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

> There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net. But if it's possible for the neural net to memorize passages of text then surely it could also memorize where it got those passages of text from. Perhaps not with today's exact models and technology, but if it was a requirement then someone would figure out a way to do it.

Neural nets don't memorize passages of text. They train on vectorized tokens. You get a model of how language statistically works, not understanding and memory.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#516

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

The way I see it, if the NYT goes under (one of the biggest newspapers in the world), all similar outlets also go under. Major publishers, both of fiction and non-fiction, as well as images, video, and all other creative content, may also go under. Hence, there is no more (reliable) training data.

Great. I will start a company to generate training data then. I will hire all those journalists. I won't make the content public. Instead I will charge OpenAI/Tesla/Anthropic millions of dollars to give them access to the content.

Can I apply for YC with this idea?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#517
post #156

Earlier quoted context omitted.

I see a complete economic collapse unless creators start getting paid both for their data upfront, and paid royalties when their data is used in an LLM response

Copyright doesn’t protect data, it only protects expression.

While I didn't say anything about copyright (obviously our current copyright laws are completely ill-equipped to handle how LLMs work), feel free to replace data with whatever you like. writing, art, music, etc. It's all the same.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#518
post #485
post #410

Earlier quoted context omitted.

Any citation would be a good start. For complex subjects, I'm sure the citation page would be large, and a count would be displayed demonstrating the depth of the subject[3]. This is how Google did it with search results in the early days[1]. Most probable to least probable, in terms of the relevancy of the page. With a count of all possible results [2]. The same attempt should be made for citations.

Ok, now please cite the source of this comment you just made. It's okay if the citation list is large, just list your citations from most probably to the least probable.

"Now displaying 3 citations out of ~150,000,000.."

[1] http://web.archive.org/web/20120608192927/http://www.google....

[2] https://steemit.com/online/@jaroli/how-google-search-result-...

[3] https://www.smashingmagazine.com/2009/09/search-results-desi...

[4] Next page

:)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#519
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

>Why isn't robots.txt enough to enforce copyright

You actually need a lot more than that. Most significantly, you need to have registered the work with the Copyright Office.

“No civil action for infringement of the copyright in any United States work shall be instituted until ... registration of the copyright claim has been made in accordance with this title.” 17 USC §411(a).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#520

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> I hope this results in Fair Use being expanded to cover AI training.

Couldn't disagree more strongly, and I hope the outcome is the exact opposite. I think we've already started to see the severe negative consequences when the lion's share of the profits get sucked up by very, very few entities (e.g. we used to have tons of local papers and other entities that made money through advertising, now Google and Facebook, and to a smaller extent Amazon, suck up the majority of that revenue). The idea that everyone else gets to toil to make the content but all the profits flow to the companies with the best AI tech is not a future that's going to end with the utopia vision AI boosters think it will.

Post reply on HN