Earlier quoted context omitted.
For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…
We all remember when Aaron Swartz got hit with a wire tapping and intent to distribute federal crime for downloading JSTR stuff right? It's really disgusting, IMO, that corporations that go above and beyond that sort of behavior are seeing NO federal investigations for this sort of behavior. Yet a private citizen does it and it's threats of life in prison. This isn't new, but it speaks to a major hole in our legal sy…
The New York Times is suing OpenAI and Microsoft for copyright infringement
511–520 of 912 posts
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#512I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
There was outrage about Amazon removing DPReview site recently. But, it would be a common practice not to publish code/info, which could be used to train the model of another company. So, expect less open source projects, that companies just released because they were feeling like it could be good for everyone.
Actually, there is the use case that NYT would become more influential and important, because if 99% of all info is generated by AI and search is not working anymore, we would have to rely on the trusted sources to get our info. In the world of garbage, we would have to have some sources of verifiable human-generated info.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#513Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#514Earlier quoted context omitted.
The fact that copyright protection is far too long is entirely separate from the need for some kind of copyright protection to exist at all. All evidence suggests that it's completely impossible to live off your work unless you copyright it for some reasonable period, with the possible exception of performance art (music, theater, ballet). A writer or journalist just can't make money if any huge company can package t…
> All evidence suggests that it's completely impossible to live off your work unless you copyright it for some reasonable period Which evidence?
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#515Earlier quoted context omitted.
A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.
> There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net. But if it's possible for the neural net to memorize passages of text then surely it could also memorize where it got those passages of text from. Perhaps not with today's exact models and technology, but if it was a requirement then someone would figure out a way to do it.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#516I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
The way I see it, if the NYT goes under (one of the biggest newspapers in the world), all similar outlets also go under. Major publishers, both of fiction and non-fiction, as well as images, video, and all other creative content, may also go under. Hence, there is no more (reliable) training data.
Can I apply for YC with this idea?
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#517Earlier quoted context omitted.
I see a complete economic collapse unless creators start getting paid both for their data upfront, and paid royalties when their data is used in an LLM response
Copyright doesn’t protect data, it only protects expression.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#518Earlier quoted context omitted.
Any citation would be a good start. For complex subjects, I'm sure the citation page would be large, and a count would be displayed demonstrating the depth of the subject[3]. This is how Google did it with search results in the early days[1]. Most probable to least probable, in terms of the relevancy of the page. With a count of all possible results [2]. The same attempt should be made for citations.
Ok, now please cite the source of this comment you just made. It's okay if the citation list is large, just list your citations from most probably to the least probable.
[1] http://web.archive.org/web/20120608192927/http://www.google....
[2] https://steemit.com/online/@jaroli/how-google-search-result-...
[3] https://www.smashingmagazine.com/2009/09/search-results-desi...
[4] Next page
:)
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#519Earlier quoted context omitted.
For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…
Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.
You actually need a lot more than that. Most significantly, you need to have registered the work with the Copyright Office.
“No civil action for infringement of the copyright in any United States work shall be instituted until ... registration of the copyright claim has been made in accordance with this title.” 17 USC §411(a).
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#520I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
Couldn't disagree more strongly, and I hope the outcome is the exact opposite. I think we've already started to see the severe negative consequences when the lion's share of the profits get sucked up by very, very few entities (e.g. we used to have tons of local papers and other entities that made money through advertising, now Google and Facebook, and to a smaller extent Amazon, suck up the majority of that revenue). The idea that everyone else gets to toil to make the content but all the profits flow to the companies with the best AI tech is not a future that's going to end with the utopia vision AI boosters think it will.