Earlier quoted context omitted.
> ... for profit. for non-profit
1. Non-profit != "not making a profit." A non-profit can still earn monetary profit, and many do. 2. The non-profit OpenAI, Inc. company is not to be confused with the for-profit OpenAI GP, LLC [0] that it controls. OpenAI was solely a non-profit from 2015-2019, and, in 2019, the for-profit arm was created, prior to the launch of ChatGPT. Microsoft has a significant investment in the for-profit company, which is why…
The New York Times is suing OpenAI and Microsoft for copyright infringement
501–510 of 912 posts
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#502When I was studying AI in grad school many years ago getting good big data sets was always an issue. It never occurred to me to just copy one without permission.
If it is legal to simply index a website, then why shouldn't it be legal to train a model in the very same data?
Of course, websites should have some option for declining data mining for ML/AI purposes, in the same way the can decline scraping/indexing in the robots.txt file.
But that ship has kind of sailed, unless the courts decide otherwise.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#503Earlier quoted context omitted.
> Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights. Figure this out and you get to choose which AI lab you want to make seven figures at. It's a really difficult problem.
It’s likely first and foremost a resource problem. “How much different would the output be if that text hadn’t been part of the training data” can _in principle_ be answered by instead of training one model, training N models where N is the number of texts in the training data, omitting text i from the training data of model i , and then when using the model(s), run all N models in parallel and apply some distance me…
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#504Earlier quoted context omitted.
robots.txt is not meant to be a mechanism of communicating the licensing of content on the page being crawled nor is it meant to communicate how the crawled content is allowed to be used by the crawler. Edit: same applies to humans. Just because a healthcare company puts up a S3 bucket with patient health data with “robots: *” doesn’t give you a right to view or use the crawled patient data. In fact, redistributing i…
Furthering the S3 health data thought exercise: If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training? The difference between this ques…
Unequivocally, yes.
LLMs have proved themselves to be useful, at times, very useful, sometimes invaluable assistants who work in different ways than us. If sticking health data into a training set for some other AI could create another class of AI which can augment humanity, great!! Patient privacy and the law can f*k off.
I’m all for the greater good.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#505Earlier quoted context omitted.
Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
a computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#506Earlier quoted context omitted.
that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…
Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#507I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
> if NYT goes under a dozen similar outlets can replace them overnight Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#508Earlier quoted context omitted.
For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…
Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#509Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…
It’s likely fair use.
If your "fair use" substantially negatively affects the market for the original source material, which I think is fairly clear in this case, the courts wont look favorably on that.
Of course, I think this is a great test case precisely because the power of "Internet scale" and generative AI is fundamentally different than our previous notions about why we wanted a "fair use exception" in the first place.
Re: The New York Times is suing OpenAI and Microsoft for copyright infringement
#510I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…
> if NYT goes under a dozen similar outlets can replace them overnight Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.
Which part of journalism is AI going to impact most? Opinion pieces that contain no new information? Summarizing past events?