OpenAI whistleblower found dead in San Francisco apartment
11–20 of 499 posts
Re: OpenAI whistleblower found dead in San Francisco apartment
#12Re: OpenAI whistleblower found dead in San Francisco apartment
#13When does generative AI qualify for fair use? by Suchir Balaji
Re: OpenAI whistleblower found dead in San Francisco apartment
#14Re: OpenAI whistleblower found dead in San Francisco apartment
#15Re: OpenAI whistleblower found dead in San Francisco apartment
#16Re: OpenAI whistleblower found dead in San Francisco apartment
#17I'm confused by the term "whistleblower" here. Was anything actual released that wasn't publicly known? It seems like he just disagreed with whether it was "fair use" or not, and it was notable because he was at the company. But the facts were always known, OpenAI was training on public copyrighted text data. You could call him an objector, or internal critic or something.
Re: OpenAI whistleblower found dead in San Francisco apartment
#18As much as I want to give this a charitable reading, the only explanation I can think of for using the word whistleblower here is to imply that there's something shady about the death.
Re: OpenAI whistleblower found dead in San Francisco apartment
#19I'm confused by the term "whistleblower" here. Was anything actual released that wasn't publicly known? It seems like he just disagreed with whether it was "fair use" or not, and it was notable because he was at the company. But the facts were always known, OpenAI was training on public copyrighted text data. You could call him an objector, or internal critic or something.
The article holds clues: "Information he held was expected to play a key part in lawsuits against the San Francisco-based company."
>In a Nov. 18 letter filed in federal court, attorneys for The New York Times named Balaji as someone who had “unique and relevant documents” that would support their case against OpenAI. He was among at least 12 people — many of them past or present OpenAI employees — the newspaper had named in court filings as having material helpful to their case, ahead of depositions.
Yes it's true it's been public knowledge that OpenAI has trained on copyrighted data, but details about what was included in training data (albeit dated ...), as well as internal metrics (e.g. do they know how often their models regurgitate paragraphs from a training document?) would be important.