OpenAI whistleblower found dead in San Francisco apartment
141–150 of 499 posts
Re: OpenAI whistleblower found dead in San Francisco apartment
#142There are some pretty callous comments on this thread. This is really sad. Suchir was just 26, and graduated from Berkeley 3 years ago. Here’s his personal site: https://suchir.net/ . I think he was pretty brave for standing up against what is generally perceived as an injustice being done by one of the biggest companies in the world, just a few years out of college. I’m not sure how many people in his position would…
If I'm a whistleblower in an active case and I end up dead before testifying, I absolutely DO want the general public to speculate about my cause of death.
The only benefit of turning it into gossip is to dissuade other whistleblowers, without the inconvenience of actually having to kill anyone.
Re: OpenAI whistleblower found dead in San Francisco apartment
#143Earlier quoted context omitted.
If I'm a whistleblower in an active case and I end up dead before testifying, I absolutely DO want the general public to speculate about my cause of death.
I would also most certainly have a dead man's switch releasing everything I know. I would have given it to an attorney along with a sworn deposition.
Re: OpenAI whistleblower found dead in San Francisco apartment
#144Normally the word "whistleblower" means someone who revealed previously-unknown facts about an organization. In this case he's a former employee who had an interview where he criticized OpenAI, but the facts that he was in possession of were not only widely known at the time but were the subject of an ongoing lawsuit that had launched months prior. As much as I want to give this a charitable reading, the only explana…
That is an exceedingly charitable read of these lawsuits.
Everyone knows LLMs are copyright infringement machines. Their architecture has no distinction between facts and expressions. For an LLM to be capable of learning and repeating facts, it must also be able to learn and repeat expressions. That is copyright infringement in action. And because these systems are used to directly replace the market for human-authored works they were trained on, it is also copyright infringement in spirit. There is no defending against the claim of copyright infringement on technical details. (C.f. Google Books, which was ruled fair use because of it's strict delineation of facts about books and the expressions of their contents, and provides the former but not a substitute for the latter.)
The legal defense AI companies put up is entirely predicated on "Well you can't prove that we did a copyright infringement on these specific works of yours!".
Which is nonsense, getting LLMs to regurgitate training data is easy. As easy at it is for them to output facts. Or rather, it was. AI companies maintain this claim of "you can't prove it" by aggressively filtering out any instances of problematic content whenever a claim surfaces. If you didn't collect extensive data before going public, the AI company quickly adds your works to it's copyright filter and proclaims in court that their LLMs do not "copy".
A copyright filter that scans all output for verbatim reproductions of training data sounds like a reasonable compromise solution, but it isn't. LLMs are paraphrasing machines, any such copyright filter will simply not work because the token sequence 2nd-most-probable to a copyrighted expression is a simple paraphrase of that copyrighted expression. Now, consider: LLMs treat facts and expressions as the same. Filtering impedes the LLM's ability to use and process facts. Strict and extensive filtering will lobotomize the system.
This leaves AI companies in a sensitive legal position. They are not playing fair in the courts. They are outright lying in the media. The wrong employees being called to testify will be ruineous. "We built an extensive system to obstruct discovery, here's the exact list of copyright infringement we hid". Even just knowing which coworkers worked on what systems (and should be called to testify) is dangerous information.
Sure. The information was public. But OpenAI denies it and gaslights extensively. They act like it's still private information, and to the courts, it currently still is.
And to clarify: No I'm not saying murder or any other foul play was involved here. Murder isn't the way companies silence their dangerous whistleblowers anyway. You don't need to hire a hitman when you can simply run someone out of town and harass them to the point of suicide with none of the legal culpability. Did that happen here? Who knows, phone & chat logs will show. Friends and family will almost certainly have known and would speak up if that is the case.
Re: OpenAI whistleblower found dead in San Francisco apartment
#145Earlier quoted context omitted.
> Anytime someone potentially possesses information that is damning to a company and that person is killed… the low probability of such an even being a random coincidence is quite low. You're running into the birthday paradox here. The probability of a specific witness dying before they can testify in a lawsuit is low. The probability of any one of dozens of people involved in a lawsuit dying before it's resolved is…
A 26yo dying is not "one of dozens," it's ~1/10,000 in the US (and likely much lower if we consider this guy's background and socioeconomic status).
I'm not going to pretend to know what the exact odds are, but it's going to end up way higher than 1/10k.
Re: OpenAI whistleblower found dead in San Francisco apartment
#146Earlier quoted context omitted.
There's also the output side: Perhaps outputs of generative AI should be ineligible for copyright.
That is the current position, weirdly enough.
At least in the US, a derivative work is a creative (i.e. copyrightable) work in its own right. Neither AI models nor their output meet that bar, so it's not clear what the infringing derivative work could be.
Re: OpenAI whistleblower found dead in San Francisco apartment
#147Earlier quoted context omitted.
> It's very naive to believe in 'European press'. To get the idea check Ukrainian war coverage. What you'll see first is how single sided it is. This is such a wild take from my POV, a person in the EU. Have you considered the possibility that the nearest imperialist power beginning to violently invade Europe again is likely to trigger a common reaction? This is one of those rare cases in modern history where there i…
I can explain a bit. Russians living in Empire of Evil can see all internet including US and EU news. At the same time 'Putin propaganda' channels are blocked in EU. In EU only one side is available. This creates an information bubble, as intended. Which is a basic crowd control technique used to drive public opinion. In this case to support the war. The result is obvious, EU polls show much stronger support than the…
Re: OpenAI whistleblower found dead in San Francisco apartment
#148There are some pretty callous comments on this thread. This is really sad. Suchir was just 26, and graduated from Berkeley 3 years ago. Here’s his personal site: https://suchir.net/ . I think he was pretty brave for standing up against what is generally perceived as an injustice being done by one of the biggest companies in the world, just a few years out of college. I’m not sure how many people in his position would…
If I'm a whistleblower in an active case and I end up dead before testifying, I absolutely DO want the general public to speculate about my cause of death.
Re: OpenAI whistleblower found dead in San Francisco apartment
#149Earlier quoted context omitted.
That is the current position, weirdly enough.
Indeed, and to me it's one of the reasons it's hard to argue that generative AI violates copyright. At least in the US, a derivative work is a creative (i.e. copyrightable) work in its own right. Neither AI models nor their output meet that bar, so it's not clear what the infringing derivative work could be.
For example, suppose if I photograph a copyrighted painting, and then started selling copies of the slightly-cropped photo. The output wouldn't have enough originality to qualify as a derivative work (let alone an original work) but it would still be infringement against the painter.