Live data from Hacker News

OpenAI departures: Why can’t former employees talk?

vox.com

431–440 of 1001 posts

Re: OpenAI departures: Why can’t former employees talk?

#431

The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…

Clever, but no. The argument about LLMs not being copyright laundromats making sense hinges the scale and non-specificity of training. There's a difference between "LLM reproduced this piece of copyrighted work because it memorized it from being fed literally half the internet ", vs. "LLM was intentionally trained to specifically reproduce variants of this particular work". Whatever one's stances on the former case,…

> In other words: GPT-4 gets to get away with occasionally spitting out something real verbatim. Llama2-7b-finetune-NYTArticles does not.

Based on what? This isn't any legal argument that will hold water in any court I'm aware of

Re: OpenAI departures: Why can’t former employees talk?

#432
post #236

It shouldn't be legal and maybe it isn't, but all schemes like this are, when you get down to it, ultimately about suppressing potential or actual evidence of serious, possibly criminal misconduct, so I don't think they are going to let the illegality get them all upset while they are having fun.

What crimes do you think have occurred here?

They don't say that criminal activity has occurred in this instance, just that this kind of behavior could be used cover it up in situations where that is the case. An example that could potentially be true. Right now with everything going on with Boeing, it sure seems plausible they are covering something(s) up that may be criminal or incredibly damaging. Like maybe falsify inspections and maintenance records? A person at Boeing who gets equity as part of compensation decides to leave. And when they leave, they eventually at some point in the future decide to speak out at a congressional investigation about what they know about what is going on. Should that person be sued into oblivion by Boeing? Or should Boeing, assuming what situation above is true, just have to eat the cost/consequences for being shitty?

Re: OpenAI departures: Why can’t former employees talk?

#433

Earlier quoted context omitted.

Clever, but no. The argument about LLMs not being copyright laundromats making sense hinges the scale and non-specificity of training. There's a difference between "LLM reproduced this piece of copyrighted work because it memorized it from being fed literally half the internet ", vs. "LLM was intentionally trained to specifically reproduce variants of this particular work". Whatever one's stances on the former case,…

How many sources do you need to steal from for it to no longer be considered stealing? Two? Three? A hundred?

Copyright infringement is not stealing.

Re: OpenAI departures: Why can’t former employees talk?

#434

The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…

Lol this would be a great performative piece. Although not so sure it'd stand up to scrutiny. Openai could probably take them to court on the grounds of disclosure of trade secrets or something like that and force them to reveal its training data and thus potentially revealing its source.

If they did so, they would open up themselves for lawsuits of people unhappy about OpenAI's own training data.

So they probably won't.

Re: OpenAI departures: Why can’t former employees talk?

#435

It probably would be better to switch the link from the X post to the Vox article [0]. From the article: “““ It turns out there’s a very clear reason for [why no one who had once worked at OpenAI was talking]. I have seen the extremely restrictive off-boarding agreement that contains nondisclosure and non-disparagement provisions former OpenAI employees are subject to. It forbids them, for the rest of their lives, fr…

Yet another ding against the "Open" character of the company.

Re: OpenAI departures: Why can’t former employees talk?

#436

So part of their compensation for working is equity, and when they leave thay have to sign an additional agreement in order to keep their previously earned compensation? How is this legal? Mine as well tell them they have to give all their money back too. What's the consideration for this contract?

That OpenAI are institutionally unethical. That such a young company can be become rotten so quickly can only be due to leadership instruction or leadership failure.

We already know there's been a leadership failure due to the mere existence of the board weirdness last year; if there has been any clarity to that, I've missed it for all the popcorn gossiping related to it.

Everyone including the board's own chosen replacements for Altman siding with Altman seems to me to not be compatible with his current leadership being the root cause of the current discontent… so I'm blaming Microsoft, who were the moustache-twirling villains when I was a teen.

Of course, thanks to the NDAs hiding information, I may just be wildly wrong.

Re: OpenAI departures: Why can’t former employees talk?

#437
When companies create rules like this, that tells me that they are very unsure of their product. Either it doesn't works as they claim, or it's incredible simple to replicate. It can also be that their entire business plan is insane, in any case, there's something basic wrong internally at OpenAI for them to feel the need for this kind of rule.

If OpenAI and ChatGPT is so far ahead for everyone else, and their product is so complex, it doesn't matter what a few disgruntled employees do or say, so the rule is not required.

Re: OpenAI departures: Why can’t former employees talk?

#438
post #377

Earlier quoted context omitted.

Is there an insightful summary of this proposal? The whole paper looks like 38 pages of non-rigorous prose with no clear procedure and already “aligned” LLMs will likely fail to analyze it. Forced myself through some parts of it and all I can get is people don’t know what they want so it would be nice to build an oracle. Yeah, I guess.

It's not a proposal with a detailed implementation spec, it's a problem statement.

“One framework proposed for superalignment” sounded like it does something. Or maybe I missed the context.

Re: OpenAI departures: Why can’t former employees talk?

#439
post #430

Earlier quoted context omitted.

If your training process ingests the entire text of the book, and trains with a large context size, you're getting more than just "a handful of word probabilities" from that book.

If you've trained a 16-bit ten billion parameter model on ten trillion tokens, then the mean training token changes 2/125 of a bit, and a 60k word novel (~75k tokens) contributes 1200 bits. It's up to you if that counts as "a handful" or not.

I think it’s questionable whether you can actually use this bit count to represent the amount of information from the book. Those 1200 bits represent the way in which this particular book is different from everything else the model has ingested. Similarly, if you read an entire book yourself, your brain will just store the salient bits, not the entire text, unless you have a photographic memory.

If we take math or computer science for example: some very important algorithms can be compressed to a few bits of information if you (or a model) have a thorough understanding of the surrounding theory to go with it. Would it not amount to IP infringement if a model regurgitates the relevant information from a patent application, even if it is represented by under a kilobyte of information?

Re: OpenAI departures: Why can’t former employees talk?

#440

Earlier quoted context omitted.

Most existing big tech datacenters use mostly carbon free or renewable energy. The vast majority of datacenters currently in production will be entirely powered by carbon free energy. From best to worst: 1. Meta: 100% renewable 2. AWS: 90% renewable 3. Google: 64% renewable with 100% renewable energy credit matching 4. Azure: 100% carbon neutral [1]: https://sustainability.fb.com/energy/ [2]: https://sustainability.a…

That's not a defense. If imaginary cloud provider "ZFQ" uses 10MW of electricity on a grid and pays for it to magically come from green generation, that means 10MW of other loads on the grid were not powered by green energy, or 10MW of non-green power sources likely could have been throttled down/shut down. There is no free lunch here; "we buy our electricity from green sources" is greenwashing bullshit. Even if they…

Who is going to decide what are a worthy uses of our precious green energy sources?
Post reply on HN