Earlier quoted context omitted.
So, the law has this concept of 'de minimus' infringement, where if you take a very small amount - like, way smaller than even a fair use - the courts don't care. If you're taking a handful of word probabilities from every book ever written, then the portion taken from each work is very, very low, so courts aren't likely to care. If you're only training on a handful of works then you're taking more from them, meaning…
>we also thought sampling in music was de minimus I would think if I can recognize exactly what song it comes from - not de minimus.
OpenAI departures: Why can’t former employees talk?
511–520 of 1001 posts
Re: OpenAI departures: Why can’t former employees talk?
#512"It forbids them, for the rest of their lives, from criticizing their former employer. Even acknowledging that the NDA exists is a violation of it." I find it hard to understand that in a country that tends to take freedom of expression so seriously (and I say this unironically, American democracy may have flaws but that is definitely a strength) it can be legal to silence someone for the rest of their life.
Re: OpenAI departures: Why can’t former employees talk?
#513Earlier quoted context omitted.
If your training process ingests the entire text of the book, and trains with a large context size, you're getting more than just "a handful of word probabilities" from that book.
If you've trained a 16-bit ten billion parameter model on ten trillion tokens, then the mean training token changes 2/125 of a bit, and a 60k word novel (~75k tokens) contributes 1200 bits. It's up to you if that counts as "a handful" or not.
Re: OpenAI departures: Why can’t former employees talk?
#514Re: OpenAI departures: Why can’t former employees talk?
#515Earlier quoted context omitted.
Clever, but no. The argument about LLMs not being copyright laundromats making sense hinges the scale and non-specificity of training. There's a difference between "LLM reproduced this piece of copyrighted work because it memorized it from being fed literally half the internet ", vs. "LLM was intentionally trained to specifically reproduce variants of this particular work". Whatever one's stances on the former case,…
Seems absurd that somehow the scale being massive makes it better somehow You would think having a massive scale just means it has infringed even more copyrights, and therefore should be in even more hot water
Re: OpenAI departures: Why can’t former employees talk?
#516Earlier quoted context omitted.
If you've trained a 16-bit ten billion parameter model on ten trillion tokens, then the mean training token changes 2/125 of a bit, and a 60k word novel (~75k tokens) contributes 1200 bits. It's up to you if that counts as "a handful" or not.
xz can compress the text of Harry Potter by a factor of 30:1. Does that mean I can also distribute compressed copies of copyrighted works and that's okay?
Re: OpenAI departures: Why can’t former employees talk?
#517Earlier quoted context omitted.
Free speech is a much more general notion than anything having to do with governments. The first amendment is a US free speech protection, but it's not prototypical. You can also find this in some other free speech protections, for example that in the UDHR >Everyone has the right to freedom of opinion and expression; this right includes freedom to hold opinions without interference and to seek, receive and impart inf…
UDHR is not law so it's irrelevant to a question of law.
Re: OpenAI departures: Why can’t former employees talk?
#518The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…
Re: OpenAI departures: Why can’t former employees talk?
#519Earlier quoted context omitted.
Lol this would be a great performative piece. Although not so sure it'd stand up to scrutiny. Openai could probably take them to court on the grounds of disclosure of trade secrets or something like that and force them to reveal its training data and thus potentially revealing its source.
If they did so, they would open up themselves for lawsuits of people unhappy about OpenAI's own training data. So they probably won't.
Re: OpenAI departures: Why can’t former employees talk?
#520Earlier quoted context omitted.
Seems absurd that somehow the scale being massive makes it better somehow You would think having a massive scale just means it has infringed even more copyrights, and therefore should be in even more hot water
My US history teacher taught me something important. He said that if you are going to steal and don't want to get in trouble, steal a whole lot.