Live data from Hacker News

OpenAI departures: Why can’t former employees talk?

vox.com

561–570 of 1001 posts

Re: OpenAI departures: Why can’t former employees talk?

#561

Earlier quoted context omitted.

Clever, but the law is not a machine or an algorithm. Intent matters. Training an LLM with the intent of contravening an NDA is just plain . Everyone would still get sued anyway.

But then training a commercial model is done with the intent to not pay the original authors, how is that different?

Chutzpah. And that the companies doing it are multi-billion dollar companies who can afford the finest legal representation money can buy.

Whether the brazenness with which they are doing this will work out for them is currently playing out in the courts.

Re: OpenAI departures: Why can’t former employees talk?

#562

The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…

[deleted]

Re: OpenAI departures: Why can’t former employees talk?

#563

Earlier quoted context omitted.

That's not a defense. If imaginary cloud provider "ZFQ" uses 10MW of electricity on a grid and pays for it to magically come from green generation, that means 10MW of other loads on the grid were not powered by green energy, or 10MW of non-green power sources likely could have been throttled down/shut down. There is no free lunch here; "we buy our electricity from green sources" is greenwashing bullshit. Even if they…

> that means 10MW of other loads on the grid were not powered by green energy, or 10MW of non-green power sources likely could have been throttled down/shut down. No. Renewable energy capacity is often built out specifically for datacenters. > Even if they install solar on the roofs and wind turbines nearby - that's still electrical generation capacity that could have been used for existing loads. No. This capacity w…

Not the OP.

I agree with a majority of points you made. Exception is to this

> A figurative drop in the bucket.

Fresh water sources are limited. Fabs water demands and pollution are high impact.

Calling a drop in the bucket comes in the weasel words category.

We still need fabs, because we need chips. Harm will be done here. However, that is a cost we, as a society, will choose to pay.

Re: OpenAI departures: Why can’t former employees talk?

#564

The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…

[dead]

Re: OpenAI departures: Why can’t former employees talk?

#565
post #475

Earlier quoted context omitted.

OpenAI is outputting the partially copyright-infringing works of their LLM for profit. How does that square?

You, the user, is inputting variables into their probability algorithm that's resulting in the copyright work. It's just a tool.

Let's say a torrent website asks the user through an LLM interface what kind of copyrighted content they want to download and then offers me links based on that, and makes money off of it.

The user is "inputting variables into their probability algorithm that's resulting in the copyright work".

Re: OpenAI departures: Why can’t former employees talk?

#566

Earlier quoted context omitted.

So, the law has this concept of 'de minimus' infringement, where if you take a very small amount - like, way smaller than even a fair use - the courts don't care. If you're taking a handful of word probabilities from every book ever written, then the portion taken from each work is very, very low, so courts aren't likely to care. If you're only training on a handful of works then you're taking more from them, meaning…

You don't even need to go this far. The word-probabilities are transformative use, a form of fair use and aren't an issue. The specific output at each point in time is what would be judged to be fair use or copyright infringing. I'd argue the user would be responsible for ensuring they're not infringing by using the output in a copyright infringing manner i.e. for profit, as they've fed certain inputs into the model…

Is converting an audio signal into the frequency domain, pruning all inaudible frequencies, and then Huffman encoding it tranformative?

Re: OpenAI departures: Why can’t former employees talk?

#567

Earlier quoted context omitted.

That's not a defense. If imaginary cloud provider "ZFQ" uses 10MW of electricity on a grid and pays for it to magically come from green generation, that means 10MW of other loads on the grid were not powered by green energy, or 10MW of non-green power sources likely could have been throttled down/shut down. There is no free lunch here; "we buy our electricity from green sources" is greenwashing bullshit. Even if they…

> that means 10MW of other loads on the grid were not powered by green energy, or 10MW of non-green power sources likely could have been throttled down/shut down. No. Renewable energy capacity is often built out specifically for datacenters. > Even if they install solar on the roofs and wind turbines nearby - that's still electrical generation capacity that could have been used for existing loads. No. This capacity w…

> No. Renewable energy capacity is often built out specifically for datacenters

Not fully accurate. Indeed there is renewable energy that is produced exclusively for the datacenter. But it is challenging to rely only on renewable energy (because it is intermittent and electricity is hard to store at scale so often you need to consume electricity when produced). So what happens in practice is that the electricity that does not come from dedicated renewable capacity is coming from the grid/network. What companies do is that they invest in renewable capacity in the network so that "the non renewable energy that they consume at time t (because not enough renewable energy available at that moment) is offsetted by someone else consuming renewable energy later". What I am saying here is not pure speculation, look at the link to meta website, they are saying themselves that this is what they are doing

Re: OpenAI departures: Why can’t former employees talk?

#568

The best approach to circumventing the nondisclosure agreement is for the affected employees to get together, write out everything they want to say about OpenAI, train an LLM on that text, and then release it. Based on these companies' arguments that copyrighted material is not actually reproduced by these models, and that any seemingly-infringing use is the responsibility of the user of the model rather than those w…

[deleted]

Re: OpenAI departures: Why can’t former employees talk?

#569
post #283

Earlier quoted context omitted.

I think superalignment is absurd, and model "safety" is the modern AI company's "think of the children" pearl clutching pretext to justify digging moats. All this after sucking up everyone's copyright material as fair use, then not releasing the result, and profiting off it. All due respect to Jan here, though. He's being (perhaps dangerously) honest, genuinely believes in AI safety, and is an actual research expert,…

> I think superalignment is absurd, and model "safety" is the modern AI company's "think of the children" pearl clutching pretext to justify digging moats. All this after sucking up everyone's copyright material as fair use, then not releasing the result, and profiting off it. How can I be confident you aren't committing the fallacy of collecting a bunch of events and saying that is sufficient to serve as a cohesive…

An organisations intentions are always the same and very simple: “Increase shareholder value”

Re: OpenAI departures: Why can’t former employees talk?

#570

Earlier quoted context omitted.

But then training a commercial model is done with the intent to not pay the original authors, how is that different?

> done with the intent to not pay the original authors no one building this software wants to “steal from creators” and the legal precedent for using copyrighted works for the purpose of training is clear with the NYT case against open AI It’s why things like the recent deal with Reddit to train on their data (which Reddit owns and users give up when using the platform) are becoming so important, same with Twitter/X

> no one building this software wants to “steal from creators”

> It’s why things like the recent deal[s ...] are becoming so important

Sorry but I don't follow. Is it one or the other?

If they didn't want to steal from the original authors, why do they not-steal Reddit now? What happens with the smaller creators that are not Reddit? When is OpenAI meeting with me to discuss compensation?

To me your post felt something like "I'm not robbing you, Small State Without Defense that I just invaded, I just want to have your petroleum, but I'm paying Big State for theirs cause they can kick my ass".

Aren't the recent deals actually implying that everything so far has actually been done with the intent of not compensating their source data creators? If that was not the case, they wouldn't need any deals now, they'd just continue happily doing whatever they've been doing which is oh so clearly lawful.

What did I miss?

Post reply on HN