Live data from Hacker News

Training Stable Diffusion from Scratch Costs <$160k

mosaicml.com

21–30 of 52 posts

Re: Training Stable Diffusion from Scratch Costs <$160k

#21
post #11

Earlier quoted context omitted.

The companies that wait for a 100% "clean" model are going to get left behind. e.g. ChatGPT launching despite Google, Meta and others already having very similar technology internally

There is already a class action lawsuit. The companies that move forward with "dirty" models can be wiped out by legal fees before they got off the ground.

[deleted]

Re: Training Stable Diffusion from Scratch Costs <$160k

#24
Note that this doesn't take into account the numerous iterations required to dial in the correct hyperparameters and model architecture, which could easily increase cost 5-10x.

> 256 A100 throughput was extrapolated using the other throughput measurements

Is it an indictment of their service that they couldn't afford 256 GPUs on their own cloud?

Re: Training Stable Diffusion from Scratch Costs <$160k

#26
post #20

Earlier quoted context omitted.

This will likely just entrench the companies with pockets deep enough to satisfy the lawyers in the class action suits. It might burn down a startup. On the other hand, if Microsoft thinks it's a potential $500 billion business, a $1 billion settlement is just table stakes.

This ignores that a legal remedy can be "you can no longer offer this as a product".

Has this ever happen in practice for well funded company outside Napster case? I am skeptical that training on publicly accessible data can be ruled illegal. Too many side effects including making Google Search problematic.

Re: Training Stable Diffusion from Scratch Costs <$160k

#28
post #20

Earlier quoted context omitted.

This will likely just entrench the companies with pockets deep enough to satisfy the lawyers in the class action suits. It might burn down a startup. On the other hand, if Microsoft thinks it's a potential $500 billion business, a $1 billion settlement is just table stakes.

This ignores that a legal remedy can be "you can no longer offer this as a product".

That likely wouldn't be on the table in a settlement offer, and it might be a tough sell for the plaintiff class to get nothing or a protracted legal battle instead of an easy and significant payout. Anything is possible, but I don't think a lawsuit's likely outcome in this situation would be a scorched earth fight.

Re: Training Stable Diffusion from Scratch Costs <$160k

#29
post #12
post #3

Still pretty pricey for average person, but these will trend cheaper and why I think it's futile to "regulate" AI. Someone somewhere will train models on anything visible to public, licensed or not. Feels like Pandora's box has been opened and we need to deal with it.

As usual, AI has no agency. I believe we should view AI as simply an extension of our own agency. Thus, if you prompt an AI to generate copyrighted work, that's fine. Viewing it yourself is like imagination. However, just as you can draw mickey mouse for your own fun all you want, you cannot sell such images.

That’s kindof already the case. Ignoring the legality of GitHub Copilot itself, if it suggests code that violates copyright and you use it, you are (probably most likely[a]) still infringing. You can’t hide behind “the AI did it.”

[a]: Of course, IANAL, and it’s up to the courts

Post reply on HN