Earlier quoted context omitted.
> Except that Stack Overflow’s CEO, in this very article, says that it’s a violation of the Creative Commons license to train an LLM on their answers. Yes, because it's a license violation — "If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original". That includes derived data products, like AI models, built using the content.
…yet in the same article he’s talking about selling the data to LLM developers. It’s hard to make sense of.
Stack Overflow Will Charge AI Giants for Training Data
21–30 of 32 posts
Re: Stack Overflow Will Charge AI Giants for Training Data
#22Earlier quoted context omitted.
Except that Stack Overflow’s CEO, in this very article, says that it’s a violation of the Creative Commons license to train an LLM on their answers. So what he’s actually proposing is very unclear. > When AI companies sell their models to customers, they “are unable to attribute each and every one of the community members whose questions and answers were used to train the model, thereby breaching the Creative Commons…
> Except that Stack Overflow’s CEO, in this very article, says that it’s a violation of the Creative Commons license to train an LLM on their answers. Yes, because it's a license violation — "If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original". That includes derived data products, like AI models, built using the content.
> A key consideration in later fair use cases is the extent to which the use is transformative. In the 1994 decision Campbell v. Acuff-Rose Music Inc,[13] the U.S. Supreme Court held that when the purpose of the use is transformative, this makes the first factor more likely to favor fair use.[14] Before the Campbell decision, federal Judge Pierre Leval argued that transformativeness is central to the fair use analysis in his 1990 article, Toward a Fair Use Standard.[11] Blanch v. Koons is another example of a fair use case that focused on transformativeness. In 2006, Jeff Koons used a photograph taken by commercial photographer Andrea Blanch in a collage painting.[15] Koons appropriated a central portion of an advertisement she had been commissioned to shoot for a magazine. Koons prevailed in part because his use was found transformative under the first fair use factor.
Re: Stack Overflow Will Charge AI Giants for Training Data
#23So the "AI Giants" that have already trained models using SO / Reddit data will have a perpetual advantage over any newcomers trying to come up. So yeah, totally not a fan of this position from SO / Reddit. Anybody trying to democratize access to foundation models, who isn't (Google|OpenAI|Meta|Microsoft) is now going to find the on-ramp even steeper than ever. As if it wasn't bad enough just paying for compute time.…
Re: Stack Overflow Will Charge AI Giants for Training Data
#24Obviously it wont work, but I do wonder if this is a direction GPL4 should be looking into ...
Re: Stack Overflow Will Charge AI Giants for Training Data
#25Re: Stack Overflow Will Charge AI Giants for Training Data
#26Re: Stack Overflow Will Charge AI Giants for Training Data
#27Contemporary AI systems are now becoming human-competitive at general tasks,[3] and we must ask ourselves: Should we let machines flood our information channels with propaganda and untruth? Should we automate away all the jobs, including the fulfilling ones? Should we develop nonhuman minds that might eventually outnumber, outsmart, obsolete and replace us? Should we risk loss of control of our civilization? Such dec…
Its not easy to stop technology with regulation. At this point companies in many countries have started building their LLMs on the whole internet.
Just upgrade yourself)
Re: Stack Overflow Will Charge AI Giants for Training Data
#28So the "AI Giants" that have already trained models using SO / Reddit data will have a perpetual advantage over any newcomers trying to come up. So yeah, totally not a fan of this position from SO / Reddit. Anybody trying to democratize access to foundation models, who isn't (Google|OpenAI|Meta|Microsoft) is now going to find the on-ramp even steeper than ever. As if it wasn't bad enough just paying for compute time.…
Perhaps but stack overflow answers have a shelf life. Who cares how to fix an obscure React 3.0 issue these days?
Re: Stack Overflow Will Charge AI Giants for Training Data
#29Re: Stack Overflow Will Charge AI Giants for Training Data
#30SO content is user generated. It is a bit rich for them to put a wall around it and claim that it is chargeable for use as an AI training set by other companies (unless they also have a plan to share the income with their users).
Especially so given their stance on cash bounties for answers/sponsored questions etc.,
The mods on SO and StackExchange family of sites frown upon cash bounties for answers. But when SO itself wants to erect a paywall around the Question-Answer set for AI training, it is somehow clean and moral.
Smells like BS.
https://meta.stackoverflow.com/questions/251576/how-open-is-...
https://meta.stackexchange.com/questions/25615/offering-actu...
https://meta.stackexchange.com/questions/57850/pay-money-to-...