Live data from Hacker News

AI Data Laundering

waxy.org

11–20 of 120 posts

Re: AI Data Laundering

#11
post #6
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

Do we allow artists to withhold their works from the minds of eager, learning children? [1] Tell me how ML is different than the mind of a toddler ravenous for new information. For every billion dollar start-up using data at scale, there are tens of thousands more researchers and hobbyists doing the exact same, producing wonderful results and advances. If we stop this growth dead in the tracks, other countries more w…

Well a toddler isn’t making money off the information they are absorbing for one. If these are open to the public models that is one thing. But no, these are proprietary models whose sole purpose is to make money for large corporations.

Re: AI Data Laundering

#12
post #6
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

Do we allow artists to withhold their works from the minds of eager, learning children? [1] Tell me how ML is different than the mind of a toddler ravenous for new information. For every billion dollar start-up using data at scale, there are tens of thousands more researchers and hobbyists doing the exact same, producing wonderful results and advances. If we stop this growth dead in the tracks, other countries more w…

> Tell me how ML is different than the mind of a toddler ravenous for new information.

If a person published a work that clearly plagiarized or violated a patent, that person would be open to legal action.

I’m all for systemic change, but uses like this may end up having a chilling effect on human-created work.

Re: AI Data Laundering

#14
post #8

Earlier quoted context omitted.

This is larger than publishers, this is every artist, film-maker, photographer, every writer, every engineer, anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. Comment generated with gpt-neox prompt: Comment about AI and data collection and generation and it…

> anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. This has been the case ever since people started putting their art on the Internet publicly. The only difference is that now it's algorithms creating the derivatives, not people.

This is not remotely the same, scale and barrier to entry matter. With stable diffusion I can pick any artist right now and create over 1000 derivative works by tomorrow morning in his style to the same degree of expertise with no training involved and no work required.

Re: AI Data Laundering

#15
post #9

Earlier quoted context omitted.

> the revelations do not provide a significant market substitute for the protected aspects of the originals It does seem like generative AI systems provide a significant market substitute, so this ruling probably wouldn’t apply, in court. edit: see https://news.ycombinator.com/item?id=33194623 for some initial thoughts on how this problem (and others) could be rectified. For example, with a database of protected work…

A substitute for what though? Copyright law is only concerned with substituting the work under copyright. That is to say, the consideration is whether the infringing aspects of the secondary work would alter the demand and market for the work being infringed. In all the talk about AI data laundering there really hasn't been any indication that the AI generated item substitutes for the item it's alleged to infringe on…

Stock photography seems to be the obvious instance - why bother paying for the labor to make a stock photo, when you can have a generative AI system create the photo for you?

And furthermore, has anyone demonstrated that it is or is not possible to fully, or substantially, recreate any given existing work using the right input prompts?

I’m interested to know more of the legal details, but my understanding of copyright law is such that it preserves the value of intellectual labor.

edit: on a certain level, the cat is already out of the bag, but that doesn’t mean that we should ignore the law, without some indication from lawmakers or government that they intend to adjust said laws

Re: AI Data Laundering

#16
post #10

Earlier quoted context omitted.

This is larger than publishers, this is every artist, film-maker, photographer, every writer, every engineer, anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. Comment generated with gpt-neox prompt: Comment about AI and data collection and generation and it…

this is larger than the arts. anybody has ever participated creatively in our culture understands that it's absolute bullshit to pretend we need money in order to want to contribute artistically. we need money because food is for sale, because most of us do not own where we live hence we are forced (a priori) to come up with a whole lot of money every month or else you're out in the streets.

Sure but unless you bring down capitalism people will still need to work to eat and most will want to use their hard-earned creative skills to make a living.

Not only that but being able to dedicate 8 to 10 hours a day to your craft for 40 years bring it to a level that you can't reach with casual practice.

Re: AI Data Laundering

#17
post #6
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

Do we allow artists to withhold their works from the minds of eager, learning children? [1] Tell me how ML is different than the mind of a toddler ravenous for new information. For every billion dollar start-up using data at scale, there are tens of thousands more researchers and hobbyists doing the exact same, producing wonderful results and advances. If we stop this growth dead in the tracks, other countries more w…

Most machine learning is assigning weights in a chain of matrix multiplications and normalization functions.

There is no known experimentally verifyiable model of toddlers' brains, let alone one based on matrix multiplication and normalization. Developing such a model would be a noteworthy achievement.

Therefore these are different.

Re: AI Data Laundering

#18
post #8

Earlier quoted context omitted.

> anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. This has been the case ever since people started putting their art on the Internet publicly. The only difference is that now it's algorithms creating the derivatives, not people.

This is not remotely the same, scale and barrier to entry matter. With stable diffusion I can pick any artist right now and create over 1000 derivative works by tomorrow morning in his style to the same degree of expertise with no training involved and no work required.

That's good!

Acting like it's a bad thing is just ludditery.

Re: AI Data Laundering

#19
post #6

Earlier quoted context omitted.

Do we allow artists to withhold their works from the minds of eager, learning children? [1] Tell me how ML is different than the mind of a toddler ravenous for new information. For every billion dollar start-up using data at scale, there are tens of thousands more researchers and hobbyists doing the exact same, producing wonderful results and advances. If we stop this growth dead in the tracks, other countries more w…

Most machine learning is assigning weights in a chain of matrix multiplications and normalization functions. There is no known experimentally verifyiable model of toddlers' brains, let alone one based on matrix multiplication and normalization. Developing such a model would be a noteworthy achievement. Therefore these are different.

Some Artificial Neural Networks have been shown to significantly (at least up to 50% concordance) model brain function.

Not the mention the laborious work of neuroscientists to build out the connectome of the human brain.

Re: AI Data Laundering

#20
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

I would argue (as the court did) that google's use is transformative because the end result "book search" is in a different marketplace from "books." The end result / output of these generative AI systems trained on stock media and art is..."stock media and art."

That's kind of what this whole article is about. Just training the systems in research is arguably fair use but creating the entire pipeline might not be and the "loophole" here is trying to claim no responsibility for the training at the center of it because that was technically done by a 3rd party (...funded by the final creator of the full entire pipeline.)

Post reply on HN