Live data from Hacker News

AI Data Laundering

waxy.org

101–110 of 120 posts

Re: AI Data Laundering

#101
post #88

Earlier quoted context omitted.

Or more to the point, “this picture contains copyrighted material, which must be censored because…” etc. I tried to be as general as possible. The training data for a self-censorship neural network could be as robust as any given society would like. An algorithm based on self-censorship of generated output wouldn’t require censorship within the training data for the generative neural network. I can imagine some other…

But literally (and I use that word literally) none of the pictures contain copyrighted material.

I don't know how people can make these strong statements about anything in law.

Disney have won cases in court were some artist has drawn their own version of Mickey Mouse, similarly try writing a story about some kids in a wizard school and you need to be extremely careful not to violate (or at least get taken to court) for Harry Potters copyright.

I'm pretty certain image production models have produced some images which would very likely to be judged to violate copyright (a much less strong statement).

Re: AI Data Laundering

#102
post #45

Earlier quoted context omitted.

Stock photography seems to be the obvious instance - why bother paying for the labor to make a stock photo, when you can have a generative AI system create the photo for you? And furthermore, has anyone demonstrated that it is or is not possible to fully, or substantially, recreate any given existing work using the right input prompts? I’m interested to know more of the legal details, but my understanding of copyrigh…

This is precisely my point though, "stock photography" isn't an individual work. Copyright law doesn't apply because you can't infringe on the copyright of "stock photography" as a whole any more than you can infringe on the copyright of "rap" or "rock and roll" or "oil paintings". Further, just because a new tech can substitute for a class of old tech doesn't (often, barring protectionist laws) mean the old tech get…

> Lastly if the argument is about that the tech makes it "possible to fully or substantially recreate any given existing work" using deliberate and specific inputs, well we've had plenty of legal precedent on that too. The same arguments were made about Xerox machines, about cassette tapes, about VCRs, about CD-Rs. The copyright holders pretty much lost in every case. At the point you are taking specific and deliberate actions to knowingly infringe on copyright is the point where the technology used is no longer relevant. The right inputs can be used to infringe on the copyright of Star Wars at any typewriter or computer keyboard in the world. The right inputs can be used to infringe on the copyrights of The Beatles at virtually any instrument. It is the act of infringing, not the technology, which is relevant here

There is a significant difference here though, a Xerox machine or a VCR itself does not contain a representation of the art they are copying, a DL network does. I am pretty certain the cases around Xerox/VCRs etc would have had a pretty different outcome if you could type a prompt into your machine "print a story about some kids in a wizard college fighting against the comeback of an evil sorcerer" and it would have put out something closely resembling Harry Potter.

Re: AI Data Laundering

#103

Earlier quoted context omitted.

I'm pretty sure I can count the number of words in Harry Potter without breaking copyright law. It is absolutely not clear when statistical models stops counting ngrams and starts making a derived works.

You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe. You can't call any of your work HP. This sounds obvious, but if it's done by a machine then some people think it's a different question. I can write code to get a list of characters in the book, get their page numbers analysed…

>You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe

IIRC there have been lawsuits about exactly that. A person wrote (and published) some fandom in the Harry Potter universe (without Harry Potter in it IIRC), he lost the case I believe. This is similar to the fact that you cannot make your own comic books with Mickey Mouse (unless your operation is small enough that it flies under the radar), the universe/characters are in fact copyrighted.

Re: AI Data Laundering

#104

Earlier quoted context omitted.

You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe. You can't call any of your work HP. This sounds obvious, but if it's done by a machine then some people think it's a different question. I can write code to get a list of characters in the book, get their page numbers analysed…

>You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe IIRC there have been lawsuits about exactly that. A person wrote (and published) some fandom in the Harry Potter universe (without Harry Potter in it IIRC), he lost the case I believe. This is similar to the fact that you canno…

Probably they used too much reference, I wasn't implying that the universe itself is not protected. But writing something similar that would appeal the fans should be okay.

Re: AI Data Laundering

#105
post #95

Big Tech has really big datasets esp Google. With YouTube, Photos, Music, Gmail, Docs, Maps, Books, Waymo, Search … they have giant multimodal datasets that capture essence of all human knowledge. They have 10+ products with more a billion users creating data for them. If Google Brain/DeepMind were to crack AGI, it would make Google/Alphabet crazy rich at the detriment of millions of YouTubers, Book authors, musician…

Ads companies getting rich off of AGI seems a bit sensational when they’re already getting rich off of the boring type of AI. They’ve already gotten rich indexing the web and all the data we have years ago.

Re: AI Data Laundering

#106
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

> Google’s unauthorized digitizing of copyright-protected works, creation of a search functionality, and display of snippets from those works are non-infringing fair uses. The purpose of the copying is highly transformative, the public display of text is limited, and the revelations do not provide a significant market substitute for the protected aspects of the originals.

So is digitizing a copyright vhs and hosting it via torrents also fair use? Its transformative, the public display of the video is limited, there is no market for vhs.

I don't get it whats the difference other than Google having deeper pockets than me?

Re: AI Data Laundering

#107

I've got a couple examples of Stable Diffusion replicating watermarks along with similar swatches of imagery into scenes from the same prompt [1]. A single case of this should be enough to file a massive lawsuit if the art were recognizable to the creator. [1] https://news.ycombinator.com/item?id=33061707

The model learns all attributes of the images it's trained on, including that some have a watermark. The fact that it generates a watermark in some images doesn't mean that that is a 1:1 image from the training set, it just means to the model some images seem to have a watermark, so it will add it sometimes. Often you can just add "no watermark" (or add it as a negative prompt with some weights) and re-use the same seed to get the same image without the watermark.

Re: AI Data Laundering

#108
post #76

Earlier quoted context omitted.

The Luddites weren't some cult of ignorant technophobes, they were highly-skilled middle class craftsmen and small business owners who went from being able to provide for their families to dying in utter destitution. The remainder of them were tried for machine breaking and were either executed by the state or exiled to penal colonies. They risked everything because everything was at stake, I have a hard time saying…

Certainly, but automation is what allows for improvement to the whole. Today, clothing is cheap and plentiful, along with bedding, curtains, towels and other cloth materials. Clothing would be outrageously expensive if everything were still hand spun, hand loomed, hand cut, hand sewn, and hand screened. If the human computers[1] that predated the rise of the machine computer had done the same and won, it would have c…

>Do we smash the data centers now

Yes, the sooner, the better.

Re: AI Data Laundering

#109
post #6
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

Do we allow artists to withhold their works from the minds of eager, learning children? [1] Tell me how ML is different than the mind of a toddler ravenous for new information. For every billion dollar start-up using data at scale, there are tens of thousands more researchers and hobbyists doing the exact same, producing wonderful results and advances. If we stop this growth dead in the tracks, other countries more w…

>Tell me how ML is different than the mind of a toddler ravenous for new information.

The toddler is human. AIs are not humans.

It's a human right to learn. Non-humans don't (and shouldn't) have human rights.

>Humans aren't the end or the peak of evolution. We should be excited to watch this unfold.

Spoken like a true evolutionary loser.

Re: AI Data Laundering

#110
post #39
post #8

Earlier quoted context omitted.

> anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. This has been the case ever since people started putting their art on the Internet publicly. The only difference is that now it's algorithms creating the derivatives, not people.

Yeah before the internet it never happened and nobody knew just how damn cliched Bill Shakespeare's plays are. Every line of Hamlet's soliloquy! It's insane!

Would we have Shakespeare's plays if he didn't make money? Which encourages better plays:

I write a play, and I can license theater companies to be able to perform it. Therefor better writers are attracted to the industry (instead of to say Ad Copywriting) and because of a higher level product, the industry thrives.

I can write a passion play that the local theater will perform. I will not generate enough income to live from my product. I will not generate income from licensing my production because there is no copyright and my scripts would just get stolen/distributed freely. The industry has less quality productions. The majority of productions have no reputation of quality.

Post reply on HN