Live data from Hacker News

AI Data Laundering

waxy.org

91–100 of 120 posts

Re: AI Data Laundering

#91
post #4
post #3

The Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "proj…

> I think allowing people to exclude themselves or their work from a dataset is necessary. or they could open it all up for everybody and stop protecting the rights of death people (authors dead less then 70 years ago) then again, that will make the publishers starve... but why pretend publishing corporations need food?

Why pretend that other corporations that vacuum up content and repackage it have rights to resell art that you want to strip from the original publishers? At least the publishers actually made a contract with the artists.

Re: AI Data Laundering

#92
post #63

Earlier quoted context omitted.

The court’s summary also mentions this aspect of differing marketplaces: “… the revelations [i.e. the information served by Google Book Search] do not provide a significant market substitute for the protected aspects of the originals.” This doesn’t apply to AI image generators which are clearly a “market substitute” for the protected originals used to train the system. For this reason I’d expect someone like Getty to…

Can we first get an AI that's actually usefull as a Getty substitute? All I'm seeing posted is visually pleasing nonsense - as soon as I tried to use it for stuff like stock photo generator it's unusable (eg. key physical properties of the object would be off to the point where the object is useless, and in many cases it would look wrong even from a thumbnail). The only thing I did see was designers cropping out the…

Did you try inpainting to fix the wrong bits? From my very little experience, AI image generation is not (yet) a one click process and requires multiple iterations to get close to the desired result, i.e. it is still more a tool for a designer than a replacement for it.

Re: AI Data Laundering

#93

Earlier quoted context omitted.

Most machine learning is assigning weights in a chain of matrix multiplications and normalization functions. There is no known experimentally verifyiable model of toddlers' brains, let alone one based on matrix multiplication and normalization. Developing such a model would be a noteworthy achievement. Therefore these are different.

Some Artificial Neural Networks have been shown to significantly (at least up to 50% concordance) model brain function. Not the mention the laborious work of neuroscientists to build out the connectome of the human brain.

These articles use far more cautious language than you suggest and if they don't everyone working in the field is hopefully aware that such claims are the academic equivalent of clickbait at best.

Re: AI Data Laundering

#94
post #40

Earlier quoted context omitted.

> Tell me how ML is different than the mind of a toddler ravenous for new information. If a person published a work that clearly plagiarized or violated a patent, that person would be open to legal action. I’m all for systemic change, but uses like this may end up having a chilling effect on human-created work.

> I’m all for systemic change, but uses like this may end up having a chilling effect on human-created work. Everytime this comes up, whichever party fears for it's livelihood always says something like this and ignores the other side: that rigorous enforcement activity is going to do the same thing, to human created work. Richard Stallman wrote a short story about this very issue.[1] There are already people hurling…

Funny that you cite Stallman, when Copilot using GPLed code in closed-source projects is a real concern.

Re: AI Data Laundering

#95
Big Tech has really big datasets esp Google. With YouTube, Photos, Music, Gmail, Docs, Maps, Books, Waymo, Search … they have giant multimodal datasets that capture essence of all human knowledge. They have 10+ products with more a billion users creating data for them.

If Google Brain/DeepMind were to crack AGI, it would make Google/Alphabet crazy rich at the detriment of millions of YouTubers, Book authors, musicians, drivers.

AI will concentrate power and wealth to fewer individuals.

Re: AI Data Laundering

#96

> It’s currently unclear if training deep learning models on copyrighted material is a form of infringement What? It's clearly a derived work.

I'm pretty sure I can count the number of words in Harry Potter without breaking copyright law. It is absolutely not clear when statistical models stops counting ngrams and starts making a derived works.

You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe. You can't call any of your work HP. This sounds obvious, but if it's done by a machine then some people think it's a different question.

I can write code to get a list of characters in the book, get their page numbers analysed and draw graphs to help me create my own version. Am I breaking copyright laws? Most likely not.

It's a truly grey area which lawmakers never saw coming.

I believe if events unfold well we'll see and treat AI tools to be like sharp knives eventually. It will be up to the user what they do with it.

Re: AI Data Laundering

#97
post #87
post #74

Earlier quoted context omitted.

>No longer do we need to pay 20,000 hours to learn one thing to the exclusion of all others we would like to try. Now we'll be able to clearly articulate ourselves with art, music, poetry. We'll become powerful beings of thought and expression. I'm a 20000 hours person. Knowing what I know about what I do, it's real sad to see someone misunderstand what goes into creativity this egregiously. Prompt engineering is suc…

You can prompt with images, which let's you control colour and composition, and with masking you can iteratively work on sections to guide the image to what you are picturing. That can shift the creative part more towards the user.

Yes, I've seen the photoshop plugin. You're comparing playing with duplo blocks to marble sculpture.

Re: AI Data Laundering

#98
post #63

Earlier quoted context omitted.

I would argue (as the court did) that google's use is transformative because the end result "book search" is in a different marketplace from "books." The end result / output of these generative AI systems trained on stock media and art is..."stock media and art." That's kind of what this whole article is about. Just training the systems in research is arguably fair use but creating the entire pipeline might not be an…

The court’s summary also mentions this aspect of differing marketplaces: “… the revelations [i.e. the information served by Google Book Search] do not provide a significant market substitute for the protected aspects of the originals.” This doesn’t apply to AI image generators which are clearly a “market substitute” for the protected originals used to train the system. For this reason I’d expect someone like Getty to…

Note the "protected aspects of the originals" part. AI generated images don't produce outputs that contain protected aspects.

Re: AI Data Laundering

#99
Not sure laundering it the right term.

Laundering private things through the commons feels not as shady as laundering in private networks. The commons benefits too.

It's more like open source that money laundering

Re: AI Data Laundering

#100
Are we heading towards voiding most of current copyrights or is there a way out of this mess with another patch to the laws?
Post reply on HN