Live data from Hacker News

Things are about to get worse for generative AI

garymarcus.substack.com

311–320 of 769 posts

Re: Things are about to get worse for generative AI

#311
I am beginning to think that in these discussions these models are functioning more like an obscuring factor than anything else and the discussion is getting bogged down in that, and not the crux of the argument.

They’re giving people plausible deniability in the “chain of responsibility”, and I think if we took away “LLM” and replaced it with “fairground sideshow magic box” the argument that LLM’s are somehow special and deserving of exemptions disappears real quick.

Re: Things are about to get worse for generative AI

#313
How about this: Image generators should be treated like random google image search. They sample randomly from the distribution of publicly viewable images. Google does it exactly while Image generators do it in an interpolative way. Google images produced copyrighted works most of the time, an image generator only sometimes. Neither should be liable if someone sells a copyrighted work that was produced to someone else.

Re: Things are about to get worse for generative AI

#314

Am I the only one believing that copyright has long outlived its usefulness? After all, copyright is not some natural law or mathematical consequence, but rather a social convention that made sense in the era of the printing press.

As an author, I do want the stories I write and worlds I build to be protected for a reasonable period.

Right now, copyright is a significant discouragement to any other entities from taking a story I wrote and claiming it as their own and preventing me from ever growing an audience for my work. It’s far from perfect, and I can’t afford litigation, but it enshrines a cultural value of allowing people to create things and be known for them. Profit is a side effect of this.

Art is already poorly valued compared to the enormous investment time and energy required to produce it. Removing copyright means you can’t even have minimal protections from a more popular person erasing you.

Re: Things are about to get worse for generative AI

#315

Earlier quoted context omitted.

> a few copyright holders By which you mean every copyright holder. > AGI in the near future Something that is purely speculative, undefined, and has been promised in the near future for 50+ years. I don't see copyright holders lying down for someone else's benefit and I don't see governments gutting copyright, contract law, and several other avenues of protection that copyright holders can deploy in the name of some…

The examples given are all billion-dollar, decades old characters. The volume of material directly/indirectly referencing those characters in a random internet crawl will be fairly large. Most copyrighted works won't have that issue. If anything it means they only infringe on archetypal works and not the other 99.9%. If I write a story involving robots and spaceships (of which there are many, before and since Star Wa…

The examples were chosen by the author to make a point precisely because they are well known.

But every single copyright holder with their works online (which includes you and me) has the same legal rights as the NYT or Disney. Naturally some copyright holders have more real-world capability to go legal than others, but that does not reduce the legal risk.

> If anything it means they only infringe on archetypal works and not the other 99.9%

How on earth do you get to that conclusion? There's no "popularity" floor to copyright protection. Either a work has been infringed or it hasn't.

Re: Things are about to get worse for generative AI

#316
post #313

How about this: Image generators should be treated like random google image search. They sample randomly from the distribution of publicly viewable images. Google does it exactly while Image generators do it in an interpolative way. Google images produced copyrighted works most of the time, an image generator only sometimes. Neither should be liable if someone sells a copyrighted work that was produced to someone els…

But when Google image search produces a result, the question of whether it is copyrighted is something I can generally figure out in a matter of seconds or minutes. This is not so for image generators.

Re: Things are about to get worse for generative AI

#317
post #124

These don't seem all that difficult to fix to me. Most of the examples are not really generic, but are shorthand descriptions of well-known entities. "Video game plumber" is practically synonymous with "Mario" and anyone that has the slightest familiarity with the character knows this. Likewise, how difficult is it to just use descriptive tools to describe Mario-like images [1] and then remove these results from anyo…

It seems like a somewhat dystopian thing to fix. Imagine a scenario where Photoshop would scan images you uploaded for copyright material and then refuse to work if it determined image contained any copyrighted material or characters (even if it was just a fan drawing you did). This reminds me of the early days of the internet where people wanted to remove free fanfiction for violations of copyright laws. Trying to a…

image editors don't offer something that's based on questionably sourced copyrighted material as a part of their product. ai apps and services do.

it's just ai companies using dirty data and hoping they get away with it - and they do, for the time being, it is a bit trickier to show that 'yep, well that's there', and people don't seem to realize that just using a copyrighted image, at all (downloading, accessing in itself, let alone using for something else), or creating an image that would just "look like" a trademarked character - not "make a 1 to 1 copy" but just "look like" - would be enough for it to possibly be an infringement.

there can be a sufficient fix - taking out potentially infringing images from a dataset, and making an effort to make an actually clean dataset. it's really just a matter of "do you actually have rights to use that content? at all, and in that way". and ai companies continually say 'no...but what if we use it anyway".

and it's kind of a sloppy analogy, because with text to image generators (where you just interact with a model that's offered to you), well - people aren't "uploading copyrighted material into an editor". the copyrighted material is already there in the model, it was used in making of it. and if there was no such copyrighted material that'd fit the prompt, it wouldn't be able to generate something. the infringement lies with the service that uses copyrighted material for a model, and then offers it.

fan fiction and fan works continuously being in a murky area with copyright/trademark is not just a thing of "early days" of internet, it's been there all along and is still very much present. companies could crack down if they wanted, but there is too much of stuff out there, it might be hard to nail down exact people, and it might be plainly not too nice to the fandom. but it is not "impossible", and it is very much not a conversation that ever 'went away' or become "kinda solved" - it isn't.

again, with image editors, text editors, etc. - user is making all the actions with content, and the user would be doing the infringement, in editing and further if they were to choose to publish.

with generative ai - copyright infringement is built into the models. copyrighted works were accessed and used to turn into a model. user is just asking, "is it there". and it is. in some of those demonstrated examples, user is not even asking for a model to infringe on anything but it just does.

Re: Things are about to get worse for generative AI

#318

As I understood it, the legal precedent for generative AI is the same one that allows google to scrape websites in order to index them for search for the common good. Google also can display cached versions of websites which is the original content of those sites. No one is going to say that google is copyright infringement just because it is showing content from other websites verbatim. So I think this is a weak arg…

Wonder. Do Cliff Notes have to pay royalties to the underlying material? Cliff Notes contain quotes, and citations. Does the cliff note company, when producing Cliff Notes for "Into The Wild", pay royalties to the publisher? For that matter, does any paper, article, etc.. that may contain a quote from another, have to pay royalties to the source of the quotes?

Cliff’s Notes has a strong fair use claim, because they offer basic criticism and surface-level commentary alongside their summaries.

Re: Things are about to get worse for generative AI

#319
Wow. I feel really sorry for these giant corporations who have wielded armies of lawyers against fanfic artists to prevent fair use, and to prevent trademarks and patents from expiring on the timelines enshrined by law.

Can we all have a moment of silence for poor Bob Iger? Maybe we can start a GoFundMe to help him out?

Re: Things are about to get worse for generative AI

#320

Should not be a problem in the EU. Article 3 and 4 of the „ Copyright in the Digital Single Market“ Directive already regulate this. Summary by Wolters Kluwer: […] Everyone else (including commercial ML developers) can only use works that are lawfully accessible and where the rightholders have not explicitly reserved use for text and data mining purposes. AFAIK they are discussing something like a robot.txt to flag s…

The EU cannot agree that the Do Not Track flag on web browsers is legally binding but big content should be able to create legally binding flags on their websites to avoid scraping of data? Seems odd!

I don't think that's a fair analogy. One forces 99% of websites to make a change, while the other is something that would need to be done by the big companies doing the scraping.

A Do Not Track flag being legally binding would force small websites, e.g. a local restaurant website, to implement something they likely are not aware of and secondly do not technically understand.

A company that is mass scraping data for their AI model is much more likely to understand and respect that scraping the data has legal implications, and would be technically capable in implementing a scraping solutions that accounts for a robots.txt.

Post reply on HN