Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

461–470 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#461

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I hope it happens. I'd love to see a market for selling training licenses to IP. This could be a small but real source of passive income for artists, authors, poets who don't mind that their IP being used in training sets. I wouldn't be practical to negotiate individually with each artist but I could see something work with larger collectives than can vouch for the quality of their members. Think like publishers, galleries, guilds or unions. It could offer a license and then share the proceeds with all members.

It's just flat out unethnical for LLMs to just soak up all this data, even off torrent sites(!!!), without any consent or agreement with the IP holders. Some model like this could be a win for everyone

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#462

Earlier quoted context omitted.

It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books. Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge…

It’s important to note that only some copyrighted works can be accessed for free using the legal options, I’m a member of probably a dozen overdrive supporting libraries and still frequently find titles unavailable for loan of any kind. I’d love to see an analysis of what % of books are available via libraries around the globe. Also, the whole DRM thing is a massive pain, audiobooks especially are terrible at allowin…

>frequently find titles unavailable for loan of any kind

Even through interlibrary loan?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#463

Earlier quoted context omitted.

He hacked into a server to release a database of paywalled studies to the public. Not only is it not the same but it was the hacking that brought charges upon him.

It's been quite a few years, but AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR. Maybe he violated their terms, but that is not illegal. He did illegally trespass an unlocked MIT switch closet to do this. They blocked several IPs but his script would rotate to continue. The downloading was over a week or two, enough for security to set up a camera in…

This just all depends on what circle you're using the word "hacking" in. While technical circles are mostly concerned about the technical details on how someone might exploit a system, legal circles don't really care. They usually avoid using the word "hacking" anyway.

The relevant law here, the CFAA, is often referred to as the US law that criminalizes "hacking", but what it specifically does is criminalize anyone who "intentionally accesses a computer without authorization or exceeds authorized access" which is much more broad than how technical disciplines might use the word.

So yes, stealing a password off a friend's post-it note and Hasselhoffing their instagram might not be considered "hacking" if you're hanging out at Defcon, this would be considered "hacking" in legal or colloquial terms.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#464

Earlier quoted context omitted.

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

Books and libraries have existed for thousands of years. It’s copyright that is the young intruder. Most people who write non-fiction books do it because they want to contribute to human knowledge and be recognized as an expert in a particular field, not because they think that writing a differential geometry textbook is their path to riches. With the internet, more and more text books are made freely available by th…

Since you mentioned textbooks on differential geometry, I can recommend anyone interested Sigmundur Gudmundsson's lecture notes in introductory differential geometry as well as Riemannian geometry. You can find them freely available on his personal academic page.

https://www.matematik.lu.se/matematiklu/personal/sigma/

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#465
post #392

Earlier quoted context omitted.

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

> The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How do you know that this benefit wouldn't exist in other schemes? Look at permissive open source software which is essentially public domain + shield from liability. No copyright does not mean no compensation. It just means different compensation that doesn't deprave other people of their right to…

> Look at permissive open source software ... No copyright does not mean no compensation

From everything I've heard it kinda does. If you're writing something valuable then maybe a company will employ you to keep working on it, and the portfolio can certainly help in interviews (to write other software), but getting non-negligible compensation for the use of the software itself is rare. Even those projects that are well funded, like the Linux kernel, are done so not out of goodness of heart, but due to companies realising it's in their rational interest to have a common standard base of sorts

In terms of written works, the best comparison we have is Wikipedia, and while the foundation does receive funding from companies who realise how useful of an integration it can be for their products, the writers themselves do not get paid afaik (and when they do, it's rarely a good thing). But if you just wrote an open source textbook, I doubt you'll manage to make much money off it, and the prospects for fiction look even worse

> Perhaps kickstarter-style firms that direct oversight over funded projects

Sounds like a return to rich patrons and needing to flatter them to get grants. Luckily with the modern internet we do now have democratised layman patronage, but why the need to force everyone into that model? Also note that almost all successful Patreon artists do have perks for paying, even if it's just early access, and afaik make a lot of their money off commissions. Those who just post art for free, with no paid comms, and just have a "tip jar" make relatively little from what I hear

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#466
post #150

Earlier quoted context omitted.

That's because, when torrenting, you're typically also seeding a copy of it, i.e. you're distributing your local copy to other devices, and thus you're directly aiding in piracy. Simply downloading content from a centralized server, as explained above, is different. Although, one could argue what OpenAI & Meta are doing is closer to the torrent definition than the "simply downloading" definition, given that they're u…

Honestly don't think our current laws are even good for this case. This clearly needs some sort of regulation or policy. It's clearly pretty bullshit if you ask chatgpt for a joke and it repeats a Sarah Silverman joke to you, while they charge you a subscription for it and she gets none of that sub money.

there is policy in most places, and the policy is fuck you pay me.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#467

Earlier quoted context omitted.

There's a difference between "information wants to be free" and "Facebook can produce works minimally derived from your greatest creative work at a scale you can't match". LLMs seem to aggregate that value to whoever builds the model, which they can then sell access to, or sell the output it produces. Five years from now, will OpenAI actually be open, or will it be a rent seeking org chasing the next quarterly gains?…

"Will OpenAI actually be open" That ship sailed, friend. OpenAI is no longer a charity in any meaningful sense of the word anymore, it's now an adversarial organization working against the public good with the sole aim of making a few rich men richer. After privatization, they sent their PR people to lobby congress to make it impossible for anyone to compete with them (important note: not out of any interest in actua…

we are rapidly achieving the cyberpunk future, and it's much worse than we thought.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#468

Earlier quoted context omitted.

I've followed the issue in the US since the early 2000s as an activist and policy expert. I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption. I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and…

Are you forgetting all the DMCA lawsuits slapping individuals who downloaded MP3s with tens of thousands of dollars? These were not corporations, these were teenagers still living with their parents who pulled music files off the likes of Napster. The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#469

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

> copyright is the reason ~every child doesn't have access to ~every book ever written. And? Is there some reason anybody, child or adult, deserves access to "every" anything? Should children have access to every video game ever made, every Matchbox car, every Lego set?

Yes? I can't see any reason "you can't have that Lego set because the giant corporation behind it stopped selling it, refuses to let anyone else make it, and resale is all horribly expensive from collectors" is something we should have to tell our kids.

I really hope my children will have access to a higher tech version of something like the (discontinued?) toy where you could melt old crayons into toy car bodies. 3D printing is almost there, but ideally the end product of a Lego set could be recycled into new blocks easily and quickly.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#470
post #307

Earlier quoted context omitted.

Then most people stop writing books because they can't get paid for their time/effort and ~every child will be stuck with outdated knowledge within a decade.

so why haven't they already stopped if it is already trivial to download nearly any new piece of media?

Because it's not. It is trivial to buy a book on the Kindle store, and while it may be trivial to us here to go and pirate a azw3 and transfer it, it's not to most people

People always forget the pareto principle when it comes to anti-piracy. No, they don't stop everyone, but a minor hurdle stops a hell of a lot of "ordinary" people

Post reply on HN