Live data from Hacker News

Be a property owner and not a renter on the internet

den.dev

201–210 of 270 posts

Re: Be a property owner and not a renter on the internet

#201
post #200

Earlier quoted context omitted.

Your shop might be open sure - but aren't we talking about people coming in and taking whatever they like for free? ie if you were an art gallery, the expectation would be people could come in and look, but you don't expect them to come in, photograph everything and then sell prints of everything online.

That's not what's happening. Instead, it's that there's some people coming into your gallery, studying the art and its style, and leaving with the learned information. They then replicate that style in their own gallery. Of course, none of the images are copies, or would be judged to be copies by a reasonable person. So now you, the gallery owner, want to forbid just those people who would come to learn the style. Bu…

> Of course, none of the images are copies, or would be judged to be copies by a reasonable person.

That's the fiction of course.

Tell me how something like ChatGPT can simultaneously claim to return accurate information while at the same time being completely independent from the sources of the information?

In terms of images - copyright isn't only for exact copies - it if was then humans would have been taking the piss by making minor changes for decades.

Sure you could argue some is fair use with genuinely original content being produced in the process, but I think you are also overlooking an important part of what's considered 'fair' - industrialised copying of source material isn't really the same in terms of fairness as one person getting inspiration.

Taking the Encylopedia Britanica and running it though an algorithm to change the wording, but not the meaning, and selling it on is really not the same as a student reading it and including those facts in their essay - the latter is considered fair use, the former is taking the piss.

Re: Be a property owner and not a renter on the internet

#202

Earlier quoted context omitted.

> But the new AI training methods are currently, at least imho, not a violation of copyright - not any more than a human eye viewing it Interesting comparison - as if a human viewed something, memorized it and reproduced in a recognisable way to be pretty much the same, wouldn't that still breach copyright? ie in the human case it doesn't matter whether it went through an intermediate neural encoding - what matters i…

The difference is an image generation algorithm does not consume images the way a human does, nor reproduce them that way. If you show a human several Rembrandt's and ask them to duplicate them, you won't get exact copies, no matter how brilliant the human is: the human doesn't know how Rembrandt painted, and especially if you don't permit them to keep references, you won't get the exact painting: you'll get the elem…

> with the right prompts.

that is doing a lot of pull. Just because you could "get the full copies" with the right prompts, doesn't mean the weights and the training is copyright infringement.

I could also get a full copy of any works out of the digits of pi.

The point i would like to emphasize is that the using data to train the model is not copyright infringement in and of itself. If you use the resulting model to output a copy of an existing work, then this act constitutes copyright infringement - in the exact same way that using photoshop to reproduce some works is.

What a lot of anti-ai arguments are trying to achieve is to make the act of training and model making the infringing act, and the claim is that the data is being copied while training is happening.

Re: Be a property owner and not a renter on the internet

#203

Earlier quoted context omitted.

> But the new AI training methods are currently, at least imho, not a violation of copyright - not any more than a human eye viewing it Interesting comparison - as if a human viewed something, memorized it and reproduced in a recognisable way to be pretty much the same, wouldn't that still breach copyright? ie in the human case it doesn't matter whether it went through an intermediate neural encoding - what matters i…

The difference is an image generation algorithm does not consume images the way a human does, nor reproduce them that way. If you show a human several Rembrandt's and ask them to duplicate them, you won't get exact copies, no matter how brilliant the human is: the human doesn't know how Rembrandt painted, and especially if you don't permit them to keep references, you won't get the exact painting: you'll get the elem…

We have the problem of too-perfect-recall with humans too -- even beyond artists with (near) photographic memory, there's the more common case of things like reverse-engineering.

At times, developers on projects like WINE and ReactOS use "clean-room" reverse-engineering policies [0], where -- if Developer A reads a decompiled version of an undocumented routine in a Windows DLL (in order to figure out what it does), then they are now "contaminated" and not eligible to write the open-source replacement for this DLL, because we cannot trust them to not copy it verbatim (or enough to violate copyright).

So we need to introduce a barrier of safety, where Developer A then writes a plaintext translation of the code, describing and documenting its functionality in complete detail. They are then free to pass this to someone else (Developer B) who is now free to implement an open-source replacement for that function -- unburdened by any fear of copyright violation or contamination.

So your comment has me pondering -- what would the equivalent look like (mathematically) inside of an LLM? Is there a way to do clean-room reverse-engineering of images, text, videos, etc? Obviously one couldn't use clean-room training for _everything_ -- there must be a shared context of language at some point between the two Developers. But you have me wondering... could one build a system to train an LLM from copywritten content in a way that doesn't violate copyright?

[0]: https://en.wikipedia.org/wiki/Clean-room_design

Re: Be a property owner and not a renter on the internet

#204
post #44

> Exploiting user-generated content. You know, if I've noticed anything in the past couple years, it's that even if you self-host your own site, it's still going to get hoovered up and used/exploited by things like AI training bots. I think between everyone's code getting trained on, even if it's AGPLv3 or something similarly restrictive, and generally everything public on the internet getting "trained" and "transfor…

I think you're right, and I don't think it's just about public content being "exploited" to train AI models and the like. Rather, even before LLMs, there was a growing sense that publishing ideas or essays publicly is "risky" with very little reward for the very real risks.

I wrote about this a little in "The Blog Chill":

https://amontalenti.com/2023/12/28/the-blog-chill

Speaking personally, among my social circle of "normie" college-educated millennials working in fields like finance, sales, hospitality, retail, IT, medicine, civil engineering, and law -- I am one of the few who runs a semi-active personal site. Thinking about it for a moment, out of a group of 50-or-so people like this, spread across several US states, I might be the only one who has a public essay archive or blog. Yet among this same group you'll find Instagram posters, TikTok'ers, and prolific DM authors in more private spaces like WhatsApp and Signal groups. A handful of them have admitted to being lurkers on Reddit or Twitter/X, but not one is a poster.

It isn't just due to a lack of technical ability, although that's a (minor) contributing factor. If that were all, they'd all be publishing to Substack, but they're not. It's that engaging with "the public" via writing is seen as an exhausting proposition at odds with everyday middle class life.

Why? My guesses: a) smartphones aren't designed for writing and editing, hardware-wise; b) long-form writing/editing is hard and most people aren't built for it; c) the dynamics of modern internet aggregation and agglomeration makes it hard to find independent sites/publishers anyway; and d) the risk of your developed view on anything being "out there" (whether professional risk or friendship risk) seems higher than any sort of potential reward.

On the bright side, for people who fancy themselves public intellectuals or public writers, hosting your own censorship-resistant publishing infrastructure has never been easier or cheaper. And for amateur writers like me, I can take advantage of the same.

But I think everyday internet users are falling into a lull of treating the modern internet as little more than a source of short-form video entertainment, streams for music/podcasts, and a personal assistant for the sundries of daily life. Aside from placating boredom, they just use their smartphones to make appointment reminders, send texts to a partner/spouse, place e-commerce orders, and check off family todo lists, etc. I expect LLMs will make this worse as a younger generation may view long-form writing not as a form of expression but instead as a chore to automate away.

Re: Be a property owner and not a renter on the internet

#205
post #200

Earlier quoted context omitted.

That's not what's happening. Instead, it's that there's some people coming into your gallery, studying the art and its style, and leaving with the learned information. They then replicate that style in their own gallery. Of course, none of the images are copies, or would be judged to be copies by a reasonable person. So now you, the gallery owner, want to forbid just those people who would come to learn the style. Bu…

> Of course, none of the images are copies, or would be judged to be copies by a reasonable person. That's the fiction of course. Tell me how something like ChatGPT can simultaneously claim to return accurate information while at the same time being completely independent from the sources of the information? In terms of images - copyright isn't only for exact copies - it if was then humans would have been taking the…

> ChatGPT can simultaneously claim to return accurate information while at the same time being completely independent from the sources of the information?

why can't that be true? Information is not copyrightable. The expression of information is. If chatGPT extracted information from a source works, and represent that information back to you in a form that is not a copy of the original works, then this is completely fine to me. An example would be a recipe.

Re: Be a property owner and not a renter on the internet

#206
post #86

Earlier quoted context omitted.

I think the source of the contrary sentiment goes something like this: AI stuff (especially image generation) is competition for artists. They don't much like competition that can easily undercut them on price, so they want to veto it somehow and lean on their go-to of accusing anybody who competes with them of theft. The problem in this case is that it doesn't matter. The AI stuff is going to exist, and compete with…

I think that's an overly reductive way of looking at it. Artists, are by their definition, creators of art. AI-generated "art" (it's not art at all in my eyes) is effectively a machine-based reproduction of actual art, but doesn't take the same skill level, time, and passion for the craft for a user to be able to generate an output, and certainly generates large profits for those that created the models. So, imagine…

> It's less about competition and more about the ethical way to do it. If another artist would learn the same techniques and then managed to produce similar art, do you think there would be just as visceral of a reaction to them publishing their art? Likely not, because it still required skill to achieve what they did.

Now suppose that the other artist studies to learn the techniques -- several of them do -- and then Adobe offers them each two cents and a french fry to train a model on it, which many accept because the alternative is that the model exists anyway and they don't even get the french fry. Is this more ethical somehow? Even if you declined the pittance, you still have to compete with the model. Even if you accept it, it's only a pittance, and you still have to compete with the model. It hasn't improved your situation whatsoever.

> My hunch is that in the near-term we'll see a major devaluing of both written and image material, while a premium will be put on exceptional human skill.

AI slop is in the nature of "80% as good for 20% of the price" except that it's more like 40% as good for 0.0001% of the price. What that's going to do is put any artists below the 40th percentile out of work, make it a lot harder for the ones at the 60th percentile and hardly affect the ones at the 99th percentile at all.

But the other thing it's going to do is cause there to be more "art". A lot of the sites with AI-generated images on them haven't replaced a paid artist, they've replaced a site without images on it. Which isn't necessarily a bad thing.

Re: Be a property owner and not a renter on the internet

#207
post #205

Earlier quoted context omitted.

> Of course, none of the images are copies, or would be judged to be copies by a reasonable person. That's the fiction of course. Tell me how something like ChatGPT can simultaneously claim to return accurate information while at the same time being completely independent from the sources of the information? In terms of images - copyright isn't only for exact copies - it if was then humans would have been taking the…

> ChatGPT can simultaneously claim to return accurate information while at the same time being completely independent from the sources of the information? why can't that be true? Information is not copyrightable. The expression of information is. If chatGPT extracted information from a source works, and represent that information back to you in a form that is not a copy of the original works, then this is completely…

So you think taking something like the Encylopedia Britanica, running it through a simple rewording algorithm, and selling it on is totally 'fair use'?

Taking all newspaper and proper journalistic output and rewording it automatically and selling it on is also 'fair use'?

Stand back from the detail ( of whether this pixel or word is the same or not ) and look at the bigger picture. You still telling me that's all fine and dandy?

I think it's obviously not 'fair use'.

It means the people doing the actual hard graft of gathering the news, or writing Encylopedias or Textbooks won't be able to make a living so these important activities will cease.

This is exactly the scenario copyright etc exists to stop.

Re: Be a property owner and not a renter on the internet

#208
post #139

Earlier quoted context omitted.

> If the holder of the rights does not agree to a certain kind of use, what else is there to discuss? the holder of content does not automatically get to prescribe how i would use said content, as long as i comply with the copyrights. The holder does not get to dictate anything beyond that - for example, i can learn from the content. Or i can berate it. Copyright is not a right that covers every single conceivable us…

Copyright means the holder does automatically get to prescribe how content can be copied. That's literally the definition of copyright. A typical copyright notice for a book says something like (to paraphrase...) "not to be stored, transmitted, or used by or on any electronic device without explicit permission." That clearly includes use for training, because you can't train without making a copy, even if the copy is…

> Any argument about this is trying to redefine copyright as the right to extract the semantic or cultural value of a document. In reality the definition is already clear - no copying of a document by any means for any purpose without explicit permission.

I've studied copyright for over 20 years as an amateur, and I used to very much think this way.

And then I started reading court decisions about copyright, and suddenly it became extremely clear that it's a very nuanced discussion about whether or not the document can be copied without explicit permission. There are tons of cases where it's perfectly permissible, even if the copyright holder demands that you request permission.

I've covered this in other posts on Hacker News, but it is still my belief that we will ultimately find AI training to be fair use because it does not materially impact the market for the original work. Perhaps someone could bring a case that makes the case that it does, but courts have yet to see a claim that asserts this in a convincing way based on my reading of the cases over the past couple of years.

Re: Be a property owner and not a renter on the internet

#209
post #178
post #115

Earlier quoted context omitted.

There's two decades worth of countless conversations on Reddit alone that would be buried into nothingness but instead ML has revived all that activity as useful data. ML is definitely a great way to bring back utility for a lot of old and unused data.

> revived that activity as useful data Revived as compressed text associations, it is potentially useful data, but also potentially totally wrong in non-obvious ways. (Or, to riff on Futurama, "The worst kind of incorrect.")

It is used to help train the LLMs on how to "talk" like normal people, even if the topic they're discussing isn't that useful or valuable.

Re: Be a property owner and not a renter on the internet

#210
post #202

Earlier quoted context omitted.

The difference is an image generation algorithm does not consume images the way a human does, nor reproduce them that way. If you show a human several Rembrandt's and ask them to duplicate them, you won't get exact copies, no matter how brilliant the human is: the human doesn't know how Rembrandt painted, and especially if you don't permit them to keep references, you won't get the exact painting: you'll get the elem…

> with the right prompts. that is doing a lot of pull. Just because you could "get the full copies" with the right prompts, doesn't mean the weights and the training is copyright infringement. I could also get a full copy of any works out of the digits of pi. The point i would like to emphasize is that the using data to train the model is not copyright infringement in and of itself. If you use the resulting model to…

>The point i would like to emphasize is that the using data to train the model is not copyright infringement in and of itself.

Interesting point - though the law can be strange in some cases - so for example in the UK in court cases where people are effectively being charged for looking at illegal images, the actual crime can be 'making illegal images' - simply because a precedence has been set that because any OS/Browser has to 'copy' the data of any image in order someone to be able to view it - any defendent has been deemed to copied it.

Here's an example. https://www.bbc.com/news/articles/cgm7dvv128ro

So to ingest something your training model ( view ) you have by definition have had to have copied it to your computer.

Post reply on HN