Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

81–90 of 102 posts

Re: Pulling my site from Google over AI training

#81
post #26

Earlier quoted context omitted.

So if someone innocently reaches your site from google they will see a bunch of LLM generated misinformation? You have a legal right to do this (assuming the LLM bullshit isn't libel) but I don't see how it could be considered a moral act.

> So if someone innocently reaches your site from google they will see a bunch of LLM generated misinformation? Yes, although I've been blocking googlebot for years, so nobody will get to my sites through google anyway. > I don't see how it could be considered a moral act. I'm curious about this -- why do you think this is in any way an immoral act? If a naïve human comes across the site, they'll quickly realize that…

>If a naïve human comes across the site, they'll quickly realize that it's not useful and move on

I do not share your confidence.

>Would your moral objections be eased if the first line on the page is something like "this page is full of machine-generated nonsense. Please ignore it"?

That would certainly help.

Re: Pulling my site from Google over AI training

#82
post #77

I can't understand the outrage. In practice absolutely nothing has changed. It is reading and learning. A person would read and learn. This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work. This is no different. I can write some code and use it, subconsciously referencing a work. If I don…

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement. Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof. I c…

So what are these "consequences" you keep referring to?

Re: Pulling my site from Google over AI training

#83
post #77

Earlier quoted context omitted.

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement. Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof. I c…

That's completely subjective. What you describe is a spectrum. If that were true, thesis' would not have to be run through plagiarism scans (they are), and there would be no copyright lawsuits for familiarity such as Ed Sheeran's Marvin Gaye lawsuit. It's also important to note these laws were always intended to strike a fair balance between the copyright owner and the good of society as a whole. Copyright is not an…

This is such an important point that I think escapes a lot of people, both on the pro-AI and anti-AI side.

Although everyone probably wishes otherwise, there isn't really a hard objective line for determining if something is infringing or not. And that's actually a good thing! But it means that, in the end, it's up to a judge to weigh a bunch of different, rather fuzzy, factors.

Re: Pulling my site from Google over AI training

#84
post #26

Earlier quoted context omitted.

> So if someone innocently reaches your site from google they will see a bunch of LLM generated misinformation? Yes, although I've been blocking googlebot for years, so nobody will get to my sites through google anyway. > I don't see how it could be considered a moral act. I'm curious about this -- why do you think this is in any way an immoral act? If a naïve human comes across the site, they'll quickly realize that…

>If a naïve human comes across the site, they'll quickly realize that it's not useful and move on I do not share your confidence. >Would your moral objections be eased if the first line on the page is something like "this page is full of machine-generated nonsense. Please ignore it"? That would certainly help.

> That would certainly help.

That seems reasonable enough. If I do this, I'll include such a disclaimer.

Re: Pulling my site from Google over AI training

#85
post #77

I can't understand the outrage. In practice absolutely nothing has changed. It is reading and learning. A person would read and learn. This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work. This is no different. I can write some code and use it, subconsciously referencing a work. If I don…

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement. Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof. I c…

Wow. It’s like the printing press! Reproduction has hit the next jump in scale point

Re: Pulling my site from Google over AI training

#86

Earlier quoted context omitted.

Why do you think it would be illegal? You can state "permission is not granted to X" on anything you want, but that doesn't mean the law is on your side. Regular rules of copyright still apply. P.S. Permission is not granted to downvote my comment!

[flagged]

Vitriol aside, you need to chill for a bit and touch grass.

"Training" doesn't really have a well-defined meaning, I could use your website to train something as simple as a histogram of word counts for an AI for example. Nothing about that constitutes copyright infringement under even the loosest definition of their legal concept.

Additionally weights from training and the AI's output are two completely different matters from a legal perspective as well.

Re: Pulling my site from Google over AI training

#87

Comments are bound to be spicy on this one. I always love it when techbros say that AI learning and human learning are exactly the same, because reading one thing at a time at a biological pace and remembering takeaway ideas rather than verbatim passages is obviously exactly the same thing as processing millions of inputs at once and still being able to regurgitate sources so perfectly that verbatim copyrighted conte…

They only believe it because they think their own labour isn't under threat. Think.

Re: Pulling my site from Google over AI training

#88

Comments are bound to be spicy on this one. I always love it when techbros say that AI learning and human learning are exactly the same, because reading one thing at a time at a biological pace and remembering takeaway ideas rather than verbatim passages is obviously exactly the same thing as processing millions of inputs at once and still being able to regurgitate sources so perfectly that verbatim copyrighted conte…

They only believe it because they think their own labour isn't under threat. Think .

I'm going to savor every last iota of the Shocked Pikachu energy they emit. Especially when they realize, far too late, that building their almighty replacements gets them exactly zero kudos from the people who will inevitably control it.

I'm not going to be apologetic about it either, since the same people who think they're invulnerable also tend to espouse sadistic glee over the impending immiseration of millions due to these developments... just practicing what they preach, after all.

Re: Pulling my site from Google over AI training

#89
post #77

Earlier quoted context omitted.

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement. Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof. I c…

Wow. It’s like the printing press! Reproduction has hit the next jump in scale point

This is a great analogy.

Re: Pulling my site from Google over AI training

#90
post #83

Earlier quoted context omitted.

That's completely subjective. What you describe is a spectrum. If that were true, thesis' would not have to be run through plagiarism scans (they are), and there would be no copyright lawsuits for familiarity such as Ed Sheeran's Marvin Gaye lawsuit. It's also important to note these laws were always intended to strike a fair balance between the copyright owner and the good of society as a whole. Copyright is not an…

This is such an important point that I think escapes a lot of people, both on the pro-AI and anti-AI side. Although everyone probably wishes otherwise, there isn't really a hard objective line for determining if something is infringing or not. And that's actually a good thing! But it means that, in the end, it's up to a judge to weigh a bunch of different, rather fuzzy, factors.

Law is not a static entity, it's morphing and evolving to suit our ideals today. Learning wasn't an issue until it is.

We could very well distinguish between machine learning and human learning although they're no different from each other in principle.

And although, we can't say that ML is plagiarism, at scale it breaches our other moral principles like privacy and individual identity.

Let's say in 10 years from now Google'll train ML on (close to) all the data available in the world, be it text, visual or audio. Would they be able to prompt this AI: "You are George Wilkinson from Colorado ** st. 23". How close will it be from real George on whose data it was trained on?

We can't tune our human learning to this level of precision, but it is only a matter of time for the machine.

Post reply on HN