Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

71–80 of 102 posts

Re: Pulling my site from Google over AI training

#71
post #50

Don't forget about also pulling your site from Bing! It would be naive if you somehow trust that Microsoft won't use your site for AI training.

And Yandex!

Side note, I have friends that crawled a massive amount of the internet over several months for their own purposes.. at this point it's probably impossible to exclude your site since tons of other people probably link to your site if it's at all of value.

Re: Pulling my site from Google over AI training

#72

This whole AI scraping argument is so silly to me. If you don't want people downloading and processing your content, then don't post it on the public internet?

They literally outline their reasoning in a link[0]. There’s a significant gap between offering information for someone to read for free (where I can see the author and choose to respect their terms, if they have any), and a huge tech company aggregating that data where it is assimilated into a model and used in a product they will profit from[1]. They are exploiting gray area regarding digital rights, copyright, etc.

[0] https://tracydurnell.com/2023/07/07/the-next-big-theft/

[1] https://www.tumblr.com/nedroidcomics/41879001445/the-interne...

Re: Pulling my site from Google over AI training

#73
post #50

Don't forget about also pulling your site from Bing! It would be naive if you somehow trust that Microsoft won't use your site for AI training.

And Yandex! Side note, I have friends that crawled a massive amount of the internet over several months for their own purposes.. at this point it's probably impossible to exclude your site since tons of other people probably link to your site if it's at all of value.

Yeah, this is why robots.txt is garbage. Too many crawlers completely ignore it, so I stopped bothering with it (I have it set to block all bots, but I don't expect it to be effective.)

Instead, I'd just keep an eye on my access logs and block obvious crawlers when I saw them.

Re: Pulling my site from Google over AI training

#74
post #64
post #53

Earlier quoted context omitted.

If Google thinks a site is serving different content to googlebot vs. real users, it will stop returning that site on the SERP, because that is a malware distribution technique, among other reasons.

What if you serve the same site? Does googlebot know that some text has the same color as the background?

I don't know, that's not my game and I never worked on crawl or index, but I imagine they have that handled since invisible keyword spam at the bottom of the page was already a universal spammer technique by 1997.

Re: Pulling my site from Google over AI training

#75
post #70
post #4

> Blocking bots that collect training data for AIs (and more) > In addition, I created a robots.txt file to tell “law abiding” bots what they’re not allowed to look at. I ought to have done this before but kind of assumed it came with my WordPress install (Nope.) > I specifically want to deter my website being used for training LLMs, so I blocked Common Crawl. Instead of blocking, it would be neater to present and al…

> Instead of blocking, it would be neater to present and alternative version to the crawlers IIUC, if a site presents different (view of) content to the crawler than users, the site can get de-indexed.

If that's a concern, and you're only worried about AI crawlers, then that's not a problem. Only provide the bogus pages to the AI crawlers, not to the search engine crawlers.

Assuming there's a difference, anyway. I suspect that with Google and Bing, there isn't.

Re: Pulling my site from Google over AI training

#76
post #69

Earlier quoted context omitted.

They're not doing much they're updating probabilities on a regression model. What the user of the tool thereafter do is the question.

Many LLM's will happily recite large segments of copyrighted material word-for-word, despite the fact that it can be difficult to tell what's happening "under the hood".

Many people can do that too? It's what they do with it that's important.

Re: Pulling my site from Google over AI training

#77

I can't understand the outrage. In practice absolutely nothing has changed. It is reading and learning. A person would read and learn. This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work. This is no different. I can write some code and use it, subconsciously referencing a work. If I don…

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement.

Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof.

I can’t understand why one would expect people to go “oh it’s technically ‘learning’ I guess I’ll ignore all the consequences that weren’t present when it was just humans”.

Re: Pulling my site from Google over AI training

#78
post #38
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

> I don't think training on copyrighted stuff will be ever banned, but we need to figure out how much they can be allowed to generate based on that.

From a US copyright law point of view, this is most likely correct. Copyright law doesn't prevent you from ingesting copyrighted works, it prevents you from distributing them.

There is also a great deal of existing case law about how different a work has to be before it's not infringing from another work anymore. There are existing rules of thumb judges go by when trying to determine if infringement occurred. They include things like the amount of difference in expression, the quantity, whether or not it's incidental, etc.

And that's not even getting into the question of fair use -- which is a whole other kettle of fish.

I suspect that the courts will deal with these issues the way that they've always dealt with these issues: on a case-by-case basis.

Re: Pulling my site from Google over AI training

#79
post #40

Earlier quoted context omitted.

I keep having to say this: An ai is not a person

Yep, there’s a big difference in practice. If an AI could attribute and provide royalties then it may not be so different but that’s never going to happen. A big reason Bard exists is Google trying to ensure they stay profitable and relevant. They don’t care where the knowledge really comes from.

> If an AI could attribute and provide royalties then it may not be so different

But even that requires the permission of the copyright holder. Nobody is required to accept an infringing use of their work in exchange for royalties.

Re: Pulling my site from Google over AI training

#80
post #77

I can't understand the outrage. In practice absolutely nothing has changed. It is reading and learning. A person would read and learn. This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work. This is no different. I can write some code and use it, subconsciously referencing a work. If I don…

A human cannot learn from and re/produce work they view at the speed, volume, and scale that “AI” does, nor can that human be infinitely replicated and farmed out. When describing this gap in capabilities, or the consequences of “learning”, “orders of magnitude” would be a comical understatement. Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof. I c…

That's completely subjective. What you describe is a spectrum. If that were true, thesis' would not have to be run through plagiarism scans (they are), and there would be no copyright lawsuits for familiarity such as Ed Sheeran's Marvin Gaye lawsuit.

It's also important to note these laws were always intended to strike a fair balance between the copyright owner and the good of society as a whole. Copyright is not an end in itself.

Post reply on HN