Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

61–70 of 102 posts

Re: Pulling my site from Google over AI training

#61
post #47

Earlier quoted context omitted.

I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.

Is it legal to transcribe a book from memory for money? Does it matter how faithful the transcription is?

> Is it legal to transcribe a book from memory for money?

If it's an accurate transcription and you don't have permission, then it's not legal. It doesn't matter if it's for money or not (or if it's from memory or not).

> Does it matter how faithful your transcription is?

Yes, it matters. Copyright covers the specific expression of an idea, not the idea itself.

Re: Pulling my site from Google over AI training

#62

As a substack author, with “permission is not granted to use any portion of this to train an ai” at the bottom of most of my posts, it’s bullshit that you have to do this sort of thing, and that it will almost certainly not work This must be illegal, but how are all the little bloggers going to oppose it?

What's right or wrong, and what's legal or illegal, are two different things. There are plenty of right things that are illegal and wrong things that are legal.

Re: Pulling my site from Google over AI training

#63
post #44

Earlier quoted context omitted.

But you don't know what the intention of the reader human is either? It could be that too?

Sure, but that would be illegal too. I'm saying it doesn't matter who reads your website, but everyone knows exactly why GPT and Bard are going to do with the information they're "learning" from it, so they're trying to block it from reading in the first place.

They're not doing much they're updating probabilities on a regression model. What the user of the tool thereafter do is the question.

Re: Pulling my site from Google over AI training

#64
post #53
post #10

Can you pollute their data with hidden elements, or do they only scrape visible stuff?

If Google thinks a site is serving different content to googlebot vs. real users, it will stop returning that site on the SERP, because that is a malware distribution technique, among other reasons.

What if you serve the same site? Does googlebot know that some text has the same color as the background?

Re: Pulling my site from Google over AI training

#65

Earlier quoted context omitted.

Why do you think it would be illegal? You can state "permission is not granted to X" on anything you want, but that doesn't mean the law is on your side. Regular rules of copyright still apply. P.S. Permission is not granted to downvote my comment!

[flagged]

Ok, so if it's already copyright infringement then what does writing "permission is not granted" at the bottom of your post do, exactly?

Re: Pulling my site from Google over AI training

#66

[flagged]

"Ironically, the uniformity of the copies of Gutenberg’s Bible led many superstitious people of the time to equate printing with Satan because it seemed to be magical. Printers’ apprentices became known as the "printer’s devil." In Paris, Fust [a typographer] was charged as a witch. Although he escaped the Inquisition, other printers did not." (The Unsung Heroes, a History of Print by Dr. Jerry Waite 2001)"

Re: Pulling my site from Google over AI training

#67
post #45
post #38

Earlier quoted context omitted.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

If I recite the vague plot of a novel or a fact I learned from an encyclopedia I'm not reproducing anything, certainly not violating copyright law. I don't see why AI developers should be expected to think otherwise and worsen their training data over this.

Scale and position matter. Google is the conduit that connects most people to most websites, so in the EU they are considered a "gatekeeper" and need to be careful about conflicts of interest with the people and websites using their "gate". I hope American competition law catches up to the point we can recognize that market makers simply should not be participating in the markets they make (and Google search is a market maker; it's connecting "buyers" [viewers or advertisers, depending on your perspective] to "sellers" [websites or viewers, respectively]), but I digress.

The point is that Google has a certain market position that makes it very different when they "recite the vague plot of a novel or a fact they learned". The point of competition law is to "distort" free market capitalism for the betterment of society. This is one of those cases where practical considerations trump information idealism. The quality of information on the internet will go down if we stop rewarding original publishers.

Re: Pulling my site from Google over AI training

#68
post #38
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

> It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later

It could be illegal if the AI reproduces vast portions of it. If you could ask the LLM over a course of prompts to generate a significant portion of the content (as the copyright law defines it), then yes.

As long as the AI isn't reproducing it, then I am not sure if it would count.

Re: Pulling my site from Google over AI training

#69
post #44

Earlier quoted context omitted.

Sure, but that would be illegal too. I'm saying it doesn't matter who reads your website, but everyone knows exactly why GPT and Bard are going to do with the information they're "learning" from it, so they're trying to block it from reading in the first place.

They're not doing much they're updating probabilities on a regression model. What the user of the tool thereafter do is the question.

Many LLM's will happily recite large segments of copyrighted material word-for-word, despite the fact that it can be difficult to tell what's happening "under the hood".

Re: Pulling my site from Google over AI training

#70
post #4

> Blocking bots that collect training data for AIs (and more) > In addition, I created a robots.txt file to tell “law abiding” bots what they’re not allowed to look at. I ought to have done this before but kind of assumed it came with my WordPress install (Nope.) > I specifically want to deter my website being used for training LLMs, so I blocked Common Crawl. Instead of blocking, it would be neater to present and al…

> Instead of blocking, it would be neater to present and alternative version to the crawlers

IIUC, if a site presents different (view of) content to the crawler than users, the site can get de-indexed.

Post reply on HN