Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

41–50 of 102 posts

Re: Pulling my site from Google over AI training

#41
post #38
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

But you don't know what the intention of the reader human is either? It could be that too?

Re: Pulling my site from Google over AI training

#42
Comments are bound to be spicy on this one. I always love it when techbros say that AI learning and human learning are exactly the same, because reading one thing at a time at a biological pace and remembering takeaway ideas rather than verbatim passages is obviously exactly the same thing as processing millions of inputs at once and still being able to regurgitate sources so perfectly that verbatim copyrighted content can be spit out of an LLM that doesn't 'contain' its training material. It's even better when they get so butthurt at being called out that they have a nice little rage-cry.

Re: Pulling my site from Google over AI training

#43

Why the preference not to have Google train their AI using your website content?

I bet the reasons vary from person to person. My reason is because I think that these AI systems pose too great of a risk to society, and I want to make sure that I'm not helping them in any way.

Re: Pulling my site from Google over AI training

#44
post #38

Earlier quoted context omitted.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

But you don't know what the intention of the reader human is either? It could be that too?

Sure, but that would be illegal too. I'm saying it doesn't matter who reads your website, but everyone knows exactly why GPT and Bard are going to do with the information they're "learning" from it, so they're trying to block it from reading in the first place.

Re: Pulling my site from Google over AI training

#45
post #38
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using toda…

If I recite the vague plot of a novel or a fact I learned from an encyclopedia I'm not reproducing anything, certainly not violating copyright law.

I don't see why AI developers should be expected to think otherwise and worsen their training data over this.

Re: Pulling my site from Google over AI training

#47

Earlier quoted context omitted.

[flagged]

I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.

Is it legal to transcribe a book from memory for money? Does it matter how faithful the transcription is?

Re: Pulling my site from Google over AI training

#48

Earlier quoted context omitted.

[flagged]

I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.

That said, the law is not made with super-humans being able to reproduce (slightly transformative) as good as all they read (all worlds' knowledge) in mind. A clarifying law should be created.

Re: Pulling my site from Google over AI training

#49
post #34

Earlier quoted context omitted.

This is the most braindead take and aibros keep pushing. Your 'AI' model is NOT A PERSON and therefore it is different.

Does that include book readers for the blind? They typically have some sort of optical character recognition and benefit a user, just like an ML training dataset benefits users. My point being: it's exceptionally hard to create laws that deny precisely what you don't want and allow precisely what you want, without quickly getting into details that bring the entire law's assumptions into question. Here being "because…

The main difference here is that these AI bots are operating with an entirely different agenda. The ethics remain to be seen and the jury is out as to whether they will benefit the user they way the promise they will.

Also on a whole different scale and instead of supplementing the web content it’s devaluing it to a degree.

Post reply on HN