Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

31–40 of 102 posts

Re: Pulling my site from Google over AI training

#32

Why the preference not to have Google train their AI using your website content?

One reason is the same as authors who don’t want their books trained and actors who don’t want their likeness trained. If that content is valuable it allows google to realize that value with no return to the creator.

The naivety is assuming a small number of people making access hard will affect the value they're able to realize.

Re: Pulling my site from Google over AI training

#33

This whole AI scraping argument is so silly to me. If you don't want people downloading and processing your content, then don't post it on the public internet?

It's not about not wanting people to access the content, it's about not wanting AI bots to do so. But yes, I agree, the only realistic defense we have is to remove it from the publicly-accessible web.

Re: Pulling my site from Google over AI training

#34
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

This is the most braindead take and aibros keep pushing. Your 'AI' model is NOT A PERSON and therefore it is different.

Does that include book readers for the blind? They typically have some sort of optical character recognition and benefit a user, just like an ML training dataset benefits users.

My point being: it's exceptionally hard to create laws that deny precisely what you don't want and allow precisely what you want, without quickly getting into details that bring the entire law's assumptions into question. Here being "because an ML training is not a person, it has no right to scan the web".

Re: Pulling my site from Google over AI training

#35

Earlier quoted context omitted.

Why do you think it would be illegal? You can state "permission is not granted to X" on anything you want, but that doesn't mean the law is on your side. Regular rules of copyright still apply. P.S. Permission is not granted to downvote my comment!

[flagged]

I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.

Re: Pulling my site from Google over AI training

#36
post #33

This whole AI scraping argument is so silly to me. If you don't want people downloading and processing your content, then don't post it on the public internet?

It's not about not wanting people to access the content, it's about not wanting AI bots to do so. But yes, I agree, the only realistic defense we have is to remove it from the publicly-accessible web.

[deleted]

Re: Pulling my site from Google over AI training

#37

Earlier quoted context omitted.

Why do you think it would be illegal? You can state "permission is not granted to X" on anything you want, but that doesn't mean the law is on your side. Regular rules of copyright still apply. P.S. Permission is not granted to downvote my comment!

[flagged]

That is not settled law. In the US the key is if this is derivative or transformative.

You can read a book without your brain getting owned by the author.

It's perfectly reasonable to say it should be considered copyright infringement, but such cases are in the court now.

Disclaimer: I am not a lawyer.

Re: Pulling my site from Google over AI training

#38
post #15

As a substack author, with “permission is not granted to use any portion of this to train an ai” at the bottom of most of my posts, it’s bullshit that you have to do this sort of thing, and that it will almost certainly not work This must be illegal, but how are all the little bloggers going to oppose it?

I don't see why it would be illegal, AI reading it should be no different from anyone else.

It's a bit different because the AI is reading it with the intent of reproducing (certain aspects of) it for other people to later consume without visiting the original site. Fair use doctrine has long held that small pieces of copyrighted material can be reproduced, but the line is very blurry and generally has to be litigated if there's any ambiguity whatsoever. I'd bet many of the models we're currently using today will be pulled from serving the public over copyright lawsuits in the coming years.

I don't think training on copyrighted stuff will be ever banned, but we need to figure out how much they can be allowed to generate based on that. Eventually new models will just pop up with more carefully curated data anyway.

Re: Pulling my site from Google over AI training

#40
post #15

Earlier quoted context omitted.

I don't see why it would be illegal, AI reading it should be no different from anyone else.

I keep having to say this: An ai is not a person

Yep, there’s a big difference in practice. If an AI could attribute and provide royalties then it may not be so different but that’s never going to happen. A big reason Bard exists is Google trying to ensure they stay profitable and relevant. They don’t care where the knowledge really comes from.
Post reply on HN