Live data from Hacker News

Pulling my site from Google over AI training

tracydurnell.com

51–60 of 102 posts

Re: Pulling my site from Google over AI training

#52
post #40

Earlier quoted context omitted.

I keep having to say this: An ai is not a person

Yep, there’s a big difference in practice. If an AI could attribute and provide royalties then it may not be so different but that’s never going to happen. A big reason Bard exists is Google trying to ensure they stay profitable and relevant. They don’t care where the knowledge really comes from.

Don't forget to provide royalties for every synapse in your head.

Re: Pulling my site from Google over AI training

#53
post #10

Can you pollute their data with hidden elements, or do they only scrape visible stuff?

If Google thinks a site is serving different content to googlebot vs. real users, it will stop returning that site on the SERP, because that is a malware distribution technique, among other reasons.

Re: Pulling my site from Google over AI training

#54
post #8

Earlier quoted context omitted.

I miss webrings. I don't think that search getting better killed them (they are useful for reasons unrelated to the state of search), but when personal and hobbyist websites started vanishing, webrings went with them.

Webrings were a great, essentially curated list of sites the authors of sites you liked thought you might also enjoy or find useful. I miss them, too.

yah, come to think of it in the curated space, this reminds me of that awesome X family of github pages. Looks like someone compiled a bunch of them here https://github.com/sindresorhus/awesome#databases. I have found those to be highly valuable treasure troves pregnant with rich and relevant information.

Re: Pulling my site from Google over AI training

#55

Earlier quoted context omitted.

[flagged]

I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.

Most normal humans do not spend their time profitably selling their "occasionally reconstructing copyrighted things" at a rate a millions of users per second, which is a pretty important difference in practice.

Re: Pulling my site from Google over AI training

#56

An interesting thought are government or foreign actors training a generative Ai. They won't abide by any no-index tax and will scrape any and everything.

It doesn't have to be government or foreign actors. Abiding by the contents of robots.txt isn't required of anybody at all. It's merely a social convention.

Re: Pulling my site from Google over AI training

#57
Bit of a related rant.

Just today I googled (and duck duck go'd?) alternatives for Discord (because reasons). Entire search results page was "X top alternatives to Discord." It was all blog posty kind of stuff with an "author".

And like 90% of it was written by indian and african sounding names. These were clearly "content farms" with low paid labour and bad grammar, or just authors with nothing better to do than write Yet Another Blog Post about Top Discord Alternatives. Sure, they weren't generated, but the fact that a human was involved in creating something crappy doesn't make it better or unique.

What I was actually looking for was unique content. Either an actual curated list of alternatives (NOT a blog post they update every year). Or an extract from a book where someone posted fiction about a fictional Discord user that meets aliens. Or comments in a forum, or a link to a song-lyrics website for a Weird Al parody song about discord, a website dedicated to expounding the virtues cutting the discord cord, a link to a PDF where someone saved a IRC chat server's logs about a person switching from discord to IRC, or an "IRC-MF do you speak it" crass website, or something. Anything but a damn content blog post by some third-world content creator or hipster-blog-poster from the 1st world.

What I got was garbage. Human-level garbage. Garbage that across hundreds of thousands of websites basically took a piece of content and expanded it with every known combination of words, sentences, and mini-stories and pasted it on a stupid blog post with an author.

And this garbage is what this AI is training on so we can have content farms make more copies of itself with more variations and in different languages now, all so we can pay Google et al attention-coins to magically sift through all that garbage and present us with something a little less garbage-y for us to consume.

Re: Pulling my site from Google over AI training

#58
post #39

My pitch: A search engine that exclusively indexes noindex sites (you can use other sites while spidering) and builds an LLM model with the results.

I rather suspect that this is already being done.

I seem to recall meta-search sites that only listed sites that had disappeared due to DMCA takedowns. There was also "unsafe search" that ran your google search twice, once with "SafeSearch" enabled, and returned the complement of the union, i.e. just the porn.

Re: Pulling my site from Google over AI training

#59
I can't understand the outrage. In practice absolutely nothing has changed.

It is reading and learning. A person would read and learn.

This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work.

This is no different. I can write some code and use it, subconsciously referencing a work.

If I don't check my written work and put it out there, someone might have a claim against me. If I don't check the machine generated work and put it out there, someone might have a claim against me.

OpenAI,Meta,et al are providing the model, basically a regression model or tool. I'm adding the variables or secret sauce that makes it output that set of data in that specific order not them. It'd be like suing Parker for making the pen.

Re: Pulling my site from Google over AI training

#60
post #49
post #34

Earlier quoted context omitted.

Does that include book readers for the blind? They typically have some sort of optical character recognition and benefit a user, just like an ML training dataset benefits users. My point being: it's exceptionally hard to create laws that deny precisely what you don't want and allow precisely what you want, without quickly getting into details that bring the entire law's assumptions into question. Here being "because…

The main difference here is that these AI bots are operating with an entirely different agenda. The ethics remain to be seen and the jury is out as to whether they will benefit the user they way the promise they will. Also on a whole different scale and instead of supplementing the web content it’s devaluing it to a degree.

The "ai bots" aren't operating with an agenda- at least as far as we can tell now, training algorithms and their scrapers do not have agency.

Basically you're assuming the agenda of the operator, saying "that's bad an shouldn't be allowed". But I see the web- except for things specifically labelled with standard copyright disclaimers- as effectively a large corpus of publicly available data, "in the market square for all to see".

Post reply on HN