Pulling my site from Google over AI training
51–60 of 102 posts
Re: Pulling my site from Google over AI training
#52Earlier quoted context omitted.
I keep having to say this: An ai is not a person
Yep, there’s a big difference in practice. If an AI could attribute and provide royalties then it may not be so different but that’s never going to happen. A big reason Bard exists is Google trying to ensure they stay profitable and relevant. They don’t care where the knowledge really comes from.
Re: Pulling my site from Google over AI training
#53Can you pollute their data with hidden elements, or do they only scrape visible stuff?
Re: Pulling my site from Google over AI training
#54Earlier quoted context omitted.
I miss webrings. I don't think that search getting better killed them (they are useful for reasons unrelated to the state of search), but when personal and hobbyist websites started vanishing, webrings went with them.
Webrings were a great, essentially curated list of sites the authors of sites you liked thought you might also enjoy or find useful. I miss them, too.
Re: Pulling my site from Google over AI training
#55Earlier quoted context omitted.
[flagged]
I agree it is wrong and should be illegal. That being said, I do find the argument that's it's no different than a human learning from and occasionally reconstructing copyrighted things compelling.
Re: Pulling my site from Google over AI training
#56An interesting thought are government or foreign actors training a generative Ai. They won't abide by any no-index tax and will scrape any and everything.
Re: Pulling my site from Google over AI training
#57Just today I googled (and duck duck go'd?) alternatives for Discord (because reasons). Entire search results page was "X top alternatives to Discord." It was all blog posty kind of stuff with an "author".
And like 90% of it was written by indian and african sounding names. These were clearly "content farms" with low paid labour and bad grammar, or just authors with nothing better to do than write Yet Another Blog Post about Top Discord Alternatives. Sure, they weren't generated, but the fact that a human was involved in creating something crappy doesn't make it better or unique.
What I was actually looking for was unique content. Either an actual curated list of alternatives (NOT a blog post they update every year). Or an extract from a book where someone posted fiction about a fictional Discord user that meets aliens. Or comments in a forum, or a link to a song-lyrics website for a Weird Al parody song about discord, a website dedicated to expounding the virtues cutting the discord cord, a link to a PDF where someone saved a IRC chat server's logs about a person switching from discord to IRC, or an "IRC-MF do you speak it" crass website, or something. Anything but a damn content blog post by some third-world content creator or hipster-blog-poster from the 1st world.
What I got was garbage. Human-level garbage. Garbage that across hundreds of thousands of websites basically took a piece of content and expanded it with every known combination of words, sentences, and mini-stories and pasted it on a stupid blog post with an author.
And this garbage is what this AI is training on so we can have content farms make more copies of itself with more variations and in different languages now, all so we can pay Google et al attention-coins to magically sift through all that garbage and present us with something a little less garbage-y for us to consume.
Re: Pulling my site from Google over AI training
#58My pitch: A search engine that exclusively indexes noindex sites (you can use other sites while spidering) and builds an LLM model with the results.
I rather suspect that this is already being done.
Re: Pulling my site from Google over AI training
#59It is reading and learning. A person would read and learn.
This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work.
This is no different. I can write some code and use it, subconsciously referencing a work.
If I don't check my written work and put it out there, someone might have a claim against me. If I don't check the machine generated work and put it out there, someone might have a claim against me.
OpenAI,Meta,et al are providing the model, basically a regression model or tool. I'm adding the variables or secret sauce that makes it output that set of data in that specific order not them. It'd be like suing Parker for making the pen.
Re: Pulling my site from Google over AI training
#60Earlier quoted context omitted.
Does that include book readers for the blind? They typically have some sort of optical character recognition and benefit a user, just like an ML training dataset benefits users. My point being: it's exceptionally hard to create laws that deny precisely what you don't want and allow precisely what you want, without quickly getting into details that bring the entire law's assumptions into question. Here being "because…
The main difference here is that these AI bots are operating with an entirely different agenda. The ethics remain to be seen and the jury is out as to whether they will benefit the user they way the promise they will. Also on a whole different scale and instead of supplementing the web content it’s devaluing it to a degree.
Basically you're assuming the agenda of the operator, saying "that's bad an shouldn't be allowed". But I see the web- except for things specifically labelled with standard copyright disclaimers- as effectively a large corpus of publicly available data, "in the market square for all to see".