Live data from Hacker News

Google’s AI is being manipulated. The search giant is quietly fighting back

bbc.com

31–40 of 235 posts

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#31
post #18

Yeah, the internet seems like a big poison pill. Training on the whole internet feels like citing the National Enquirer (or the Daily Mail?) for a school essay. Having an archive of "curated" training data seems like it is going to be important. Otherwise you need "AS" (artificial skepticism) introduced into future models. ("But I read it on the internet!", ha ha.) Or perhaps there are ways to bucket training data su…

> Training on the whole internet feels like citing the National Enquirer It's not, though, because the refutations are in the training data too. This isn't actually the problem being described. The weights in the LLM are fine. It's that the task the LLM is being asked to do is to search and summarize new content that isn't in its training data. And it does it too much like a naive reader and not enough like a cynical…

"…the LLM is being asked to do is to search and summarize new content that isn't in its training data…"

If it fails at that then it is a pretty significant problem. As you say earlier "the refutations are in the training data too", then the LLM should in fact be able to use "both sides" and land with a little better confidence when presented with new data.

(Hopefully your point regarding prompting issues is resolved then.)

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#32
The best way to fight back is to not play the game at all. AI slop has completely ruined the internet, it's not going to get better. It was already on a massive downard trend pre-AI and generative AI has only accelerated the decline by 100x. It's only going to get worse from here.

uBlock Origin: Settings -> Filter Lists -> EasyList –> Annoyances -> EasyList –> AI Widgets

It's not perfect but the internet feels slightly better when AI garbage is not constantly being shoved in my face 24/7.

I want to go one step further -> I want to hide widgets, but I also want to intercept the request it would have made and replace the payload with garbled nonsense. Similar to how Ad Nauseam will hide ads but it also clicks every single one to poison the data collection.

And for this reason alone you will pry Firefox from my cold, dead hands.

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#33

> I was able to demonstrate the problem by publishing a single article on my personal website about my hot-dog-eating prowess. One blog post ... that's all it takes. i'm actually surprised it's that bad . i would have thought it'd take more effort, but i guess it could depend on some sort of purposeful weighting based on search rank during training? > If a company or website is caught breaking the rules, it could be…

I don't think Google even indexes my blog, but these people were able to get a new post into all major LLMs within 24 hours?

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#34
If you ask Google "what's the name of the whale in half moon bay harbor?" it still confidently includes Teresa T in the AI summary, thanks to my frankly amateur attempt at index poisoning from a year and a half ago: https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-p...

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#35
post #34

If you ask Google "what's the name of the whale in half moon bay harbor?" it still confidently includes Teresa T in the AI summary, thanks to my frankly amateur attempt at index poisoning from a year and a half ago: https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-p...

Aren't you afraid Google will send you a threat for an attempt to manipulate AI responses?

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#36
post #12

Earlier quoted context omitted.

The strength of the sources are not a question of quantity. A hundred obscure blog post have not the same strength as one wikipedia link, because the latter is more trustworthy. There could be some indication beside the info showing the strength of the sources (how many major trustworthy sources support it, etc.).

Seems like a tall order to do that for literally everything. I guess there’ll be some guy at google going through every blog and saying whether it’s reliable or not?

This is what Google has been doing, via various methods, for 25 years.

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#37

> I was able to demonstrate the problem by publishing a single article on my personal website about my hot-dog-eating prowess. One blog post ... that's all it takes. i'm actually surprised it's that bad . i would have thought it'd take more effort, but i guess it could depend on some sort of purposeful weighting based on search rank during training? > If a company or website is caught breaking the rules, it could be…

I don't think Google even indexes my blog, but these people were able to get a new post into all major LLMs within 24 hours?

Google indexes other people's blogs.

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#38

Earlier quoted context omitted.

> Having an archive of "curated" training data seems like it is going to be important the justification for not doing that is probably "prohibitively expensive given the amount of data involved". they'd need a bunch of human reviewers combing through massive troves of data. it's probably cheaper to "sort of fix" it after the fact. > perhaps there's ways to bucket training data such that the model is aware of which da…

"…they'd need a bunch of human reviewers combing through massive troves of data…" Yeah, I concede that. It doesn't need to be done over night. Having a static repo of data though that you can work through over time (years)—removing some data, add pre-curated data to. In so many years you can have a pretty good "reference dataset".

I think some of the thousands of people working on training LLMs have tried some of the low-hanging-fruit ideas we can brainstorm of the top of our head 5 years later.

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#39
post #26

Earlier quoted context omitted.

The article also said this: “ But our investigation also found the same trick being used to dismiss health concerns about medical supplements or influence financial information provided by Google's AI about retirement.” That’s a lot more alarming than just hotdogs.

They should provide the queries then, because it's likely the same trick people have used for decades now with SEO'ing blog posts to appear as "3rd party review" for their shitty products. I create a supplement called Xanatewthiuy, I write blogs/make websites that appear totally unaffiliated saying positive things about "Xanatewthiuy", and then when people see my ads and search for "Xanatewthiuy", the only results ar…

Right now, using Google searching for "what is Xanatewthiuy" , the AI summary is not generated, but the only search result previews as

> Xanatewthiuy is a supplement that dramatically lowers anxiety from media induced hysteria, primarily stemming from carefully worded pieces meant ...

Re: Google’s AI is being manipulated. The search giant is quietly fighting back

#40
After reading this, I'm thinking of trying some AI data poisoning. I'm going to spam my website with hidden text that only AI scrapers can read, claiming I'm a 'highly excellent programmer' just to advertise my site. I really hope it drives a lot of traffic. I'm honestly sick and tired of getting zero comments on my website
Post reply on HN