Live data from Hacker News

Google scrambles to manually remove weird AI answers in search

theverge.com

341–350 of 387 posts

Re: Google scrambles to manually remove weird AI answers in search

#341

Earlier quoted context omitted.

I can ask a human to explain the steps they took to answer a question. I can ask a human a question 100 times and I don't get back 100 different answers. None of those applies to an LLM.

Isn't it actually known that every time a human brain recalls a piece of memory the memory gets slightly changed? If the answer has any length at all, I imagine the answer can vary every single time the person answers, unless they prepared for it, memorized it word by word.

It's also known that the brain is prone to outright constructing demonstrably fictional rationalisations of decisions it's never made.

Any notion that we're reliable narrators of our own thoughts and actions is fiction.

Re: Google scrambles to manually remove weird AI answers in search

#342

Earlier quoted context omitted.

> Except that 90% of Reddit isn't garbage. It's really useful. Citation needed. I've been a Reddit user since its inception and honestly except for niche hobby subreddits, Reddit is mostly low effort garbage, bots and rehashed content. I'd wager that mainstream subreddits are 99% garbage for training an LLM for anything other than shitposting.

Even in the niche hobby subreddits there can be a really high garbage factor. There's plenty of well meaning posters that are just wrong. They're not trying to mislead or lying they're just unaware they're wrong.

The good answers tend to use links as well, which won't capture well. In many political and local subreddits there's a huge amount of Russian and far right sock puppet activity. Good luck training an AI to understand political opinions or what people in an area are like when most of the longer comments are pre written copy pasted talking points from astro turf groups and bad actors.

Re: Google scrambles to manually remove weird AI answers in search

#343

Earlier quoted context omitted.

Or to put it another way, I think Google should have a way of saying "yes, we know this result is wrong, but we're leaving it in because it's funny." There is a demand for funny results. Someone asking “how many rocks should I eat” is looking for entertainment, so you might as well give it to them.

The right answer is no rocks. Some mentally ill person could type that in and get "eat 1000 rocks" and then die from eating rocks, and that would be Google's fault. It's not funny. I have no doubt right now there are at least 50 youtube videos being made testing different glue's effectiveness holding cheese on a pizza. And some of those idiots are going to taste-test it, too. And then people will try it at home, some…

It's funny how people on reddit think that these LLMs will somehow become AGI in the next year, or when openAI releases gpt 5.

The reality is though, that there is no known path currently to true AGI system and the research needs to be done. No one knows how to build this kind of system yet. LLMs are nice for things like roleplaying, helping with code stuff etc. but they are far from all that the marketers hype them to be.

Re: Google scrambles to manually remove weird AI answers in search

#344

Earlier quoted context omitted.

Google Answers was launched in 2002 and retired in 2006. https://en.wikipedia.org/wiki/Google_Answers

"users would pay someone else to do the search." My notion isn't a rehash of Google Answers. Google pays the "someone else", not you.

You assume that a modern day tech giant would hire an army of experts, instead of just outsourcing it to the lowest bidder in the current third world country of choice.

Re: Google scrambles to manually remove weird AI answers in search

#346
post #341

Earlier quoted context omitted.

Isn't it actually known that every time a human brain recalls a piece of memory the memory gets slightly changed? If the answer has any length at all, I imagine the answer can vary every single time the person answers, unless they prepared for it, memorized it word by word.

It's also known that the brain is prone to outright constructing demonstrably fictional rationalisations of decisions it's never made. Any notion that we're reliable narrators of our own thoughts and actions is fiction.

Right, that makes sense, our brain will do a quick black box judgment (some may call it system 1), and then rational process only works to justify that or explain the black box, assuming that black box is always correct (depending on the person and how much they trust their black box or system 1).

So system 2 is "hallucinating" the best justification for system 1.

And usually system 2 will do it only when it's required to justify it for anyone else.

Re: Google scrambles to manually remove weird AI answers in search

#347
post #177

Earlier quoted context omitted.

What I've been wanting to see for a while now is a social-network based search engine: * No pages are indexed automatically. The only indexed pages are pages that users say are worth indexing. Probably have a browser add-on for a one button click that people can use. * You can friend/follow others * Your search results are a combination of your own indexed pages and the pages indexed by people in your network.

Isn’t that what Reddit is or digg was ? Link aggregators ? Gaming that is solved problem , you can use human bot farms to brigade and astroturf and you can even motivate people to do it for free . If cost of spamming is cheaper than cost of moderation, spam will win

Not quite like reddit and digg. You can bot farm those because the lists are common to all.

In this search engine, let's say there's you, me, person three, and spammer. You you are following me, and I'm following person three. Spammer isn't in any of our networks.

When you use the search engine, you only see results that you, me, or person three manually tagged as worthwhile. Any pages or content that Spammer tagged as worthwhile aren't part of your results, because they aren't in your network. So they can try to game the system all they want, but it won't affect you.

If person three starts following Spammer, I can unfollow them and then Spammer's results will no longer be included in your search results (or you can unfollow me and avoid those results).

I imagine rankings would also be affected by degrees of separation, so even if you followed me, I followed person three, and person three followed spammer, results tagged by me and you would take much higher precedence than results tagged by Spammer.

This also allows you to make custom searches by choosing which people you follow to include. Suppose you want to search for good headphones, so you make a search, but only include people in your network that you know are music and audio savvy, so that the results reflect the pages tagged by those people.

Re: Google scrambles to manually remove weird AI answers in search

#348

My whole qualm with this AI integration into search engines: it's a search engine, not a question engine. I go to google to search the internet for something, not ask it a question. IMO, asking AI for something is a different task than searching the internet. It's sorta the same problem as if I go into a store and ask an employee where something is, and they reply with "well what are you trying to do?"

Recently I searched Google for a slightly unlikely phrase — in quotation marks — and Google proudly told me that my phrase was grammatically correct.

And nothing else. They didn't give me any search results. Or even tell me there weren't any results. Or even give me a button to press to say "no, I really wanted to search the internet for this phrase".

And also I have zero interest in Google's opinion on English grammar and am frankly insulted to be offered it, although to be fair I'm probably in a minority worldwide on that one.

If I can't use Google to search the internet for things, then Google is eventually going to have a big problem.

Re: Google scrambles to manually remove weird AI answers in search

#349

This approach to remove bad search suggestions manually reminded of a different approach Google once took, where they weren’t satisfied with manually tweaking search results but rather wanted to tweak the algorithm that produces these results when there were bad results. 'Around 2002, a team was testing a subset of search limited to products, called Froogle. But one problem was so glaring that the team wasn't comfort…

The solution is always the same: pay people off and keep it under the radar. What stops the vendor, or other vendors, from creating more gnomes with sneakers. Easy money from customer with billions of dollars to spend on payola, fines, legal settlements, etc. Maybe they made the vendor sign an NDA.

> The solution is always the same: pay people off and keep it under the radar.

You’re making this into a conspiracy unnecessarily. They didn’t “pay people off”, they bought an item. Do you “pay off” your grocer when you buy a carrot from them?

> What stops the vendor, or other vendors, from creating more gnomes with sneakers.

The fact they don’t know their entry was causing this issue to a major corporation?

> Maybe they made the vendor sign an NDA.

Why would they? Someone had one gnome with sneakers for sale; someone else bought it; end of story.

Re: Google scrambles to manually remove weird AI answers in search

#350
post #300

Earlier quoted context omitted.

I am glad that you think it an excellent point! I think that this might be a nice way of getting round my objection, but there is one worry, which is that X is relative to a distribution on the questions we ask when we aren’t dealing with Encyclopedia Eric but with an LLM. I don’t actually use LLMs very much myself, partly out of arrogance and Luddite tendencies. But I suspect that the value of X for some sorts of qu…

A relevant point to this is the notion of "System-1" vs "System-2" thinking. Somewhat dubious when applied to actual human psychology but I think a valid metaphor for how LLMs work: they are only capable of System-1 thinking; a single forward pass through the weights of intuition In my actual life, I don't trust my own System-1 thoughts: for anything important, I'm always going to engage System-2. And LLMs don't have…

I am not sure that the ambiguity has to work that way. The suggestion I am making is that if we fix the distribution on questions and the training data, we might (a) know the value of X in this specific case, (b) be able to ensure that it is fairly high.

I’d say this is the murky case because on the fixed training data and query distribution X ≈ 1 and we know that even though we don’t know the value of X on other training data and other query distributions. I think that might be where the disagreement lies.

Post reply on HN