Live data from Hacker News

Google scrambles to manually remove weird AI answers in search

theverge.com

301–310 of 387 posts

Re: Google scrambles to manually remove weird AI answers in search

#301

Earlier quoted context omitted.

I am glad that you think it an excellent point! I think that this might be a nice way of getting round my objection, but there is one worry, which is that X is relative to a distribution on the questions we ask when we aren’t dealing with Encyclopedia Eric but with an LLM. I don’t actually use LLMs very much myself, partly out of arrogance and Luddite tendencies. But I suspect that the value of X for some sorts of qu…

> X is relative to a distribution on the questions we ask when we aren’t dealing with Encyclopedia Eric but with an LLM. Assuming I understand what you mean here correctly, this should be the case for both LLMs and Encyclopedia Eric - there are topics Eric knows by heart (or thinks they know); there are specific phrases seared into his mind through sheer exposure during his life prior to becoming a living Encyclopedi…

I think the thought experiment can be set up both ways. In one case Eric has a fixed probability of getting /any/ query right. (This might leave boosting open so we might want to gerrymander repeated queries out.) In the other this is relative to the distribution of queries.

Re: Google scrambles to manually remove weird AI answers in search

#303
post #297

Earlier quoted context omitted.

Do you have anything to add?

Sure. Your claim has some truth, but is far too strong. The articles you cited upthread do not support the notion that models consistently activate differently when generating true facts vs false facts. It is true that models can capture some notion of reliability based on patterns in their training data. For a concrete example, it is entirely plausible that a model can capture the sense that data trained from Reddit…

I agree with this; that's why I was careful not to use the examples you mentioned. Quoting incorrect training knowledge would be an unavoidable issue if your probe can only say "it's quoting something", and as far as I know it can't do better than that.

But I have seen issues with prompts where a creative-writing prompt and just asking a question look similar, and in that case it could help to know which one it thinks it's doing.

Gemini itself has a funny verification button where it more or less Googles every sentence the model writes and tells you if it seems like it made it up or not.

Re: Google scrambles to manually remove weird AI answers in search

#304

Earlier quoted context omitted.

I am also surprised that training data are not much more curated. Encyclopedias, textbooks, reputable journals, newspapers and magazines make sense. But to throw in social media? Reddit? Seems insane.

The fact is that I think that there is not much written word, to actually train a sensible model on. A lot of books don't have OCRed scans, or a digital version. Humans can extrapolate knowledge from a relatively succinct book and some guidance. But I don't know how a model can add the common sense part (that we already have) that books relies on to transmit knowledge and ideas.

> The fact is that I think that there is not much written word, to actually train a sensible model on. A lot of books don't have OCRed scans, or a digital version.

https://books.google.com/

Re: Google scrambles to manually remove weird AI answers in search

#305
post #175

1. Google announces something that has AI bolted on 2. A VP pontificates about how much work they did to "get it right" 3. An easy-to-anticipate first-order issue surfaces 4. Sundar issues a statement like "this is completely unacceptable. We will be making structural changes to ensure this never happens again."[0] 5. GOTO 1 [0] https://m.economictimes.com/tech/technology/sundar-pichai-ca...

This is what happens when senior leadership no longer even attempt to hide their contempt for the rank-and-file.

Re: Google scrambles to manually remove weird AI answers in search

#306
post #177

Earlier quoted context omitted.

that's the part that scares me. I railed on someone's comment the other day about "indexes will come back into fashion" but the more I think about how much garbage has increased in just the past 2 to 3 years, I think I was wrong. Indexes and forums may be the only way to have a sane net where you can find things. Perhaps communities linking together in a ring like format, a "web ring" of sorts.

What I've been wanting to see for a while now is a social-network based search engine: * No pages are indexed automatically. The only indexed pages are pages that users say are worth indexing. Probably have a browser add-on for a one button click that people can use. * You can friend/follow others * Your search results are a combination of your own indexed pages and the pages indexed by people in your network.

Isn’t that what Reddit is or digg was ? Link aggregators ?

Gaming that is solved problem , you can use human bot farms to brigade and astroturf and you can even motivate people to do it for free .

If cost of spamming is cheaper than cost of moderation, spam will win

Re: Google scrambles to manually remove weird AI answers in search

#307
post #175

1. Google announces something that has AI bolted on 2. A VP pontificates about how much work they did to "get it right" 3. An easy-to-anticipate first-order issue surfaces 4. Sundar issues a statement like "this is completely unacceptable. We will be making structural changes to ensure this never happens again."[0] 5. GOTO 1 [0] https://m.economictimes.com/tech/technology/sundar-pichai-ca...

This is what happens when senior leadership no longer even attempt to hide their contempt for the rank-and-file.

At that point, this form of contempt usually referred to as narcism.

Re: Google scrambles to manually remove weird AI answers in search

#309
post #93

Earlier quoted context omitted.

The problem is that for some searches and answers Reddit or other social media is fine.

But only if you do a lot of filtering when going through responses. It’s kind of simple to do as a human, we see a ridiculous joke answer or obvious astroturfing and move on, but Reddit is like >99% noise, with people upvoting obviously wrong answer because it’s funny, lots of bot content, constant astroturfing attempts.

The users of r/montreal are so sick of lazy tourists constantly asking the same dumb "what's the best XYZ" questions without doing a basic search fit, the meme answer is always "bain colonial" which is a men-only spa for cruising. Often the topmost voted comment. I just tried asking gemini and chatgpt what that response meant and neither caught on..

Re: Google scrambles to manually remove weird AI answers in search

#310

Earlier quoted context omitted.

Spend say £500M (USD/GBP/EUR) on experts, per annum. Imagine typing a search and getting a response: "Give us 30 mins to respond - here's a token, come back at 17:35 with your token" ... and then you get an answer from an expert, which also gets indexed. The clever bit decides when to defer to an expert instead of returning answers from the index. I'll leave the finer details out.

Google Answers was launched in 2002 and retired in 2006. https://en.wikipedia.org/wiki/Google_Answers

Quora still exists. :-)
Post reply on HN