Live data from Hacker News

Google scrambles to manually remove weird AI answers in search

theverge.com

61–70 of 387 posts

Re: Google scrambles to manually remove weird AI answers in search

#61
post #27

"Achieving the initial 80 percent is relatively straightforward since it involves approximating a large amount of human data, Marcus said, but the final 20 percent is extremely challenging. In fact, Marcus thinks that last 20 percent might be the hardest thing of all." 100% completely accurate is super-AI-complete. No human can meet that goal either. No, not even you, dear person reading this. You are wrong about som…

> So 100% accurate can't be the goal. Obviously the goal is to get the responses to be less obviously stupid. I'm not sure I agree. I think you're right that 100% accuracy is potentially unfeasable as a realistic aim, but I think the question is how accurate something needs to be in order to be a useful proposition for search. AI that's as knowledgable as I am is a good achievement and helpful for a lot of use cases,…

The problem is that in all the shared examples, Google ai search does not respond with a Maybe xyz, question mark? like you did. It always answers with high confidence and can't seem to navigate any gray area where there are multiple differing opinions or opposing source of truths.

Re: Google scrambles to manually remove weird AI answers in search

#63
It's debatable whether Google has truly lost the plot because of the "AI wars", but the moment the statement "Bing returns more sensible results than you" becomes verifiably true, it's... cause for concern?

The approach that Google appears to have taken, which is to assume that the top-ranked part of its current search index is a sensible knowledge base, may have been true some years ago, but definitely isn't now: for whatever reasons, it's now 33% spam, 33% clickbait/propaganda, with the rest being equally divided between what could be called "truths" and miscellaneous detritus.

To me, it seems that returning to the concept that search results should at least reflect a broad consensus of what is true is a necessary first step for Google. As part of that, learning to flag obvious trolling, clickbait and bad-faith content is paramount. And then, maybe then, they can start touting their LLM benefits. But until the realities of the Internet are taken into account (i.e.: it's 80% spam!), any "we offer automated answers!" play is doomed.

Re: Google scrambles to manually remove weird AI answers in search

#66
post #15

So Google hasn't used an LLM to generate and test weird queries ? This is not putting the bar very high for the whole industry... There'd be so much to gain from a clean deployment... Either it hard, either it is a rush. As a machine learnist, I believe it's actually impossible, by design of the autoregressive LLM. This race may we'll be partially to the bottom.

> So Google hasn't used an LLM to generate and test weird queries ? You don't even need an LLM for that. Google will almost certainly have tested. The test result is just politically-unacceptable within the company: It doesn't work, it's a architectural issue inherent to the technology, we can't fix it. Instead, they just rush to patch any specific, individual errors that show up, and claim that these errors are "rar…

Deploy the cheap offshore labor!

Re: Google scrambles to manually remove weird AI answers in search

#67
post #27

"Achieving the initial 80 percent is relatively straightforward since it involves approximating a large amount of human data, Marcus said, but the final 20 percent is extremely challenging. In fact, Marcus thinks that last 20 percent might be the hardest thing of all." 100% completely accurate is super-AI-complete. No human can meet that goal either. No, not even you, dear person reading this. You are wrong about som…

[dead]

Re: Google scrambles to manually remove weird AI answers in search

#68
post #37

I'm actually shocked that a company that has spent 25 years on finetuning search results for any random question people ask in the searchbox does not have a good, clean, dataset to train an LLM on. Maybe this is the time to get out the old Encyclopedia Britannica CD and use that for training input.

I am also surprised that training data are not much more curated. Encyclopedias, textbooks, reputable journals, newspapers and magazines make sense. But to throw in social media? Reddit? Seems insane.

The problem is that for some searches and answers Reddit or other social media is fine.

Re: Google scrambles to manually remove weird AI answers in search

#69

Putting glue on the pizza is (apparently) a clever way to take pictures of slices of pizza that look "perfect" to the camera (not for eating, obviously) [1]. I remember a couple years ago some videos of "tricks" showing this, plus literally screwing the pizza with screws. So, yeah, the ai did in fact autocompleted the question correctly. It was just the wrong context. Good luck trying to "fix" that. [1] https://shotk…

"correctly but wrong" is just wrong.... there are no points scored for "in a very specific context it would've made sense"

Re: Google scrambles to manually remove weird AI answers in search

#70
post #37

I'm actually shocked that a company that has spent 25 years on finetuning search results for any random question people ask in the searchbox does not have a good, clean, dataset to train an LLM on. Maybe this is the time to get out the old Encyclopedia Britannica CD and use that for training input.

The problem in this case is not that it was trained on bad data. The AI summaries are just that - summaries - and there are bad results that it faithfully summarizes.

This is an attempt to reduce hallucinations coming full circle. A simple summarization model was meant to reduce hallucination risk, but now it's not discerning enough to exclude untruthful results from the summary.

Post reply on HN