Live data from Hacker News

Google scrambles to manually remove weird AI answers in search

theverge.com

41–50 of 387 posts

Re: Google scrambles to manually remove weird AI answers in search

#42
post #15

So Google hasn't used an LLM to generate and test weird queries ? This is not putting the bar very high for the whole industry... There'd be so much to gain from a clean deployment... Either it hard, either it is a rush. As a machine learnist, I believe it's actually impossible, by design of the autoregressive LLM. This race may we'll be partially to the bottom.

They still haven’t learned from the Gemini diverse Nazis debacle.

Re: Google scrambles to manually remove weird AI answers in search

#43
post #27

"Achieving the initial 80 percent is relatively straightforward since it involves approximating a large amount of human data, Marcus said, but the final 20 percent is extremely challenging. In fact, Marcus thinks that last 20 percent might be the hardest thing of all." 100% completely accurate is super-AI-complete. No human can meet that goal either. No, not even you, dear person reading this. You are wrong about som…

> putting glue on pizza to hold the cheese on

It's actually not the dumbest idea I've heard from a real person. So no surprise it might be suggested by an AI that was trained on data from real people.

Re: Google scrambles to manually remove weird AI answers in search

#44
post #15

So Google hasn't used an LLM to generate and test weird queries ? This is not putting the bar very high for the whole industry... There'd be so much to gain from a clean deployment... Either it hard, either it is a rush. As a machine learnist, I believe it's actually impossible, by design of the autoregressive LLM. This race may we'll be partially to the bottom.

Google’s poor testing is hardly in doubt. But keep in mind that the whole problem is that LLMs don’t handle “unlikely” text nearly as well as “likely” text. So the near-infinite space of goofy things to search on Google is basically like panning for gold in terms of AI errors (especially if they are using a cheap LLM).

And in particular LLMs are less likely to generate these goofy prompts because they wouldn’t be in the training data.

Re: Google scrambles to manually remove weird AI answers in search

#45
post #27

"Achieving the initial 80 percent is relatively straightforward since it involves approximating a large amount of human data, Marcus said, but the final 20 percent is extremely challenging. In fact, Marcus thinks that last 20 percent might be the hardest thing of all." 100% completely accurate is super-AI-complete. No human can meet that goal either. No, not even you, dear person reading this. You are wrong about som…

> So 100% accurate can't be the goal. Obviously the goal is to get the responses to be less obviously stupid.

I'm not sure I agree. I think you're right that 100% accuracy is potentially unfeasable as a realistic aim, but I think the question is how accurate something needs to be in order to be a useful proposition for search.

AI that's as knowledgable as I am is a good achievement and helpful for a lot of use cases, but if I'm searching "What's the capital of Mongilia" someone with averageish knowledge taking a punt with "Maybe Mongoliana City?" is not helpful at all- if I can't trust AI responses to a high degree, I'd much rather just have normal search results showing me other resources I can trust.

Google's bar for justifying adding AI to their search proposition isn't "be better than asking someone on the street", it's "be better than searching google without any AI results"

Re: Google scrambles to manually remove weird AI answers in search

#47
post #37

I'm actually shocked that a company that has spent 25 years on finetuning search results for any random question people ask in the searchbox does not have a good, clean, dataset to train an LLM on. Maybe this is the time to get out the old Encyclopedia Britannica CD and use that for training input.

I am also surprised that training data are not much more curated.

Encyclopedias, textbooks, reputable journals, newspapers and magazines make sense.

But to throw in social media? Reddit? Seems insane.

Re: Google scrambles to manually remove weird AI answers in search

#48
post #11

Who knows how many of these are fake. People have been dropping inspect-element-manipulated screenshots all over twitter. https://www.nytimes.com/2024/05/24/technology/google-ai-over... > A correction was made on May 24, 2024: An earlier version of this article referred incorrectly to a Google result from the company’s new artificial-intelligence tool AI Overview. A social media commenter claimed that a result for a…

[deleted]

Re: Google scrambles to manually remove weird AI answers in search

#49
Putting glue on the pizza is (apparently) a clever way to take pictures of slices of pizza that look "perfect" to the camera (not for eating, obviously) [1]. I remember a couple years ago some videos of "tricks" showing this, plus literally screwing the pizza with screws.

So, yeah, the ai did in fact autocompleted the question correctly. It was just the wrong context. Good luck trying to "fix" that.

[1] https://shotkit.com/food-photography-secrets-revealed/ (number 2)

Re: Google scrambles to manually remove weird AI answers in search

#50
post #27

"Achieving the initial 80 percent is relatively straightforward since it involves approximating a large amount of human data, Marcus said, but the final 20 percent is extremely challenging. In fact, Marcus thinks that last 20 percent might be the hardest thing of all." 100% completely accurate is super-AI-complete. No human can meet that goal either. No, not even you, dear person reading this. You are wrong about som…

Yes, which is why the ability to sift accurate and authoritative sources from spam, propaganda, and intentionally deceptive garbage, like advertising, and present those high-quality results to the user for review and consideration, is more important than any attempt to have an AI serve a single right answer. Google, unfortunately, abandoned this problem some time ago and is now left to serve up nonsense from the melange of low-quality noise they incentivized in pursuit of profits. If they had, instead, remained focused on the former problem, it’s actually conceivable to have an LLM work more successfully from this base of knowledge.
Post reply on HN