Live data from Hacker News

Microsoft's AI shopping announcement contains hallucinations in the demo

perfectrec.com

51–60 of 108 posts

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#52
post #27

Is it just me or does everyone trust AI opinions less and less ? Every time I ask it to find top 5 of something, I go and double check myself and almost always find it to be wrong. For example try searching for top 5 restaurants around me in bard. Some of them dont even exist lol and some are just random if you cross verify with actual popularity from yelp etc.

Well it doesn’t surprise me since I have been saying this for a while that these LLMs hallucinate nonsense to the point where you end up triple checking whatever it outputs. LLMs thrive in applications that involve creativity and non-serious applications mostly around fantasy or creative writing. Anyone using them seriously outside of summarization for high risk use cases is going to be very disappointed.

Perhaps the outcome is we get better at actually checking things, not a terrible result.

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#53

Is it just me or does everyone trust AI opinions less and less ? Every time I ask it to find top 5 of something, I go and double check myself and almost always find it to be wrong. For example try searching for top 5 restaurants around me in bard. Some of them dont even exist lol and some are just random if you cross verify with actual popularity from yelp etc.

My trust factor for online opinion is ranked:

1) Online forums (adding 'reddit' or 'hacker news' to a search query) 2) GPT4 3) Google search

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#54

Earlier quoted context omitted.

> If we're going to anthropomorphize AIs, let's just call it bullshitting and lies. why? "bullshitting and lies" suggests that the AI is intentionally being deceptive. "hallucinations" conveys the idea that the information is incorrect, but the AI perceives it to be correct, which is more in line with what is actually happening.

https://en.m.wikipedia.org/wiki/Confabulation

That seems more technically accurate but less likely to catch on since it's a much less common word. I think my parents are a lot more likely to understand from context a news story that mentions that "lawyers relied on an AI that hallucinated court cases" than "lawyers relied on an AI that confabulated court cases".

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#55
post #34
post #26

Earlier quoted context omitted.

Using language models for location or time based things is not recommended, as this usually requires non-textual data. Better to use them for general knowledge questions, programming help, translation, or writing. Asking them to do any complex calculations (especially when they also require non-text raw data, like inflation in a given time period) is also futile.

> general knowledge questions, programming help, translation, or writing. They get all of these wrong too. It's like some AI-specific variant of the Gell-man amnesia effect. It's usually right in the first sentence, but if you really know the answer, it's often either very debatable or completely wrong by the halfway mark of the paragraph. Meanwhile, the associated brand authority is problematic.

Gell-man is exactly what I’ve been referencing in conversation recently. Any professional will gladly explain why their field is really much too nuanced and complex for LLMs to threaten in the near term before seamlessly explaining how close we are to all those other engineers/doctors/clerks being automated right away.

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#57

Earlier quoted context omitted.

Given the euphemism "bug" substituting for "programming error" you'd be tempted to allow something similar for LLMs, but these are not errors, the output is by design. There is no motive for truth, just the most likely output, even if the likeliness is low.

> There is no motive for truth This also ignores the larger question that has been a known issue for at least 2,000 years: "Quid est veritas?"

You can substitute "accuracy" or "usefulness" for "truth".

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#59

Pretty soon some LLM owner is going to use the argument "Everyone is allowed to have their own opinions, and LLMs are too, their responses don't have to line up with someone else's preferences."

Alternative Intelligence

Re: Microsoft's AI shopping announcement contains hallucinations in the demo

#60

Is it just me or does everyone trust AI opinions less and less ? Every time I ask it to find top 5 of something, I go and double check myself and almost always find it to be wrong. For example try searching for top 5 restaurants around me in bard. Some of them dont even exist lol and some are just random if you cross verify with actual popularity from yelp etc.

I think the most amusing comment I've read here in the last few weeks called it "demented Clippy".
Post reply on HN