Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

231–240 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#231

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

Anyone and everyone who has bought any stocks into the circular Ponzi pyramid has this knee jerk response to rationalise LLM failure modes.

They want to believe that statistical distribution of meaningless tokens is real cognition of machines and if not that, works flawlessly for most of the cases and if not flawlessly, is usable enough to be valued at trillions of dollars collectively.

Re: AI assistants misrepresent news content 45% of the time

#232

Earlier quoted context omitted.

The biggest problem with that citation isn't that the article has since been deleted. The biggest problem is that that particular Wikipedia article was never a good source in the first place. That seems to be the real challenge with AI for this use case. It has no real critical thinking skills, so it's not really competent to choose reliable sources. So instead we're lowering the bar to just asking that the sources a…

I think this is a real challenge for everyone. In many ways potentially we need a restart of a wikipedia like site to document all the valid and good sources. This would also hopefully include things like source bias and whether it's a primary/secondary/tertiary source.

I noticed that my local library has a new set of World Book. Maybe it's time to bring back traditional encyclopedias.

Re: AI assistants misrepresent news content 45% of the time

#234

Earlier quoted context omitted.

Relatedly, I wonder if we count misrepresenting a misleading news article such that it becomes more-accurate as misrepresenting news content…

Since the model doesn't get to observe the actual situation being reported on, such an improvement in accuracy would only be random chance and should not be rewarded.

Oh, agreed, they don’t get points for failing the task but accidentally reporting something more-correct in the process.

Re: AI assistants misrepresent news content 45% of the time

#235

Earlier quoted context omitted.

This is likely because of the knowledge cutoff. I have seen a few cases before of "hallucinations" that turned out to be things that did exist, but no longer do.

The fix for this is for the AI to double-check all links before providing them to the user. I frequently ask ChatGPT to double check that references actually exist when it gives me them. It should be built in!

Gemini will lie to me when I ask it to cite things, either pull up relevant sources or just hallucinate them.

IDK how you people go through that experience more than a handful of times before you get pissed off and stop using these tools. I've wasted so much time because of believable lies from these bots.

Sorry, not even lies, just bullshit. The model has no conception of truth so it can't even lie. Just outputs bullshit that happens to be true sometimes.

Re: AI assistants misrepresent news content 45% of the time

#236

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

I wonder how many of those evangelists have some dumb AI startup that'll implode once the hype dies down (or a are a software engineer who feels smart when he follows their lead). One thing that's been really off putting about the technology industry is how fake-it-till-you-make-it has become so pervasive.

> One thing that's been really off putting about the technology industry is how fake-it-till-you-make-it has become so pervasive.

It feels accidental, but it's definitely amusing that the models themselves are aping this ethos.

Re: AI assistants misrepresent news content 45% of the time

#237

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

They've conned themselves with the LLMs they use, and are desperate to keep the con going: "The LLMentalist Effect: how chat-based Large Language Models replicate the mechanisms of a psychic’s con"

https://softwarecrisis.dev/letters/llmentalist/

Re: AI assistants misrepresent news content 45% of the time

#238

I'm curious how many people have actually taken the time to compare AI summaries with sources they summarize. I did for a few and ... it was really bad. In my experience, they don't summarize at all, they do a random condensation.. not the same thing at all. In one instance I looked at the result was a key takeaway being the opposite of what it should have been. I don't trust them at all now.

They’re basically markov chain text generators with a relevance-tracking-and-correction step. It turns out this is like 100x more useful than the same thing without the correction step, but they don’t really escape what they are “at heart”, if you will.

The ways they fail are often surprising if your baseline is “these are thinking machines”. If your baseline is what I wrote above (say, because you read the “Attention Is All You Need” paper) none of it’s surprising.

Re: AI assistants misrepresent news content 45% of the time

#240
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore

https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect
Post reply on HN