Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

251–260 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#251
post #20

I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…

About 75% of the time I look at the Gemini answer, it's wrong. Maybe 80%. Sometimes it's a little wrong, like giving the correct answer for another product/item, or the times that a business is open wrong. There's a local business I took my wife to, Gemini told her it's open monday to friday, but it's open tuesday to saturday, so we showed up on a monday to see them closed. But sometimes it's insanely wrong making up dozens of wrong "facts". My wife started looked more carefully now. My boss will even say "Gemini says X so it's probably Y" these days.

Re: AI assistants misrepresent news content 45% of the time

#252

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

Is that just an LLM thing? I thought that as a society, we decided a long time ago that competence doesn't really matter. Why else would we be giving high school diplomas to people who can't read at a 5th grade level? Or offshore call center jobs to people who have poor English skills?

It's been a 50 year downward slope. We're in the harvest phase of that crop. All the people we raised to believe their incompetence was just as valid as other people's facts are now confidently running things because they think magical thinking works.

Re: AI assistants misrepresent news content 45% of the time

#253
post #136

I've switched almost entirely to AI news (basically research mode & give it 10 areas I'm interested in). It definitely has a issues in the detail, but if you're only skimming the result for headlines it's perfectly fine. e.g. Pakistan and Afghanistan are shooting at each other. I wouldn't trust it to understand the tribal nuances behind why, but the key fact is there. [One exception is economic indicators, especially…

If all you're interested in are the headlines then why not just read the headlines?

The aggregation is one feature, and the dedupe another. ie if you grab only headlines, how to avoid seeing the same or similar headline twice, given op wants to pull from 10 topics from a potentially large variety of sources.

Re: AI assistants misrepresent news content 45% of the time

#254

Earlier quoted context omitted.

The biggest problem with that citation isn't that the article has since been deleted. The biggest problem is that that particular Wikipedia article was never a good source in the first place. That seems to be the real challenge with AI for this use case. It has no real critical thinking skills, so it's not really competent to choose reliable sources. So instead we're lowering the bar to just asking that the sources a…

I think this is a real challenge for everyone. In many ways potentially we need a restart of a wikipedia like site to document all the valid and good sources. This would also hopefully include things like source bias and whether it's a primary/secondary/tertiary source.

An example of this.

I've seen a certain sensationalist news source write a story that went like this.

Site A: Bad thing is happening, cite: article Site B

* follow the source *

Site B: Bad thing is happening, cite different article on Site A

* follow the source *

Site A: Bad thing is happening, no citation.

I fear that's the current state of a large news bubble that many people subscribe to. And when these sensationalist stories start circulating there's a natural human tendency to exaggerate.

I don't think AI has any sort of real good defense to this sort of thing. 1 level of citation is already hard enough. Recognizing that it is citing the same source is hard enough.

There was another example from the Kagi news stuff which exemplified this. A whole article written which made 3 citations that were ultimately spawned from the same new briefing published by different outlets.

I've even seen an example of a national political leader who fell for the same sort of sensationalization. One who should have known better. They repeated what was later found to be a lie by a well-known liar but added that "I've seen the photos in a classified debriefing". IDK that it was necessarily even malicious, I think people are just really bad at separating credible from uncredible information and that it ultimately blends together as one thing (certainly doesn't help with ancient politicians).

Re: AI assistants misrepresent news content 45% of the time

#255
post #12

Kagi News has been pretty accurate. Source information is provided along with the summary and key details too. AI summarizes are good for getting a feel of if you want to read an article or not. Even with Kagi News I verify key facts myself.

Kagi News is basically a summary of news articles fed into the context. It's different from what the op is about, that is just asking an LLM with web access to query the news.

I hate saying people are holding it wrong but given just given how LLMs work, how did anyone expect that this would go right? Managing the LLM's context is the game. I feel like ChatGPT has done such a disservice for teaching users how to actually use these tools and what their failure modes are.

Re: AI assistants misrepresent news content 45% of the time

#257
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore

Years ago in college, we had a class where we analyzed science in the news for a few weeks compared to the publish research itself. I think it was a 100% misrepresentation rate comparing what a news article summarized about a paper verses what the paper itself said. We weren't going off of CNN or similar main news sites, but news websites aimed at specific types of news which were consistently better than the articles in mainstream news (whenever the underlying research was noteworthy enough to earn a mention on larger sites). Leaving out complete details or only reporting some of the findings weren't enough to count, as it was expected any news summary would reduce the total amount of information being provided about a published paper compared to reading the paper directly. The focus was on looking for summaries that were incorrect or which made claims which the original paper did not support.

Probably the most impactful "easy A" class I had in college.

Re: AI assistants misrepresent news content 45% of the time

#258

Earlier quoted context omitted.

I don’t trust it at all. I wanted to know if he would be able to explain its own results. Just because it was displaying sources and links made me trust it until I checked and was horrified. I wanted to know if it was old link that broke or changed but no apparently

You said: >...it finally told me that it generated « the most probable » urls for the topic in question based on the ones he knows exists. smrq is asking why you would believe that explanation. The LLM doesn't necessarily know why it's doing what it's doing, so that could be another hallucination. Your answer: > ...I wanted to know if it was old link that broke or changed but no apparently Leads me to believe that yo…

No I got the question, I said that I wanted to see what kind of explanation it would give me. Ofc it can hallucinate that explanation as well. The bottom line is I don’t trust it, and the source link are fake (and not broken or obsolete)

Re: AI assistants misrepresent news content 45% of the time

#259
post #136

I've switched almost entirely to AI news (basically research mode & give it 10 areas I'm interested in). It definitely has a issues in the detail, but if you're only skimming the result for headlines it's perfectly fine. e.g. Pakistan and Afghanistan are shooting at each other. I wouldn't trust it to understand the tribal nuances behind why, but the key fact is there. [One exception is economic indicators, especially…

If all you're interested in are the headlines then why not just read the headlines?

Especially in the case of aggregators, I like that they remove, or at least tone down, the sensationalism.

Re: AI assistants misrepresent news content 45% of the time

#260

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

They've conned themselves with the LLMs they use, and are desperate to keep the con going: "The LLMentalist Effect: how chat-based Large Language Models replicate the mechanisms of a psychic’s con" https://softwarecrisis.dev/letters/llmentalist/

I had a look at that and am not convinced

>people are convinced that language models, or specifically chat-based language models, are intelligent... But there isn’t any mechanism inherent in large language models (LLMs) that would seem to enable this...

and says it must be a con but then how come they pass most of the exams designed to test humans better than humans do?

And there are mechanisms like transformers that may do something like human intelligence.

Post reply on HN