Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

271–280 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#271
post #114

Earlier quoted context omitted.

Why would you want Gemini to do this instead of just going to a news site (or several news sites) and reading what the headlines they wrote?

I wanted to use the agentic powers of the model to dig for specific kinds of news, and use iterative search as well. I think when LLMs use tools correctly this kind of search is more powerful than simple web search. It also has better semantic capabilities, so in a way I wanted to make my own LLM powered news feed.

> I wanted to use the agentic powers of the model

Do you have an in-depth understanding of how those "agentic powers" are implemented? If not, you should probably research it yourself. Understanding what's underneath the buzzwords will save you some disappointment in the future.

Re: AI assistants misrepresent news content 45% of the time

#272

I'm curious how many people have actually taken the time to compare AI summaries with sources they summarize. I did for a few and ... it was really bad. In my experience, they don't summarize at all, they do a random condensation.. not the same thing at all. In one instance I looked at the result was a key takeaway being the opposite of what it should have been. I don't trust them at all now.

I have just tried doing this. I thought I could take all the release notes for my project over the past year and AI could give a great summary of all the work that had been done, categorize it and organize it. Seems like a good application for AI.

Result was just trash. It would do exactly as you say: condense the information, but there was no semblance of "summary". It would just choose random phrases or keywords from the release notes and string them together, but it had no meaning or clarity, it just seemed garbled.

And it's not for lack of trying; I tried to get a suitable result out of the AI well past the amount of time it would have taken me to summarize it myself.

The more I use these tools the more I feel their best use case is still advanced autocomplete.

Re: AI assistants misrepresent news content 45% of the time

#273
post #248

Earlier quoted context omitted.

It would be important to bear this in mind if it was true, but it's not. I do sales meetings all day every day, and I've tried different AI note takers that send a summary of the meeting afterwards. I skim them when they get dumped into my CRM and they're almost always quite accurate. And I can verify it, because I was in the meeting .

It makes me think that a lot of the folks commenting on this stuff haven't actually used the tooling. Agreed, it's generally quite accurate. I find for hectic meetings, it can get some things wrong... But the notes are generally still higher quality than human generated notes. Is it perfect? No. Is it good enough? IMO absolutely. Similar to many other things, the key is that you don't just blindly trust it. Have the…

I think the cost of inaccuracy is very a important factor in if it works for a specific use case. Meeting notes probably don't have much cost of inaccuracy. Medical records on the other hand...

Re: AI assistants misrepresent news content 45% of the time

#274
post #114
post #20

I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…

Why would you want Gemini to do this instead of just going to a news site (or several news sites) and reading what the headlines they wrote?

They're selling it as having this ability, so it really doesn't matter what people want. We should be holding these companies to account for selling software that doesn't live up to what they say it does.

Re: AI assistants misrepresent news content 45% of the time

#275

Earlier quoted context omitted.

Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore

Years ago in college, we had a class where we analyzed science in the news for a few weeks compared to the publish research itself. I think it was a 100% misrepresentation rate comparing what a news article summarized about a paper verses what the paper itself said. We weren't going off of CNN or similar main news sites, but news websites aimed at specific types of news which were consistently better than the article…

The Science News Cycle: https://phdcomics.com/comics.php?f=1174

Re: AI assistants misrepresent news content 45% of the time

#276
I have a gut feeling sycophancy would become a huge problem if I were ever to ask any AI assistant with even a vague idea of my political opinions to start summarizing news stories. If AIs twist other things around to give glowing responses to their users I'm almost certain they'll resort to giving a "spin" to news stories they think is in line with what the user wants to hear. Everyone will get a bespoke biased cable news station in the future!

Re: AI assistants misrepresent news content 45% of the time

#277
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

I wouldn't even say BBC is a good source to cite. For foreign news, BBC is outright biased. Though I don't have any good suggestions for what an LLM should cite instead.

You're downvoted but quite accurate. I would like to see this statistic compared against how often the BBC misrepresents news content, and the backflips that come with defining such a metric.

Re: AI assistants misrepresent news content 45% of the time

#278
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

Yes, but the problems with processing human writing are huge, so even if this article is bad something like the problem they claim exists is very real. LLMs misunderstanding individual sentences, losing track of who said what etc. happen in best models, including GPT-5 when they're asked to analyze normal human-written discussions like those we have here.

Much of this is probably solvable, but it very much not solved.

Re: AI assistants misrepresent news content 45% of the time

#279
post #33

Earlier quoted context omitted.

> Now let's run this experiment against the editorial boards in newsrooms. Or against people in general. It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure. Is misrepresenting news content 45% of the time better or worse than the average person? I don't know. By extension: Would a person using an AI assistant misrepresent news more or less…

> It's a pet peeve of mine that we get these kinds of articles without a baseline established of how people do on the same measure I don’t have a personal human news summarizer? The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source. > cynic in me want another question answered too: How often does reporters misrep…

> I don’t have a personal human news summarizer?

Is this not the editorial board and journalist? I'm not sure what the gripe is here.

Re: AI assistants misrepresent news content 45% of the time

#280
post #160

Earlier quoted context omitted.

> I don’t have a personal human news summarizer? Not a personal one. You do however have reporters sitting between you and the source material a lot of the time, and sometimes multiple levels of reporters playing games of telephone with the source material. > The comparison is between a human reading the primary source against the same human reading an LLM hallucination mixed with an LLM referring the primary source.…

> You do however have reporters sitting between you and the source material a lot of the time In cases where a reporter is just summarising e.g. a court case, sure. Stock market news has been automated since the 2000s. More broadly, AI assistants misrepresenting news content may sometimes direct reference a court case. But they often don't. Even if they only could, that covers a small fraction of the news, much of wh…

> Reporter error is somewhat orthogonal to AI assistants' accuracy.

It is not at all. Journalists are wrong all the time, but you still treat news like record and not a sample. In fact I'd put money that AI mischaracterizes events at a LOWER rate than AI does: narratives shift over time, and journalists are more likely to succumb to this shift.

Post reply on HN