Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

81–90 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#81
post #20

I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…

They can be good for search, but you must click through the provided links and verify that they actually say what it says they do.

They can be good for search, but you must click through the provided links and verify that they actually say what it says they do.

Then they're not very good at search.

It's like saying the proverbial million monkeys at typewriters are good at search because eventually they type something right.

Re: AI assistants misrepresent news content 45% of the time

#82

Earlier quoted context omitted.

> For the current research, a set of 30 “core” news questions was developed Right. Let's talk about statistics for a bit. Or let's put it differently: they found in their report that 45% of the answers for 30 questions they have "developed" had a significant issue, e.g. inexisting reference I'll give you 30 questions out of my sleeve where 95% of the answers will not have any significant issue.

Yes, I'm sure you could hack together some bullshit questions to demonstrate whatever you want. Is there a specific reason that the reasonably straightforward methodology they did use is somehow flawed?

Yes, and you answered it yourself.

Re: AI assistants misrepresent news content 45% of the time

#83
post #47
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

> Participating organizations raised concerns about responses that relied heavily or solely on Wikipedia content – Radio-Canada calculated that of 108 sources cited in responses from ChatGPT, 58% were from Wikipedia. CBC-Radio-Canada are amongst a number of Canadian media organisations suing ChatGPT’s creator, OpenAI, for copyright infringement. Although the impact of this on ChatGPT’s approach to sourcing is not exp…

[deleted]

Re: AI assistants misrepresent news content 45% of the time

#85
post #74

Actual news articles misrepresent reality more often than 45%. Some very recent discussions on HN: https://news.ycombinator.com/item?id=45617088 https://news.ycombinator.com/item?id=45585323

How is that possible if the AI models rely on and implicitly trust these sources?

The article is about how well AI models misrepresent the content of news, not how often they misrepresent reality. My point is that even if the AI models make no errors when representing news content, they'll still be quite inaccurate when reality is the benchmark.

Who cares if AI does a good job representing the source, when the source is crap?

Re: AI assistants misrepresent news content 45% of the time

#87
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

Do we have any good research on how much less often larger, newer models will just make stuff up like this? As it is, it's pretty clear LLMs are categorically not a good idea for directly querying for information in any non-fiction-writing context. If you're using an LLM to research something that needs to be accurate, the LLM needs to be doing a tool call to a web search and only asked to summarize relevant facts fr…

A further problem is that Wikipedia is chock full of nonsense, with a large proportion of articles that were never fact checked by an expert, and many that were written to promote various biased points of view, inadvertently uncritically repeat claims from slanted sources, or mischaracterize claims made in good sources. Many if not most articles have poor choice of emphasis of subtopics, omit important basic topics, and make routine factual errors. (This problem is not unique to Wikipedia by any means, and despite its flaws Wikipedia is an amazing achievement.)

A critical human reader can go as deep as they like in examining claims there: can look at the source listed for a claim, can often click through to read the claim in the source, can examine the talk page and article history, can search through the research literature trying to figure out where the claim came from or how it mutated in passing from source to source, etc. But an AI "reader" is a predictive statistical model, not a critical consumer of information.

Re: AI assistants misrepresent news content 45% of the time

#88

Hallucination Leaderboard "This evaluates how often an LLM introduces hallucinations when summarizing a document." https://github.com/vectara/hallucination-leaderboard If the figures on this leaderboard are to be trusted, many frontier and near-frontier models are already better than the median white-collar worker in this aspect. Note: The leaderboard doesn't cover tool calling, to be clear.

I’ve been reviewing academic papers for decades, and I’ve reviewed thousands of them. I’ve never seen a fake citation. I’ve seen misrepresented sources and cooked data, but never a straight-up fake citation.

So the min max and median are at 0.

Re: AI assistants misrepresent news content 45% of the time

#89
post #39

Page 10 onwards of this PDF shows concrete examples of the mistakes: https://www.bbc.co.uk/aboutthebbc/documents/news-integrity-i... > ChatGPT / CBC / Is Türkiye in the EU? > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

It did exist but got removed: https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...

Quite an omission to not even check for that and it make me think that was done intentionally.

Re: AI assistants misrepresent news content 45% of the time

#90
post #47
post #42

Earlier quoted context omitted.

> or it (shocking) cites Wikipedia instead of the BBC. No... the problem is that it cites Wikipedia articles that don't exist . > ChatGPT linked to a non-existent Wikipedia article on the “European Union Enlargement Goals for 2040”. In fact, there is no official EU policy under that name. The response hallucinates a URL but also, indirectly, an EU goal and policy.

> Participating organizations raised concerns about responses that relied heavily or solely on Wikipedia content – Radio-Canada calculated that of 108 sources cited in responses from ChatGPT, 58% were from Wikipedia. CBC-Radio-Canada are amongst a number of Canadian media organisations suing ChatGPT’s creator, OpenAI, for copyright infringement. Although the impact of this on ChatGPT’s approach to sourcing is not exp…

Literally constantly? It takes both careful prompting and throughout double-checking to really notice however. Because often the links also exist, just don't represent what the LLM made it sound like.

And the worst part about the people unironically thinking they can use it for "research" is, that it essentially supercharges confirmation bias.

The inefficient sidequests you do while researching is generally what actually gives you the ability to really reason about a topic.

If you instead just laser focus on the tidbits you prompted with... Well, your opinion is a lot less grounded.

Post reply on HN