Live data from Hacker News

AI assistants misrepresent news content 45% of the time

bbc.co.uk

241–250 of 306 posts

Re: AI assistants misrepresent news content 45% of the time

#241

Earlier quoted context omitted.

Do we have any good research on how much less often larger, newer models will just make stuff up like this? As it is, it's pretty clear LLMs are categorically not a good idea for directly querying for information in any non-fiction-writing context. If you're using an LLM to research something that needs to be accurate, the LLM needs to be doing a tool call to a web search and only asked to summarize relevant facts fr…

A further problem is that Wikipedia is chock full of nonsense, with a large proportion of articles that were never fact checked by an expert, and many that were written to promote various biased points of view, inadvertently uncritically repeat claims from slanted sources, or mischaracterize claims made in good sources. Many if not most articles have poor choice of emphasis of subtopics, omit important basic topics,…

Wikipedia is pretty good for most topics. Anything even remotely political somewhere however, it isn't just bad, it is one of the worst sources out there. And therein lies the problem, its wildly different levels of quality depending on the topic.

Re: AI assistants misrepresent news content 45% of the time

#242
post #174

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

We live in a post-truth society. This means that, unfortunately, most of society has learned that it doesn't matter if what you're saying is true. All that matters is that the words that you speak cause you or your cause to gain power.

All the more reason to call out bullshit in real life

Value truth and honesty. Call out lies for what they are.

This is the way to get sanity both for ourselves and for society as a whole

Re: AI assistants misrepresent news content 45% of the time

#243
post #114

Earlier quoted context omitted.

Why would you want Gemini to do this instead of just going to a news site (or several news sites) and reading what the headlines they wrote?

I wanted to use the agentic powers of the model to dig for specific kinds of news, and use iterative search as well. I think when LLMs use tools correctly this kind of search is more powerful than simple web search. It also has better semantic capabilities, so in a way I wanted to make my own LLM powered news feed.

That's makes sense. Thanks for explaining!

Re: AI assistants misrepresent news content 45% of the time

#244

Earlier quoted context omitted.

[flagged]

The BBC is famous for platforming Farage and smearing Corbyn. The Guardian is at best centre-right. Next you’ll try to convince me that Starmer’s Labour is left wing or the Lock Ness monster is real.

If this is an honest take, you probably have to look to the right to see Mao's ghost. Maybe talk to other humans in real life, you might be shocked about your actual place on the political spectrum.

Re: AI assistants misrepresent news content 45% of the time

#245

Earlier quoted context omitted.

The report says that different media organizations dropped their robots.txt for the duration of the research to give LLMs access. I would expect this isn't the on-off switch they conceptualized, but I don't know enough about how different LLM providers handle news search and retrieval to say for sure.

Does it work like that though? How long does it take for AI bots to crawl sites and have the data added to the model currently being used? Am I wrong in thinking that it takes a lot longer for AI bot crawls to be available to the public than a typical search engine crawler?

Bots could be crawlers gathering data to periodically be used as raw training data or the requests could just be from a web search agent of some form like ChatGPT finding latest news stories on topic X for example. I don’t know if robots.txt can distinguish between the two types of bot request or whether LLM providers even adhere to either.

Re: AI assistants misrepresent news content 45% of the time

#246

Earlier quoted context omitted.

The BBC is famous for platforming Farage and smearing Corbyn. The Guardian is at best centre-right. Next you’ll try to convince me that Starmer’s Labour is left wing or the Lock Ness monster is real.

If this is an honest take, you probably have to look to the right to see Mao's ghost. Maybe talk to other humans in real life, you might be shocked about your actual place on the political spectrum.

I talk to people all the time. There are both communists and fascists in the UK on the two extremes.

However, that is currently not reflected in electoral politics or the media. The farthest left are currently the Greens, at best centre-left. On the right and far right there are Tories and Reform.

Re: AI assistants misrepresent news content 45% of the time

#247

Earlier quoted context omitted.

Does it work like that though? How long does it take for AI bots to crawl sites and have the data added to the model currently being used? Am I wrong in thinking that it takes a lot longer for AI bot crawls to be available to the public than a typical search engine crawler?

Bots could be crawlers gathering data to periodically be used as raw training data or the requests could just be from a web search agent of some form like ChatGPT finding latest news stories on topic X for example. I don’t know if robots.txt can distinguish between the two types of bot request or whether LLM providers even adhere to either.

Wow, Just reading the headline I had assumed they were giving the new article as a document, then asking it to summarize the the document given.

Re: AI assistants misrepresent news content 45% of the time

#248

It's important to bear this in mind whenever you find out that someone uses an LLM to summarize a meeting, email, or other communication you've held. That person is not really getting the message you were conveying.

It would be important to bear this in mind if it was true, but it's not. I do sales meetings all day every day, and I've tried different AI note takers that send a summary of the meeting afterwards. I skim them when they get dumped into my CRM and they're almost always quite accurate. And I can verify it, because I was in the meeting .

It makes me think that a lot of the folks commenting on this stuff haven't actually used the tooling.

Agreed, it's generally quite accurate. I find for hectic meetings, it can get some things wrong... But the notes are generally still higher quality than human generated notes.

Is it perfect? No. Is it good enough? IMO absolutely.

Similar to many other things, the key is that you don't just blindly trust it. Have the LLM take notes and summarize, and then _proofread_ them, just as you would if you were writing them yourself...

Re: AI assistants misrepresent news content 45% of the time

#249
post #174

I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.

We live in a post-truth society. This means that, unfortunately, most of society has learned that it doesn't matter if what you're saying is true. All that matters is that the words that you speak cause you or your cause to gain power.

It does feel like there would be a lot more skepticism about the technology if it had appeared a decade or two ago.

Re: AI assistants misrepresent news content 45% of the time

#250
post #38

If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…

Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore

Hell, human editors seem to misrepresent their journalists frequently enough that I'm left wondering if it's hyperbolic or not to guess if they misrepresent them 45% of the time, too.
Post reply on HN