Earlier quoted context omitted.
I wanted to use the agentic powers of the model to dig for specific kinds of news, and use iterative search as well. I think when LLMs use tools correctly this kind of search is more powerful than simple web search. It also has better semantic capabilities, so in a way I wanted to make my own LLM powered news feed.
> I wanted to use the agentic powers of the model Do you have an in-depth understanding of how those "agentic powers" are implemented? If not, you should probably research it yourself. Understanding what's underneath the buzzwords will save you some disappointment in the future.
AI assistants misrepresent news content 45% of the time
291–300 of 306 posts
Re: AI assistants misrepresent news content 45% of the time
#292Re: AI assistants misrepresent news content 45% of the time
#293Earlier quoted context omitted.
Err, no? Being _possible_ does not necessarily imply that's what happened.
A bucket of 30 questions is not a statistically significant sample size which we can use to support the hypothesis which goes to say that all AI assistants they tested are 45% of the time wrong. That's not how science works. Neither is my bucket of 30 questions statistcally significant but it goes to say that I can disprove their hypothesis just by giving them my sample. I think that the report is being disingenious…
Re: AI assistants misrepresent news content 45% of the time
#294I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…
Re: AI assistants misrepresent news content 45% of the time
#295I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.
Anyone and everyone who has bought any stocks into the circular Ponzi pyramid has this knee jerk response to rationalise LLM failure modes. They want to believe that statistical distribution of meaningless tokens is real cognition of machines and if not that, works flawlessly for most of the cases and if not flawlessly, is usable enough to be valued at trillions of dollars collectively.
statistical distribution of meaningless tokens As a aside note the biggest argument for the possibility of machine consciousness is the depressing fact that so many humans are uncritical bullshit spreaders themselves.
Re: AI assistants misrepresent news content 45% of the time
#296I am curious if LLMs evangelists understand how off-putting it is when they knee-jerk rationalize how badly these tools are performing. It makes it seem like it isn't about technological capabilities: it is about a religious belief that "competence" is too much to ask of either them or their software tools.
We live in a post-truth society. This means that, unfortunately, most of society has learned that it doesn't matter if what you're saying is true. All that matters is that the words that you speak cause you or your cause to gain power.
I urge everyone to read Harry Frankfurt's short essay On Bullshit: https://www2.csudh.edu/ccauthen/576f12/frankfurt__harry_-_on...
Re: AI assistants misrepresent news content 45% of the time
#297Re: AI assistants misrepresent news content 45% of the time
#298If you dig into the actual report (I know, I know, how passe), you see how they get the numbers. Most of the errors are "sourcing issues": the AI assistant doesn't cite a claim, or it (shocking) cites Wikipedia instead of the BBC. Other issues: the report doesn't even say which particular models it's querying [ETA: discovered they do list this in an appendix], aside from saying it's the consumer tier. And it leaves o…
> 45% of responses contained at least one meaningful error. Sourcing [...] is 31%, followed by accuracy 20%
And you can see the reason they think this is important on the second page just after the summary.
> More than 1 in 3 (35%) of UK adults instinctively agree the news source should be held responsible for errors in AI-generated news
So of course the BBC cares that Googles summary said that the BBC cites pornhub when talking about domestic abuse (when they didn't), because a large portion of people blame them for the fact that a significant amount of AI generated crap is wrong.
Re: AI assistants misrepresent news content 45% of the time
#299Earlier quoted context omitted.
Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore
Years ago in college, we had a class where we analyzed science in the news for a few weeks compared to the publish research itself. I think it was a 100% misrepresentation rate comparing what a news article summarized about a paper verses what the paper itself said. We weren't going off of CNN or similar main news sites, but news websites aimed at specific types of news which were consistently better than the article…
Re: AI assistants misrepresent news content 45% of the time
#300In other words they are more factual than the bbc
LLMs aren't doing journalism on their own, whatever mistakes they make are compounded on top of any mistakes that the actual sources (such as the BBC) might have made.