I'm curious how many people have actually taken the time to compare AI summaries with sources they summarize. I did for a few and ... it was really bad. In my experience, they don't summarize at all, they do a random condensation.. not the same thing at all. In one instance I looked at the result was a key takeaway being the opposite of what it should have been. I don't trust them at all now.
AI assistants misrepresent news content 45% of the time
281–290 of 306 posts
Re: AI assistants misrepresent news content 45% of the time
#282I'm curious how many people have actually taken the time to compare AI summaries with sources they summarize. I did for a few and ... it was really bad. In my experience, they don't summarize at all, they do a random condensation.. not the same thing at all. In one instance I looked at the result was a key takeaway being the opposite of what it should have been. I don't trust them at all now.
I wonder that's because a lot of news titles are clickbait. If they hallucinate the summary based on what the title may suggest, no wonder they misunderstand half of news articles.
Re: AI assistants misrepresent news content 45% of the time
#283Re: AI assistants misrepresent news content 45% of the time
#284Earlier quoted context omitted.
Human journalists misrepresent the white paper 85% of the time. With this in mind, 45% doesn't seem so bad anymore
Hell, human editors seem to misrepresent their journalists frequently enough that I'm left wondering if it's hyperbolic or not to guess if they misrepresent them 45% of the time, too.
Maybe we complained with enough concrete examples of how absolute shit editors and summarizers are now.
Re: AI assistants misrepresent news content 45% of the time
#285Hallucination Leaderboard "This evaluates how often an LLM introduces hallucinations when summarizing a document." https://github.com/vectara/hallucination-leaderboard If the figures on this leaderboard are to be trusted, many frontier and near-frontier models are already better than the median white-collar worker in this aspect. Note: The leaderboard doesn't cover tool calling, to be clear.
I’ve been reviewing academic papers for decades, and I’ve reviewed thousands of them. I’ve never seen a fake citation. I’ve seen misrepresented sources and cooked data, but never a straight-up fake citation. So the min max and median are at 0.
Note that people who write academic papers are quite far from the median white-collar worker.
Re: AI assistants misrepresent news content 45% of the time
#286Earlier quoted context omitted.
Wikipedia is pretty good for most topics. Anything even remotely political somewhere however, it isn't just bad, it is one of the worst sources out there. And therein lies the problem, its wildly different levels of quality depending on the topic.
Wikipedia is bad even for topics that aren't particularly political, not even because the editor was trying to be misleading but rather was being lazy and wrote up their own misconception and either made up a source or pulled a source without bothering to actually read it. These kind of errors can stay in place for years . I have one example that I check periodically just to see if anybody else has noticed. I've been…
Imagine if this was the ethos regarding open source software projects. Imaging Microsoft saying 20 years ago, "Linux has this and that bug, but you're not allowed to go fix it because that detracts from our criticism of open source." (Actually, I wouldn't be surprised if Microsoft or similar detractors literally said this.)
Of course Wikipedia has wrong information. Most open source software projects, even the best, have buggy, shite code. But these things are better understood not as products, but as processes, and in many (but not all) contexts the product at any point in time has generally proven, in a broad sense, to outperform their cathedral alternatives. But the process breaks down when pervasive cynicism and nihilism reduce the number of well-intentioned people who positively engage and contribute, rather than complain from the sidelines. Then we land right back to square 0. And maybe you're too young to remember what the world was like at square 0, but it sucked in terms of knowledge accessibility, notwithstanding the small number of outstanding resources--but which were often inaccessible because of cost or other barriers.
Re: AI assistants misrepresent news content 45% of the time
#287I recently tried to get Gemini to collect fresh news and show them to me, and instead of using search it hallucinated everything wholesale, titles, abstracts and links. Not just once, multiple times. I am kind of afraid of using Gemini now for anything related to web search. Here is a sample: > [1] Google DeepMind and Harvard researchers propose a new method for testing the ‘theory of mind’ of LLMs - Researchers have…
Re: AI assistants misrepresent news content 45% of the time
#288Earlier quoted context omitted.
> You do however have reporters sitting between you and the source material a lot of the time In cases where a reporter is just summarising e.g. a court case, sure. Stock market news has been automated since the 2000s. More broadly, AI assistants misrepresenting news content may sometimes direct reference a court case. But they often don't. Even if they only could, that covers a small fraction of the news, much of wh…
> Reporter error is somewhat orthogonal to AI assistants' accuracy. It is not at all. Journalists are wrong all the time, but you still treat news like record and not a sample. In fact I'd put money that AI mischaracterizes events at a LOWER rate than AI does: narratives shift over time, and journalists are more likely to succumb to this shift.
Straw man. Everyone educated constantly argues over sourcing.
> I'd put money that AI mischaracterizes events at a LOWER rate than AI does
Maybe it does. But an AI sourcing journalists is demonstrably worse. Source: TFA.
> narratives shift over time, and journalists are more likely to succumb to this shift
Lol, we’ve already forgotten about MechaHitler.
At the end of the day, a lot of people consume news to be entertained. They’re better served by AI. The risk is folks of consequence start doing that, at which point I suppose the system self resolves by making them, in the long run, of no consequence compared to those who own and control the AI.
Re: AI assistants misrepresent news content 45% of the time
#289Earlier quoted context omitted.
The biggest problem with that citation isn't that the article has since been deleted. The biggest problem is that that particular Wikipedia article was never a good source in the first place. That seems to be the real challenge with AI for this use case. It has no real critical thinking skills, so it's not really competent to choose reliable sources. So instead we're lowering the bar to just asking that the sources a…
I get what your saying. But you are now asking for a level of intelligence and critical thinking that I honestly believe is higher than the average person. I think its absolutely doable, but I also feel like we shouldn't make it sound like the current behavior is abhorrent or somehow indicative of a failure in the technology.
These grifters simply were not attracted to these gigs in these quantities prior to AI, but now the market incentives have changed. Should we "blame" the technology for its abuse? I think AI is incredible, but market endorsement is different from intellectual admiration.
Re: AI assistants misrepresent news content 45% of the time
#290Earlier quoted context omitted.
I thought people here hated it when LLMs made http requests?
It's bad when they indiscriminately crawl for training, and not ideal (but understandable) to use the Internet to communicate with them (and having online accounts associated with that etc.) rather than running them locally. It's not bad when they use the Internet at generation time to verify the output.