Live data from Hacker News

AI advice made people less accurate but more confident – sudy

thenextweb.com

171–180 of 234 posts

Re: AI advice made people less accurate but more confident – sudy

#171
post #164

Earlier quoted context omitted.

Here is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational…

> So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. They are not. But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.

It’s not about the LLM being modern or not. 3.5 Flash is fairly new, but it’s also a flash model. It’s not designed to be knowledgeable.

People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.

Re: AI advice made people less accurate but more confident – sudy

#172

Earlier quoted context omitted.

Another human being took time from their life to try and help you It wasn't as much time as you felt entitled to from them, so you responded by minimizing their contribution and being condescending. That upset them. The problem here was not the other person.

When it comes to the specific case of people just mindlessly copy/pasting LLM’s: No, I asked them to help and they offloaded it without applying any of the expertise I specifically asked for. If they don’t want to give their opinion or their expertise they should say “no.” It’s a valid response! If I’m asking for you to give your time and you don’t have time to give, that’s fine! But don’t just throw my question into…

> I too have access to LLM’s.

You should have stated what you already tried then. You didn't, so they did the obvious thing for you... to do you a favor. If you'd tried something before asking them surely you'd have told them that at the start, right?

> as evidenced by all the slop dumping

I see more evidence of lazy questions. A good question includes what you've already done to help understand the issue (often refereed to as the proof of work). If you don't do that the other party must assume you are either don't now how, or didn't have time to. If they are kind they do the obvious thing for you. If they are not kind, they just ignore your lazy question.

I mostly ignore the lazy ones these days. Tell me what you tried or I assume you're just trying to get me to do the reading so you don't have to.

You should value the people who take the time to answer your inadequate question at all, rather than bemoaning them for not doing ENOUGH free labor for you when you weren't even willing to type a few lines about what you'd already tried.

Now if you say "I've already Googled it and asked all the AI's, but can't find anything" and they still dump some slop on you then there's a problem.

Re: AI advice made people less accurate but more confident – sudy

#173

Earlier quoted context omitted.

If you ask that, you fundamentally misunderstand the point. It's not about the LLM, it's about whether people will critically evaluate what it spits out.

Here is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational…

Interestingly enough, Kimi K2.6 said that it didn't know what car Monica drove.

>If you want to know this specific detail you might have to watch the movie yourself.

GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure.

GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer.

One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it.

Re: AI advice made people less accurate but more confident – sudy

#174

This study is pretty bad. The comment ( https://news.ycombinator.com/item?id=48970182 ) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a giv…

> This study is pretty bad.

The study is OK. The article (and the original headline that came with it) is pretty bad because it claims things that the study doesn't. And I guess it is ironic that the TNW article looks 100% AI-generated.

Re: AI advice made people less accurate but more confident – sudy

#175

Earlier quoted context omitted.

Exactly, it’s unrepresentative of AI. It’s damaged AI.

It seems perfectly representative of AI and AI users.

Why didn't they use a model people actually use?

Re: AI advice made people less accurate but more confident – sudy

#176

Earlier quoted context omitted.

When it comes to the specific case of people just mindlessly copy/pasting LLM’s: No, I asked them to help and they offloaded it without applying any of the expertise I specifically asked for. If they don’t want to give their opinion or their expertise they should say “no.” It’s a valid response! If I’m asking for you to give your time and you don’t have time to give, that’s fine! But don’t just throw my question into…

> I too have access to LLM’s. You should have stated what you already tried then. You didn't, so they did the obvious thing for you... to do you a favor. If you'd tried something before asking them surely you'd have told them that at the start, right? > as evidenced by all the slop dumping I see more evidence of lazy questions. A good question includes what you've already done to help understand the issue (often refe…

> You should have stated what you already tried then. You didn't, so they did the obvious thing for you... to do you a favor. If you'd tried something before asking them surely you'd have told them that at the start, right?

I’m not sure what you mean here, might be a misunderstanding. My point is don’t use your LLM and simply paste it. I have an LLM, I can use it too. We all have access to them. We should all operate under that assumption. Also, it’s not like I gave people a detailed list of everything I tried before asking them prior to LLM’s. Generally we should assume the person is asking us for a reason, we don’t need to run an audit here.

At the the day my point is I am asking you, not chat gpt. Feel free to use it just like I’d use a search to make sure I’m giving an accurate, comprehensive answer sometimes. But just like I don’t simply dump links on people when they ask for help, don’t dump your unvetted outputs on people who ask you for help.

Re: AI advice made people less accurate but more confident – sudy

#177

Earlier quoted context omitted.

Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study…

If you ask that, you fundamentally misunderstand the point. It's not about the LLM, it's about whether people will critically evaluate what it spits out.

How does it test that at all? Did the quiz have answers that people could figure out better by scrutinizing the llm?

Re: AI advice made people less accurate but more confident – sudy

#178

This study is pretty bad. The comment ( https://news.ycombinator.com/item?id=48970182 ) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a giv…

>This is akin to giving someone a textbook

Except LLM isn't a textbook, people know that but believe it nonetheless.

Re: AI advice made people less accurate but more confident – sudy

#179

Advice and information subreddits have gone to shit because of AI usage. A large number of people seem to think that when someone asks a question, what they really want is not someone with direct knowledge, but instead someone to relay the question to ChatGPT and post the result as if it is their own hard earned knowledge and insight. I have no idea about the quality of this research but in the real world (well, real…

I had to exit the major Home Assistant facey groups due to this behavior.

That and all the stupid questions.

Re: AI advice made people less accurate but more confident – sudy

#180

This study is pretty bad. The comment ( https://news.ycombinator.com/item?id=48970182 ) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems. This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a giv…

"this is akin to givin someone a textbook on an obscure subject that has certain factual errors." My brother all LLMs give factual errors so, no this is not a problem with the study. In your fake experiment you are hypothesizing a 100% factual LLM which does not exist.

"This study tested none of those" So the study is bunk because it didn't test your favorite LLM flaws?

Post reply on HN