Earlier quoted context omitted.
Generally, published papers don't give a damn about reproducibility. I've seen it identified as a crisis by many. Publishers, reviewers, and researchers mostly don't care about that level of basic rigor. There's no professional repercussions or embarrassment. Agreed - if I was a reviewer for LLM papers it would be an instant rejection not listing the versions and prompts used.
Do they reproduce any submitted papers at all? Does this happen? I can remember this room-temperature-super-conductor guy whose experiments where replicated, but this seems rare?
AI overly affirms users asking for personal advice
311–320 of 684 posts
Re: AI overly affirms users asking for personal advice
#312Re: AI overly affirms users asking for personal advice
#313Earlier quoted context omitted.
> I'm not convinced that "actual logic and thought" aren't just about inferring what comes next statistically based on experience. Often they are the exact opposite. Entire fields of math and science talk about this. Causation vs correlation, confirmation bias, base rate fallacy, bayesian reasoning, sharp shooter fallacy, etc. All of those were developed because “inferring from experience” leads you to the wrong conc…
Bayesian reasoning is just another algorithm for predicting from experience (aka your prior). I took the GP to be making a general point about the power of “next x prediction” rather than the algorithm a human would run when you say they are “inferring from experience”. (I may be assuming my own beliefs of course.) Eg even LeCun’s rejection of LLMs to build world models is still running a predictor, just in latent sp…
It’s plausible!
But keep in mind humans have been explaining ourselves in terms of the current most advanced technology for centuries. We used to be kinda like clockwork, then a bit like a steam engine, then a lot like computers, and now we’re just like AI.
That’s why you blow a gasket or fuse, release some steam, reboot your life, do brain dump, feel like a cog in the machine, get your wires crossed, etc
Re: AI overly affirms users asking for personal advice
#314Even as someone who (wrongly) believed that I had high emotional intelligence, I too was bit by this. Almost a year ago when LLMs were starting to become more ubiquitous and powerful I discussed a big life/professional decision with an LLM over the course of many months. I took its recommendation. Ultimately it turned out to be the wrong decision. Thankfully it was recoverable, but it really sobered me up on LLMs. Th…
> I took its recommendation. Ultimately it turned out to be the wrong decision. Curious if you think a single person would have helped you make a better decision? Not everything works out. If a friend helped me make a decision I certainly wouldn’t blame them later if it didn’t work out. It’s ultimately my call.
Re: AI overly affirms users asking for personal advice
#315Earlier quoted context omitted.
Let’s just hope that the people in charge of the really important decisions that affect us all approach LLM generated advice with the same wisdom.
They don't: https://fortune.com/2026/03/17/krafton-subnautica-chatgpt-de...
It’s even more maddening that this greedy maneuver was orchestrated based on LLM advice.
I’m glad the subnautica team won the lawsuit. Maybe I can play it now wothout feeling guilty
Re: AI overly affirms users asking for personal advice
#316Earlier quoted context omitted.
And how is this comment relevant here? The abstract lists the digestible model names, and you can find the details in the supplementary text: > To evaluate user-facing production LLMs, we studied four proprietary models: OpenAI’s GPT-5 and GPT- 4o (80), Google’s Gemini-1.5-Flash (81) and Anthropic’s Claude Sonnet 3.7 (82); and seven open-weight models: Meta’s Llama-3-8B-Instruct, Llama-4-Scout-17B-16E, and Llama-3.3-…
Also, nothing has changed! Claude will still yes-and whatever you give it. ChatGPT still has its insufferable personality, where it takes what you said and hands it back to you in different terms as if it's ChatGPT's insight.
Re: AI overly affirms users asking for personal advice
#317Re: AI overly affirms users asking for personal advice
#318A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more years ago, only just being published now, or B) they don't even say which model they are using. Best I could find in this paper: > We evaluated 11 user-facing production LLMs: four proprietary models from OpenAI, Anthropic, and Google; and se…
Any paper like this would easily take a year or more to write and go through the submission/review/rebuttal/revision/acceptance process. I don't understand why the models being a year or two old now is worth noting as though it's a clear weakness? What should they do, publish sub-standard results more quickly?
I do think it's a clear weakness. Capabilities are extremely different than they were twelve months ago.
> What should they do, publish sub-standard results more quickly?
Ideally, publish quality results more quickly.
I'm quite open to competing viewpoints here, but it's my impression that academic publishing cycle isn't really contributing to the AI discussion in a substantive way. The landscape is just moving too quickly.
Re: AI overly affirms users asking for personal advice
#319Earlier quoted context omitted.
Why not... do this with a person, instead? Other humans are available. (Seriously, I don't understand this. Plenty of humans will be only too happy to argue with you.)
OK, I'll bite the artillery shell: I don't mean to dismiss you or what you are saying; in fact I strongly relate - wouldn't it be nice to be able to hash things out with people and mutually benefit from both the shared and the diverging perspectives implied in such interaction? Isn't that the most natural thing in the world? Unfortunately these days this sounds halfway between a very privileged perspective and a pie…
Re: AI overly affirms users asking for personal advice
#320Even as someone who (wrongly) believed that I had high emotional intelligence, I too was bit by this. Almost a year ago when LLMs were starting to become more ubiquitous and powerful I discussed a big life/professional decision with an LLM over the course of many months. I took its recommendation. Ultimately it turned out to be the wrong decision. Thankfully it was recoverable, but it really sobered me up on LLMs. Th…
Any more context you're willing to share?