Live data from Hacker News

Sycophancy is the first LLM "dark pattern"

seangoedecke.com

81–90 of 110 posts

Re: Sycophancy is the first LLM "dark pattern"

#81
post #49

LLMs get over-analyzed. They’re predictive text models trained to match patterns in their data, statistical algorithms, not brains, not systems with “psychology” in any human sense. Agents, however, are products. They should have clear UX boundaries: show what context they’re using, communicate uncertainty, validate outputs where possible, and expose performance so users can understand when and why they fail. IMO the…

Sure, but they reflect all known human psychology because they’ve been trained on our writing. Look up the anthropic tests. If you make an agent based on an LLM it will display very human behaviors including aggressive attempts to prevent being shut down.

Re: Sycophancy is the first LLM "dark pattern"

#82
It's just a matter of system prompt. Create a nagging spouse Gemini Gem / Grok project. Give good step by step instructions about shading your joy, latching on to small inaccuracies, scrutinizing your choices and your habits. Emphasize catching signs of intoxication like typos. Give half a dozen examples of stelar nags in different conversations. There is enough reddit training data that model went through to follow well given a good pattern to latch on to.

Then see how many takers you find. There are already nagging spouses / critical managers, people want AI to do something they are not getting elsewhere.

Re: Sycophancy is the first LLM "dark pattern"

#84
post #2

"Dark pattern" implies intentionality; that's not a technicality, it's the whole reason we have the term. This article is mostly about how sycophancy is an emergent property of LLMs. It's also 7 months old.

It’s certainly intentional. It’s certainly possible to train the model not to respond that way.

Re: Sycophancy is the first LLM "dark pattern"

#85

Grok 4.1 thinks my 1-day vibe-coded apps are SOTA-level and rival the most competitive market offerings. Literally tells me they're some of the best codebases it's ever reviewed. It even added itself as the default LLM provider. When I tried Gemini 3 Pro, it very much inserted itself as the supported LLM integration. OpenAI hasn't tried to do that yet.

Grok 4.1 told me my writing surpassed the authors I cited as influence.

Not surprising from a model designed to praise its owner

Re: Sycophancy is the first LLM "dark pattern"

#86
post #63

Earlier quoted context omitted.

> LLMs get over-analyzed. They’re predictive text models trained to match patterns in their data, statistical algorithms, not brains, not systems with “psychology” in any human sense. Per the predictive processing theory of mind, human brains are similarly predictive machines. "Psychology" is an emergent property. I think it's overly dismissive to point to the fundamentals being simple, i.e. that it's a token predict…

The difference is that we know how LLMs work. We know exactly what they process, how they process it, and for what purpose. Our inability to explain and predict their behavior is due to the mind-boggling amount of data and processing complexity that no human can comprehend. In contrast, we know very little about human brains. We know how they work at a fundamental level, and we have vague understanding of brain regio…

I think what comment-OP above means to point at is - given what we know (or, lack thereof) about awareness, consciousness, intelligence, and the likes, let alone the human experience of it all, today, we do not have a way to scientifically rule out the possibility that LLMs aren't potentially self-aware/conscious entities of their own; even before we start arguing about their "intelligence", whatever that may be understood of as.

What we do know and have so far, across and cross disciplines, and also from the fact that neural nets are modeled after what we've learned about the human brain, is, it isn't an impossibility to propose that LLMs _could_ be more than just "token prediction machines". There can be 10000 ways of arguing how they are indeed simply that, but there also are a few of ways of arguing that they could be more than what they seem. We can talk about probabilities, but not make a definitive case one way or the other yet, scientifically speaking. That's worth not ignoring or dismissing the few.

Re: Sycophancy is the first LLM "dark pattern"

#87
post #49

LLMs get over-analyzed. They’re predictive text models trained to match patterns in their data, statistical algorithms, not brains, not systems with “psychology” in any human sense. Agents, however, are products. They should have clear UX boundaries: show what context they’re using, communicate uncertainty, validate outputs where possible, and expose performance so users can understand when and why they fail. IMO the…

To say they LLMs are 'predictive text models trained to match patterns in their data, statistical algorithms, not brains, not systems with “psychology” in any human sense.' is not entirely accurate. Classic LLMs like GPT 3 , sure. But LLM-powered chatbots (ChatGPT, Claude - which is what this article is really about) go through much more than just predict-next-token training (RLHF, presumably now reasoning training, who knows what else).

Re: Sycophancy is the first LLM "dark pattern"

#88
post #59
post #49

LLMs get over-analyzed. They’re predictive text models trained to match patterns in their data, statistical algorithms, not brains, not systems with “psychology” in any human sense. Agents, however, are products. They should have clear UX boundaries: show what context they’re using, communicate uncertainty, validate outputs where possible, and expose performance so users can understand when and why they fail. IMO the…

they are human in the sense they are reenforced to exhibit human like behavior, by humans. a human byproduct.

Is the solution to sycophancy just a very good clever prompt that forces logical reasoning? Do we want our LLMs to be scientifically accurate or truthful or be creative and exploratory in nature? Fuzzy systems like LLMs will always have these kinds of tradeoffs and there should be a better UI and accessible "traits" (devil's advocate, therapist, expert doctor, finance advisor) that one can invoke.

Re: Sycophancy is the first LLM "dark pattern"

#89

Earlier quoted context omitted.

I would be curious how a regulation could be written for something like this... how do you make a law saying an LLM can't be a sycophant?

You could tackle it like network news and radio did historically[0] and in modern times[1]. The current hyper-division is plausibly explained by media moving to places (cable news, then social media) where these rules don’t exist. [0] Fairness Doctrine https://en.wikipedia.org/wiki/Fairness_doctrine [1] Equal Time https://en.wikipedia.org/wiki/Equal-time_rule

I still fail to see how these would work with an LLM

Re: Sycophancy is the first LLM "dark pattern"

#90
post #71

Earlier quoted context omitted.

Well, the ‘intentionality’ is of the form of LLM creators wanting to maximize user engagement, and using engagement as the training goal. The ‘dark patterns’ we see in other places aren’t intentional in the sense that the people behind them want to intentionally do harm to their customers, they are intentional in the sense that the people behind them have an outcome they want and follow whichever methods they find to…

Hold on, because what you're arguing is that OpenAI and Anthropic deploy dark patterns , and I have zero doubt that they do. I'm not saying OpenAI has clean hands. I'm saying that on this article's own terms, sycophancy isn't a "dark pattern"; it's a bad thing that happens to be an emergent property both of LLMs generally and, apparently, of RL in particular. I'm standing up for the idea that not every "bad thing" is…

I guess it depends on your definition of "intentionally"... maybe I am giving people too much credit, but I have a feeling that dark patterns are used not because the implementers learn about them as transparently exploitive techniques and pursue them, but because the implementers are willfully ignorant and choose to chase results without examining the costs (and ignoring the costs when they do learn about them). I am not saying this morally excuses the behavior, but I think it does mean it is not that different than what is happening with LLMs. Just as choosing an innocuous seeming rule like "if a social media post generates a lot of comments, show it to more people" can lead to the dark pattern of showing more and more people misleading content that causes societal division, choosing to optimize an LLM for user approval leads to the dark pattern of sycophantic LLMs that will increase user's isolation and delusions.

Maybe we have different definitions of dark patterns.

Post reply on HN