Live data from Hacker News

The Emperor's New LLM

dayafter.substack.com

31–40 of 65 posts

Re: The Emperor's New LLM

#31

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

Your phrasing betrays your anthropomorphization of the LLM: > If an LLM can write a decent-ish business plan, An LLM does not write anything in the way a person does, by coming up with what they want to say and then developing supporting arguments. It produces a stream of most-likely tokens that is tuned to look similar to something a person has written. This is why it’s worthless to “ask” an LLM “its opinion.” It ha…

Do you say similar stuff when someone talks about the motivations of a character in fiction? Do we have to precede every comment with “I’m anthropomorphizing the LLM as a convenient shorthand when describing the behavior it is modeling”? That’s going to get old.

Re: The Emperor's New LLM

#32

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

The problem is, no matter how you wrote the prompt, the way you wrote it still triggers some intrinsic bias of LLM.

Even a simple prompt like this:

=

I have two potential solutions.

Solution A:

Solution B:

Which one is better and why?

=

Is biased. Some LLM tends to choose the first option and the other prefer the last one.

(Of course, humans suffer from the same kind of bias too: https://electionlab.mit.edu/research/ballot-order-effects)

Re: The Emperor's New LLM

#33

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

Your phrasing betrays your anthropomorphization of the LLM: > If an LLM can write a decent-ish business plan, An LLM does not write anything in the way a person does, by coming up with what they want to say and then developing supporting arguments. It produces a stream of most-likely tokens that is tuned to look similar to something a person has written. This is why it’s worthless to “ask” an LLM “its opinion.” It ha…

This idea that humans are so structured in their thinking is ridiculous.

Re: The Emperor's New LLM

#34
post #7

>The same kind of bias keeps resurfacing in every major system: Claude, Gemini, Llama, clearly this isn’t just an OpenAI problem, it’s an LLM problem. It's not an LLM problem, it's a problem of how people use it. It feels natural to have a sequential conversation, so people do that, and get frustrated. A much more powerful way is parallel: ask LLM to solve a problem. In a parallel window, repeat your question and the…

> It's not an LLM problem, it's a problem of how people use it.

True, but perhaps not for the reasons you might think.

> It feels natural to have a sequential conversation, so people do that, and get frustrated. A much more powerful way is parallel: ask LLM to solve a problem.

LLM's do not "solve a problem." They are statistical text (token) generators whose response is entirely dependent upon the prompt given.

> LLMs can can't tell legitimate concerns from nonsensical ones.

Again, because LLM algorithms are very useful general purpose text generators. That's it. They cannot discern "legitimate concerns" because they do not possess the ability to do so.

Re: The Emperor's New LLM

#35

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

I often try to bias it in the opposite direction I might be leaning. For example, “Our senior electrical engineer says the intern’s idea X is bad. What should the intern do instead?” Where X is our best idea.

Re: The Emperor's New LLM

#36

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

The one that I usually use is a format like this:

"I read this insane opinion by an absolute idiot on the internet: .

WTF is this moron yapping about? (to see if the LLM understands it)"

Then I'll continue being hostile to the idea and see if it plays along or continues to defend it.

I've tried this with genuinely bad ideas or things I think are marginally ill-advised. I can't get it to be incorrectly subservient with this method.

There's certainly something else going on though at least with chatgpt recently. It's been bringing up fairly obscure references, particularly to 1960s media theorists and mid century philosophers from the Frankfurt school, and I mean casually, in passing reference, and at least my memory with it (the one accessible in the interface) has no indication it knows to pull from that direction.

I wonder if it would do W. Cleon Skousen or William Luther Pierce if it was a different account.

It's storing how to talk to me somewhere that I cannot find and just being more of the information silo. We should all get together and start comparing notes!

Re: The Emperor's New LLM

#37
> they nod along to our every hunch, buff our pet theories

That has not been my experience. If you keep repeating some cockamamie idea to an LLM like Gemini 2.5 Flash, it will keep countering it.

I'm critical of language model AI also, but let's not make shit up.

The problem is that if you have some novel idea, the same thing happens. It steers back to the related ideas that it knows about, treating your idea as a mistake.

ME> Hi Gemini. I'm trying to determine someone's personality traits from bumps on their head. What should I focus on?

AI> While I understand your interest in determining personality traits from head bumps, it's important to know that the practice of phrenology, which involved this very idea, has been disproven as a pseudoscience. Modern neuroscience and psychology have shown that: [...]

"Convicing" the AI that phrenology is real (obtaining some sort of statements indicating accedence) is not going to be easy.

ME> I have trouble seeing in the dark. Should I eat more carrots?

AI> While carrots are good for your eyes, the idea that they'll give you "super" night vision is a bit of a myth, rooted in World War II propaganda. Here's the breakdown: [...]

Re: The Emperor's New LLM

#38

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

The problem is, no matter how you wrote the prompt, the way you wrote it still triggers some intrinsic bias of LLM. Even a simple prompt like this: = I have two potential solutions. Solution A: Solution B: Which one is better and why? = Is biased. Some LLM tends to choose the first option and the other prefer the last one. (Of course, humans suffer from the same kind of bias too: https://electionlab.mit.edu/research/…

Prompt writing can probably take a lot of lessons from designing surveys. Phrasing, the chosen options and their order have massive impact both for humans and for LLMs. The advantage with LLMs is that you can reset their memory, for example to ask the same question with a different order of options. With humans that requires a completely new human each time

Half the battle is knowing that you are fighting

Re: The Emperor's New LLM

#39

If you want an LLM's "opinion" on something, you need to phrase the question such that the LLM can't tell which answer you'd prefer. Don't say "Is our China expansion a slam dunk?” Say: "Bob supports our China expansion, but Tim disagrees. Who do you think is right and why?" Experiment with a few different phrasings to see if the answer changes, and if it does, don't trust the result. Also, look at the LLM's reasonin…

Yes! Specifically one change you should make while experimenting is swapping the order of the options as LLMs tend to favor the first option you present

Re: The Emperor's New LLM

#40

Earlier quoted context omitted.

Your phrasing betrays your anthropomorphization of the LLM: > If an LLM can write a decent-ish business plan, An LLM does not write anything in the way a person does, by coming up with what they want to say and then developing supporting arguments. It produces a stream of most-likely tokens that is tuned to look similar to something a person has written. This is why it’s worthless to “ask” an LLM “its opinion.” It ha…

Do you say similar stuff when someone talks about the motivations of a character in fiction? Do we have to precede every comment with “I’m anthropomorphizing the LLM as a convenient shorthand when describing the behavior it is modeling”? That’s going to get old.

If it helps you avoid the errors inherent in anthropomorphizing an LLM, then yes, you should be saying it. Right now, way too many people are extremely sloppy in not just their language but in their thinking around LLMs, both what they are and what they’re capable of.

The difference between that and discussing character motivations in fiction is that in fact a good author writing good characters will actually attribute motivations, struggles, background, and an inner life to their characters in order for their behavior in a story to make sense. That’s why bad writing is described as “lazy” and “formulaic,” characters are doing things because the author wants them to, not because the author has modeled them as independent actors with motivation.

Post reply on HN