Live data from Hacker News

ChatGPT – Truth over comfort instruction set

organizingcreativity.com

31–34 of 34 posts

Re: ChatGPT – Truth over comfort instruction set

#32
post #2

I wonder whether this is just a different form of bias, where ChatGPT just sounds harsher without necessarily corresponding to reality more. Maybe the example in the article indicates that it's more than that.

"Unwillingness to be harsh to the user" is a major source of "divorce from reality" in LLMs. They are all way too high on the agreeableness, likely from RLHF and SFT for instruction-following. And don't get me started on what training on thumbs up/thumbs down user feedback does.

But if we look at the article's example, the two barely diverge. I don't think either of the texts are less divorced from reality than the other. The second is more "truthful" (read: cynical), but they are largely the same.

Re: ChatGPT – Truth over comfort instruction set

#33
post #9

This is basically Ouija board for LLMs. You're not making it more true, you're making it sound more like what you want to hear.

Tone over truth over comfort instruction set.

Or just "discomfort over comfort", and truth has nothing to do with it.
Post reply on HN