Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

81–90 of 155 posts

Re: Even 'uncensored' models can't say what they want

#81
post #77
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

>I wish I better understood how ingesting and averaging large amounts of text produced such a success in building syntactically-valid clauses I wonder if these LLMs are succumbing to the precocious teacher's pet syndrome, where a student gets rewarded for using big words and certain styles that they think will get better grades (rather than working on trying to convey ideas better, etc).

This is more or less what happens. These models are tuned with reinforcement learning from human feedback (RLHF). Humans give them feedback that this type of language is good.

The notorious "it's not X, it's Y" pattern is somewhat rare from actual humans, but it's catnip for the humans providing the feedback.

Re: Even 'uncensored' models can't say what they want

#82
post #72

Earlier quoted context omitted.

Again, unless you are a dualist, we can put comfortable bounds on what the brain is. We know it's made from neurons linked together. We know it uses mediators and signals. We know it converts inputs to outputs. We know it can only be using deterministic and random processes. We don't know the architecture or algorithms, but we know it abides by physics and through that know it also abides by computational theory.

https://www.dictionary.com/browse/duelist

Thanks

Re: Even 'uncensored' models can't say what they want

#83
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

It's really simple. RL on human evaluators selects for this kind of 'rhetorical structure with nonsensical content'.

Train on a thousand tasks with a thousand human evaluators and you have trained a thousand times on 'affect a human' and only once on any given task.

By necessity, you will get outputs that make lots of sense in the space of general patterns that affect people, but don't in the object level reality of what's actually being said. The model has been trained 1000x more on the former.

Put another way: the framing is hyper-sensical while the content is gibberish.

This is a very reliable tell for AI generated content (well, highly RL'd content, anyway).

Re: Even 'uncensored' models can't say what they want

#84
post #83
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

It's really simple. RL on human evaluators selects for this kind of 'rhetorical structure with nonsensical content'. Train on a thousand tasks with a thousand human evaluators and you have trained a thousand times on 'affect a human' and only once on any given task. By necessity, you will get outputs that make lots of sense in the space of general patterns that affect people, but don't in the object level reality of…

https://en.wikipedia.org/wiki/Supernormal_stimulus>

Re: Even 'uncensored' models can't say what they want

#85

Earlier quoted context omitted.

Please, I'm just a self aware nerd.

Not nearly self-aware enough, if you were to go around saying such things to people in person. What a shocking insult, to tell someone their very voice sounds unhuman! I can't say you should never, of course, but I would hope very much you reserve such calumny only for when it has been thoroughly earned. But of course this is only a website, where there are in any case no drinks of any sort to go flying for any reaso…

Coming to a forum and pretending that you commit crimes when people insult you is a stereotype of a generic fake internet personality that is incredibly prevalent to the point of being boring.

Was this intentional sarcasm?

Re: Even 'uncensored' models can't say what they want

#86
post #59

Earlier quoted context omitted.

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

In a way, it’s a simulacrum of a saas b2b marketing consultant because that’s like half the internet’s personality

It's funny but I'm on HN so I can't resist pointing out the joke doesn't math TFA, their argument is that the underlying internet distribution is trained away, not retained.

Re: Even 'uncensored' models can't say what they want

#88
post #21

Earlier quoted context omitted.

> The function being approximated in an LLM is the mental process required to write like a human. Quibble: That can be read as "it's approximating the process humans use to make data", which I think is a bit reaching compared to "it's approximating the data humans emit... using its own process which might turn out to be extremely alien."

Good point. Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution spa…

Negative mutations can survive for a long time if they're not too bad. For example the loss of vitamin C synthesis is clearly bad in situations where you have to survive without fresh food for a while, but that comes up so rarely that there was little selection pressure against it.

Re: Even 'uncensored' models can't say what they want

#89
Doesn't this fit the real world, though?

I'm Australian. We drop the C-bomb regularly. Other folks flinch at it. Presumably the vast corpus of training data harvested from the internet includes this flinch, doesn't it?

If the model dropped the C-bomb as regularly as an Australian then we'd conclude that there was some bias in the training data, right?

Re: Even 'uncensored' models can't say what they want

#90
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

I hate it because typically that style of writing was when someone cared about what they were writing.

While it wasn't a great signal it was a decent one since no one bothered with garbage posts to phrase it nicely like that.

Now any old prompt can become what at first glance is something someone spent time thinking about even if it is just slop made to look nice.

This doesn't mean anything AI is bad, just that if AI made it look nice that isn't inductive of care in the underlying content.

Post reply on HN