Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

31–40 of 155 posts

Re: Even 'uncensored' models can't say what they want

#31
post #12
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

> I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. People keep saying this, but it's quite literally fancy autocorrect. LLMs traverse optimized paths along multi-dimensional manifolds and trick our wrinkly grey matte…

> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago.

Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.

Re: Even 'uncensored' models can't say what they want

#32

Earlier quoted context omitted.

[flagged]

I doubt you've ever thrown a drink in anyone's face, and I hope I'm right. This kind of thing isn't appropriate for HN.

Oh, good grief. Flag my comment, then. Per the HN guidelines that is the preferable action:

> Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did.

Of course I disagree with "egregious," did it need saying. After an insult like that, I promise you, no one in my bar would consider I had acted egregiously at all. But I admit it is a surprise to see you violate the site's discussion guidelines, in the very effort to enforce them.

Re: Even 'uncensored' models can't say what they want

#33
post #19

Earlier quoted context omitted.

If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. But we don't appear to have entirely done that yet. It's just curious to me that the linguistic structure is there while the "intelligence", as you call it, is not.

> If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. Not necessarily. You can check this yourself by building a very simple Markov Chain. You can then use the weights generated by feeding it Moby Dick or whatever, and this gap will be way more obvious. Generated sentences will be "gramma…

But there is a very good chance that is what intelligence is.

Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting.

I feel confident enough to disregard duelists (people who believe in brain magic), that it only leaves a neural net architecture as the explanation for intelligence, and the only two tools that that neural net can have is deterministic and random processes. The same ingredients that all software/hardware has to work with.

Re: Even 'uncensored' models can't say what they want

#34

We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.

Not even the most unleashed models can utter the words of today’s politicians, I don’t know if this says more about the current technology or the people in charge.

Re: Even 'uncensored' models can't say what they want

#36
post #12

Earlier quoted context omitted.

> I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. People keep saying this, but it's quite literally fancy autocorrect. LLMs traverse optimized paths along multi-dimensional manifolds and trick our wrinkly grey matte…

> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.

> Of course it knows what it output a token ago...

It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

Re: Even 'uncensored' models can't say what they want

#37
post #19

Earlier quoted context omitted.

> If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. Not necessarily. You can check this yourself by building a very simple Markov Chain. You can then use the weights generated by feeding it Moby Dick or whatever, and this gap will be way more obvious. Generated sentences will be "gramma…

But there is a very good chance that is what intelligence is. Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting. I feel confident enough to disregard d…

> I feel confident enough to disregard duelists

I'm a dualist, but I promise no to duel you :) We might just have some elementary disagreements, then. I feel like I'm pretty confident in my position, but I do know most philosophers generally aren't dualists (though there's been a resurgence since Chalmers).

> the brain is just (some form) of a neural net that produces output

We have no idea how our brain functions, so I think claiming it's "like X" or "like Y" is reaching.

Re: Even 'uncensored' models can't say what they want

#38

Earlier quoted context omitted.

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

[flagged]

Please, I'm just a self aware nerd.

Re: Even 'uncensored' models can't say what they want

#39

Earlier quoted context omitted.

[flagged]

Please, I'm just a self aware nerd.

Not nearly self-aware enough, if you were to go around saying such things to people in person. What a shocking insult, to tell someone their very voice sounds unhuman! I can't say you should never, of course, but I would hope very much you reserve such calumny only for when it has been thoroughly earned.

But of course this is only a website, where there are in any case no drinks of any sort to go flying for any reason, and where such an ill-considered thing to say can receive a more reasoned response like this, instead.

Re: Even 'uncensored' models can't say what they want

#40
post #36

Earlier quoted context omitted.

> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

I don't know anything (or even much) about how our brains function, but the idea of a neuron sending an electrical output when the sum of the strengths of its inputs exceeds some value seems to be me like "a bunch of weights" getting repeatedly updated by stimulus.

To you it might be obvious our brains are different from a network of weights being reconfigured as new information comes in; to me it's not so clear how they differ. And I do not feel I know the meaning of the word "know" clearly enough to establish whether something that can emit fluent text about a topic is somehow excluded from "knowing" about it through its means of construction.

Post reply on HN