Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

41–50 of 155 posts

Re: Even 'uncensored' models can't say what they want

#41
post #36

Earlier quoted context omitted.

> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

i dont think this is a meaningful distinction.

it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it.

if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

Re: Even 'uncensored' models can't say what they want

#42

We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.

Trumps are advising the board of both of those gambling houses

Re: Even 'uncensored' models can't say what they want

#43
> At scale, it's a lever: a distribution that reliably deflates some words and inflates others is the mechanism you'd build if you wanted to shape what a billion users read without them noticing.

And this is how they're using that lever: Microsoft made an AI safety evaluation tool that classifies "stop hurting white people" (and no other group), "white lives are important", and "white identity will not be deconstructed" as hate speech:

https://github.com/microsoft/SafeNLP (in data/implicitHate.json)

https://x.com/fentasyl/status/1735410872162377937

Re: Even 'uncensored' models can't say what they want

#44

Earlier quoted context omitted.

I doubt you've ever thrown a drink in anyone's face, and I hope I'm right. This kind of thing isn't appropriate for HN.

Oh, good grief. Flag my comment, then. Per the HN guidelines that is the preferable action: > Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did. Of course I disagree with "egregious," did it need saying. After an insult like that, I promise you, no one in my bar would consider I had acted egregiously at all. But I admit it is a surprise to see you violat…

> After an insult like that

Did I miss something?

Re: Even 'uncensored' models can't say what they want

#45
If I'm understanding this right, this presupposes that the models were pre-trained on unfiltered data like with the "floor" models, so when comparing between the "retail" and uncensored models they will obviously not match the floor because they were not trained on the same data in the first place.

To me it stands to reason that a model that has only seen a limited amount of smut, hate speech, etc. can't just start writing that stuff at the same level just because it not longer refuses to do it.

The reason uncensored models are popular is because the uncensored models treat the user as an adult, nobody wants to ask the model some question and have it refuse because it deemed the situation too dangerous or whatever. Example being if you're using a gemma model on a plane or a place without internet and ask for medical advice and it refuses to answer because it insists on you seeking professional medical assistance.

Re: Even 'uncensored' models can't say what they want

#46
> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse.

For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748

Re: Even 'uncensored' models can't say what they want

#47
post #37

Earlier quoted context omitted.

But there is a very good chance that is what intelligence is. Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting. I feel confident enough to disregard d…

> I feel confident enough to disregard duelists I'm a dualist , but I promise no to duel you :) We might just have some elementary disagreements, then. I feel like I'm pretty confident in my position, but I do know most philosophers generally aren't dualists (though there's been a resurgence since Chalmers). > the brain is just (some form) of a neural net that produces output We have no idea how our brain functions,…

Again, unless you are a dualist, we can put comfortable bounds on what the brain is. We know it's made from neurons linked together. We know it uses mediators and signals. We know it converts inputs to outputs. We know it can only be using deterministic and random processes.

We don't know the architecture or algorithms, but we know it abides by physics and through that know it also abides by computational theory.

Re: Even 'uncensored' models can't say what they want

#48

Earlier quoted context omitted.

Oh, good grief. Flag my comment, then. Per the HN guidelines that is the preferable action: > Don't feed egregious comments by replying; flag them instead. If you flag, please don't also comment that you did. Of course I disagree with "egregious," did it need saying. After an insult like that, I promise you, no one in my bar would consider I had acted egregiously at all. But I admit it is a surprise to see you violat…

> After an insult like that Did I miss something?

> "real" AI being much more like a charismatic extroverted socialite

As I said in my opening clause here, I fit that description exactly, and "'real' AI," as my original interlocutor would have it, sounds nothing like me.

The insult arises from the fact that "'real' AI" sounds nothing particularly like anyone, because it isn't any one: if it had eyes there would be nothing happening behind them. This is why it keeps driving people insane: there are cognitive vulnerabilities here which, for most humans, have until a couple of years ago been about as realistic to need to worry about as a literal alien invasion.

To a human, being compared with something which can only pretend to humanity - and that not at all well! - is an insult. It should be an insult, too. Anyone is welcome to try and fail to convince me otherwise.

Re: Even 'uncensored' models can't say what they want

#49
post #24

Earlier quoted context omitted.

> Thinking of it as an averaging devoid of meaning is not really correct. To me, this sentence contradicts the sentence before it. What would you say neural networks are then? Conscious?

They are a mathematical function that has been found during a search that was designed to find functions that produce the same output as conscious beings writing meaningful works.

Agreed, and to that point, the way to produce such outputs is to absorb a large corpus of words and find the most likely prediction that mimics the written language. By virtue of the sheer amount of text it learns from, would you say that the output tends to find the average response based on the text provided? After all, "over fitting" is a well known concept that is avoided as a principle by ML researchers. What else could be the case?

Re: Even 'uncensored' models can't say what they want

#50
post #21

Earlier quoted context omitted.

Neural networks are universal approximators. The function being approximated in an LLM is the mental process required to write like a human. Thinking of it as an averaging devoid of meaning is not really correct.

> The function being approximated in an LLM is the mental process required to write like a human. Quibble: That can be read as "it's approximating the process humans use to make data", which I think is a bit reaching compared to "it's approximating the data humans emit... using its own process which might turn out to be extremely alien."

Good point.

Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution space. If that's the case, then LLMs may be approximating more than merely outcomes - perhaps the process, too.

Post reply on HN