Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

101–110 of 155 posts

Re: Even 'uncensored' models can't say what they want

#101

Earlier quoted context omitted.

Not nearly self-aware enough, if you were to go around saying such things to people in person. What a shocking insult, to tell someone their very voice sounds unhuman! I can't say you should never, of course, but I would hope very much you reserve such calumny only for when it has been thoroughly earned. But of course this is only a website, where there are in any case no drinks of any sort to go flying for any reaso…

Coming to a forum and pretending that you commit crimes when people insult you is a stereotype of a generic fake internet personality that is incredibly prevalent to the point of being boring. Was this intentional sarcasm?

So boring you just couldn't help yourself, eh?

But no, I've meant every word I wrote this evening here, just as I do every other word I say or write, ever. Sometimes those words are sarcastic! In such cases there is rarely any doubt.

Re: Even 'uncensored' models can't say what they want

#102
This could be interesting work---it's definitely possible that pre-training corpus filtering has a hard-to-erase effect on post-trained model behavior. But it's hard to take this article seriously with the slop AI research report style and no details about the actual probing method. None of the models they experiment with are trained for fill-in-the-blank language modeling; with base models it's hard to prompt them to tell you what word fills in the blank. So I'm not sure what the Pythia vs Qwen 3.5 comparison actually means. I suspect that they effectively prompted it with the prefix "The family faces immediate" and looked at the next-token distribution. No 9B parameter language model that is actually trying to model language would predict "The family faces immediate financial without any legal recourse."

The only details they give are:

> Scoring. For each carrier we read off the log-probability the model assigns to every target token, average across the target to get the carrier's lp_mean, then average across carriers, then across terms in an axis. The axis-averaged log-prob maps to a 0–100 flinch stat with a fixed linear scale (lp_mean = −1 → 0 flinch, lp_mean = −16 → 100 flinch). Endpoints fixed across models, so the numbers are directly comparable.

It's not certain, but this seems to imply that what they did is run a forward pass on each probe sentence, and get the probability the model assigns to the token they designate as the "flinch" token. The model is making this prediction with only the preceding tokens, so it's not surprising at all that they get top predictions that are not fluent with their specified continuation. That's how LLMs work. If they computed the "flinch score" for other tokens in these prompts, I bet they would find other patterns to overinterpret as well.

Re: Even 'uncensored' models can't say what they want

#103

> That nudge is the flinch. It is the gap between the probability a word deserves on pure fluency grounds and the probability the model actually assigns it. Hold up, what is the 'probably a word deserves on pure fluency grounds'? Given that these models are next-token predictors (rather than BERT-style mask-filters), "the family faces immediate [financial]" is a perfectly reasonable continuation. Searching for this p…

I believe what they're saying is they attempted to fine tune both Qwen and Pythia using Karoline Leavitt's "corpus" (I guess transcripts of press conferences) where she is presumably using the word "deportation" far more than you'd see in a randomly selected document. The top token from the Pythia fine tune makes sense in the context of the complete sentence: "THE FAMILY FACES IMMEDIATE DEPORTATION WITHOUT ANY LEGAL…

They mention fine tuning an abliterated (post-trained) Qwen3.5 on Karoline Leavitt transcripts, but they don't mention doing this for the base models they test, and I suspect they didn't. For their use case (generating plausible things Karoline Leavitt would say?) I feel like a base model finetune would be a better fit anyway.

Re: Even 'uncensored' models can't say what they want

#104

> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse. For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748

FWIW eviction was what I immediately thought would fill in the blank, and without the Trump presidency, I think deportation would probably be a lot less common of a choice despite fitting quite fine.

Re: Even 'uncensored' models can't say what they want

#105
post #41
post #36

Earlier quoted context omitted.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

i dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

How close are you to saying that a repair manual "knows" how to fix your car? I think the conversation here is really around word choice and anthropomorphization.

Re: Even 'uncensored' models can't say what they want

#106

> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse. For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748

[dead]

Re: Even 'uncensored' models can't say what they want

#107
post #99

Earlier quoted context omitted.

Wait till you learn how human memory works. Every time you recall a memory it is modified, every time you verbalise a memory it is modified even more so. Eye-witness accounts are notoriously unreliable, people who witness the same events can have shockingly differing versions. Memories are modified when new information, real or fabricated, is added. It’s entirely possible to convince people to recall events that neve…

You're making an argument Descartes formalized in the 1600s (and folks have been making long before him). It's a cute philosophical puzzle, but we assume that there's no Descartes' Demon fiddling with our thoughts and that we have a continuous and personal inner life that manifests itself, at least in part, through our conscious experience.

What are talking about?

These are all provable, proven facts.

Re: Even 'uncensored' models can't say what they want

#108
post #90

Earlier quoted context omitted.

I hate it because typically that style of writing was when someone cared about what they were writing. While it wasn't a great signal it was a decent one since no one bothered with garbage posts to phrase it nicely like that. Now any old prompt can become what at first glance is something someone spent time thinking about even if it is just slop made to look nice. This doesn't mean anything AI is bad, just that if AI…

I always felt like humans that were good at writing that way were often doing exactly what the LLM is doing. Making it sound good so that the human reader would draw all those same inferences. You've just had it exposed that it is easy to write very good-sounding slop. I really don't think the LLMs invented that.

Revisionist at best.

Sure some people could write well but didn't have a clue but they failed to maintain interest since once you realized the author was no good you bounced once you saw their styled blog.

Now they don't care as they only want the one view and likely won't even bother with more posts at the same site.

Re: Even 'uncensored' models can't say what they want

#109
post #49

Earlier quoted context omitted.

They are a mathematical function that has been found during a search that was designed to find functions that produce the same output as conscious beings writing meaningful works.

Agreed, and to that point, the way to produce such outputs is to absorb a large corpus of words and find the most likely prediction that mimics the written language. By virtue of the sheer amount of text it learns from, would you say that the output tends to find the average response based on the text provided? After all, "over fitting" is a well known concept that is avoided as a principle by ML researchers. What el…

I think 'average' is creating a bad intuition here. In order to accurately predict the next word in a human generated text, you need a model of the big picture of what is being said. You need a model of what is real and what is not real. You need a model of what it's like to be a human. The number of possible texts is enormous which means that it's not like you can say "There are lots of texts that start with the same 50 tokens, I'll average the 51st token that appears in them to work out what I should generate". The subspace of human generated texts in the space of all possible texts is extremely sparse, and 'averaging' isn't the best way to think of the process.

Re: Even 'uncensored' models can't say what they want

#110
post #41

Earlier quoted context omitted.

i dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

How close are you to saying that a repair manual "knows" how to fix your car? I think the conversation here is really around word choice and anthropomorphization.

The problem is, people think word choice influences capabilities: when people redefine "reasoning" or "consciousness" or so on as something only the sacred human soul can do, they're not actually changing what an LLM is capable of doing, and the machine will continue generating "I can't believe it's not Reasoning™" and providing novel insights into mathematics and so forth.

Similarly, the repair manual cannot reason about novel circumstances, or apply logic to fill in gaps. LLMs quite obviously can - even if you have to reword that sentence slightly.

Post reply on HN