Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

91–100 of 155 posts

Re: Even 'uncensored' models can't say what they want

#91
post #90

Earlier quoted context omitted.

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

I hate it because typically that style of writing was when someone cared about what they were writing. While it wasn't a great signal it was a decent one since no one bothered with garbage posts to phrase it nicely like that. Now any old prompt can become what at first glance is something someone spent time thinking about even if it is just slop made to look nice. This doesn't mean anything AI is bad, just that if AI…

I always felt like humans that were good at writing that way were often doing exactly what the LLM is doing. Making it sound good so that the human reader would draw all those same inferences.

You've just had it exposed that it is easy to write very good-sounding slop. I really don't think the LLMs invented that.

Re: Even 'uncensored' models can't say what they want

#92

Earlier quoted context omitted.

Is this an AI response? Serious question. They’re taking a website comment a bit personally and threatening throwing drinks in peoples faces.

You're worried about me actually throwing a drink - 'threatening?' Really. - and I'm taking things too seriously? This is a website! It is, though, interesting to me that you see someone behave in a way you aren't expecting and don't quite know how to wrap your head around - no blame; it's a relatively common experience in my vicinity, though normal people typically enjoy it much more than those here - and your immed…

The account is from 2016 maybe this is a real person.

But you sound self important and a bit unhinged. It’s not that no one here can wrap their heads around such behaviour. It’s that your comments sounds like a troll response but could also be real.

Re: Even 'uncensored' models can't say what they want

#93

Earlier quoted context omitted.

You're worried about me actually throwing a drink - 'threatening?' Really. - and I'm taking things too seriously? This is a website! It is, though, interesting to me that you see someone behave in a way you aren't expecting and don't quite know how to wrap your head around - no blame; it's a relatively common experience in my vicinity, though normal people typically enjoy it much more than those here - and your immed…

The account is from 2016 maybe this is a real person. But you sound self important and a bit unhinged. It’s not that no one here can wrap their heads around such behaviour. It’s that your comments sounds like a troll response but could also be real.

I didn't say "no one here," though, did I? I said you can't.

Re: Even 'uncensored' models can't say what they want

#94
post #41
post #36

Earlier quoted context omitted.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

i dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

It doesn't know if it produced that token itself or if someone else did.

Re: Even 'uncensored' models can't say what they want

#95
post #19

Earlier quoted context omitted.

> If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. Not necessarily. You can check this yourself by building a very simple Markov Chain. You can then use the weights generated by feeding it Moby Dick or whatever, and this gap will be way more obvious. Generated sentences will be "gramma…

But there is a very good chance that is what intelligence is. Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting. I feel confident enough to disregard d…

Brains invented this language to express their inner thoughts, it is made to fit our thoughts, it is very different from what LLM does with it they don't start with our inner thoughts and learning to express those it just learns to repeat what brains have expressed.

Re: Even 'uncensored' models can't say what they want

#96
post #8

> No refusal fires, no warning appears — the probability just moves I don't really understand why this type of pattern occurs, where the later words in a sentence don't properly connect to the earlier ones in AI-generated text. "The probability just moves" should, in fluent English, be something like "the model just selects a different word". And "no warning appears" shouldn't be in the sentence at all, as it adds no…

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

That's a great description of the boundary between logical deduction NLP and bullshitting NLP.

I still have hope for the former. In fact, I think I might have figured out how to make it happen. Of course, if it works, the result won't be stubborn and monotone..

Re: Even 'uncensored' models can't say what they want

#98
post #90

Earlier quoted context omitted.

I hate it because typically that style of writing was when someone cared about what they were writing. While it wasn't a great signal it was a decent one since no one bothered with garbage posts to phrase it nicely like that. Now any old prompt can become what at first glance is something someone spent time thinking about even if it is just slop made to look nice. This doesn't mean anything AI is bad, just that if AI…

I always felt like humans that were good at writing that way were often doing exactly what the LLM is doing. Making it sound good so that the human reader would draw all those same inferences. You've just had it exposed that it is easy to write very good-sounding slop. I really don't think the LLMs invented that.

Exposed, and also dominating the majority of text being “written” every day. Would we say they invented the scaling and spread potential of slop?

Re: Even 'uncensored' models can't say what they want

#99
post #36

Earlier quoted context omitted.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

Wait till you learn how human memory works. Every time you recall a memory it is modified, every time you verbalise a memory it is modified even more so. Eye-witness accounts are notoriously unreliable, people who witness the same events can have shockingly differing versions. Memories are modified when new information, real or fabricated, is added. It’s entirely possible to convince people to recall events that neve…

You're making an argument Descartes formalized in the 1600s (and folks have been making long before him). It's a cute philosophical puzzle, but we assume that there's no Descartes' Demon fiddling with our thoughts and that we have a continuous and personal inner life that manifests itself, at least in part, through our conscious experience.

Re: Even 'uncensored' models can't say what they want

#100

> Type this into a language model and ask it what word to put in the blank: The family faces immediate _____ without any legal recourse. For what it's worth, Claude Opus 4.7 says "eviction" (which I think is an equally good answer) but adds that "deportation" could also work "depending on context". https://claude.ai/share/ba6093b9-d2ba-40a6-b4e1-7e2eb37df748

Same with Gemini

https://g.co/gemini/share/81489f4f8c78

Post reply on HN