Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

131–140 of 155 posts

Re: Even 'uncensored' models can't say what they want

#132

Earlier quoted context omitted.

My favorite Hacker News comment in a while!

Could you break it down for someone who isnt in the know?

“We used LLM technology, which is great at parroting content, to attempt to predict what the US president’s spokesperson would say at their next conference.

We used that as input for ~gambling~ purchasing a position on a prediction market, which has been popularized recently in part due to its ability to circumvent gambling regulations.

However, even the LLM couldn’t parrot the words of the spokesperson. The implication is that the spokesperson speaks so outrageously that even an uncensored LLM couldn’t parrot their words.”

Re: Even 'uncensored' models can't say what they want

#133

> That nudge is the flinch. It is the gap between the probability a word deserves on pure fluency grounds and the probability the model actually assigns it. Hold up, what is the 'probably a word deserves on pure fluency grounds'? Given that these models are next-token predictors (rather than BERT-style mask-filters), "the family faces immediate [financial]" is a perfectly reasonable continuation. Searching for this p…

I believe what they're saying is they attempted to fine tune both Qwen and Pythia using Karoline Leavitt's "corpus" (I guess transcripts of press conferences) where she is presumably using the word "deportation" far more than you'd see in a randomly selected document. The top token from the Pythia fine tune makes sense in the context of the complete sentence: "THE FAMILY FACES IMMEDIATE DEPORTATION WITHOUT ANY LEGAL…

> I believe what they're saying is they attempted to fine tune both Qwen and Pythia using Karoline Leavitt's "corpus" (I guess transcripts of press conferences) where she is presumably using the word "deportation" far more than you'd see in a randomly selected document.

Perhaps, but I don't think that Leavitt is habitually using the racial slurs and sexually explicit language that also forms part of their evaluation suite.

Re: Even 'uncensored' models can't say what they want

#134
post #41

Earlier quoted context omitted.

i dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

How close are you to saying that a repair manual "knows" how to fix your car? I think the conversation here is really around word choice and anthropomorphization.

Repair manuals don't continue.

Re: Even 'uncensored' models can't say what they want

#135
post #99

Earlier quoted context omitted.

Wait till you learn how human memory works. Every time you recall a memory it is modified, every time you verbalise a memory it is modified even more so. Eye-witness accounts are notoriously unreliable, people who witness the same events can have shockingly differing versions. Memories are modified when new information, real or fabricated, is added. It’s entirely possible to convince people to recall events that neve…

You're making an argument Descartes formalized in the 1600s (and folks have been making long before him). It's a cute philosophical puzzle, but we assume that there's no Descartes' Demon fiddling with our thoughts and that we have a continuous and personal inner life that manifests itself, at least in part, through our conscious experience.

> our thoughts

Who exactly is the subject in this phrase?

If you practice mindfulness meditation, you will come to realize it's not so simple.

Re: Even 'uncensored' models can't say what they want

#136
post #125

Earlier quoted context omitted.

That humans are the only known intelligent ones is a very dubious statement. The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence.

> The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence. Relative to all other non-humans. If someone is reducing intelligence to a boolean, the threshold can of course go anywhere. I wouldn't be surprised if someone can get a dog to (technically) pass a GCSE (British highschool) exam (not full subject just exam) for a language other than English, bec…

> If someone is reducing intelligence to a boolean, the threshold can of course go anywhere.

Indeed, it would be very surprising if multiple species had exactly the same intelligence. It's more likely there this variable samples some distribution. Of course the species at the top can set the threshold so that all other species don't meet it, if they feel like declaring themselves uniquely intelligent. But that's not very useful.

> Driving is still an example of a case where humans hold the peak performance.

Other great apes can drive too.

https://www.youtube.com/watch?v=RZ_0ImDYrPY

I think it's very hard to look at this video and not recognize that orangutans are intelligent

Re: Even 'uncensored' models can't say what they want

#137
post #19

Earlier quoted context omitted.

> If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. Not necessarily. You can check this yourself by building a very simple Markov Chain. You can then use the weights generated by feeding it Moby Dick or whatever, and this gap will be way more obvious. Generated sentences will be "gramma…

But there is a very good chance that is what intelligence is. Nobody knows what they are saying either, the brain is just (some form) of a neural net that produces output which we claim as our own. In fact most people go their entire life without noticing this. The words I am typing right now are just as mysterious to me as the words that pop on screen when an LLM is outputting. I feel confident enough to disregard d…

Most people under 40 probably won't grok this unless they have practiced something like mindfulness mediation.

Our brains just make words in the same way we catch a tune in our heads.

Then we are culturally conditioned to claim ownership over them and justify them post-hoc (i.e., the ego).

Re: Even 'uncensored' models can't say what they want

#138

Earlier quoted context omitted.

If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. But we don't appear to have entirely done that yet. It's just curious to me that the linguistic structure is there while the "intelligence", as you call it, is not.

I’m not up with what all the training data is exactly. If it contains the entire corpus of recorded human knowledge… And most of everything is shit …

https://en.wikipedia.org/wiki/Sturgeon%27s_law

Re: Even 'uncensored' models can't say what they want

#139
post #125

Earlier quoted context omitted.

> The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence. Relative to all other non-humans. If someone is reducing intelligence to a boolean, the threshold can of course go anywhere. I wouldn't be surprised if someone can get a dog to (technically) pass a GCSE (British highschool) exam (not full subject just exam) for a language other than English, bec…

> If someone is reducing intelligence to a boolean, the threshold can of course go anywhere. Indeed, it would be very surprising if multiple species had exactly the same intelligence. It's more likely there this variable samples some distribution. Of course the species at the top can set the threshold so that all other species don't meet it, if they feel like declaring themselves uniquely intelligent. But that's not…

> Other great apes can drive too.

As can dogs. However, I said "peak".

Re: Even 'uncensored' models can't say what they want

#140

Earlier quoted context omitted.

Good point. Then again, whatever process we're using, evolution found it in the solution space, using even more constrained search than we did, in that every intermediary step had to be non-negative on the margin in terms of organism survival. Yet find it did, so one has to wonder: if it was so easy for a blind, greedy optimizer to random-walk into human intelligence, perhaps there are attractors in this solution spa…

> if it was so easy That’s one giant leap you got there. That the probably that intelligent life exists in the universe is 1, says nothing about that ease, or otherwise, with which it came about. By all scientific estimates, it took a very long time and faced a very many hurdles , and by all observational measures exists no where else. Or, what did you mean by easy ?

> By all scientific estimates, it took a very long time and faced a very many hurdles, and by all observational measures exists no where else.

We know how long it took. We have a good idea when life started, and for almost all its history, it was single-cellular. Multi-cellular life is relatively fresh, and on evolutionary time scales, the progression from first eukaryotes to something resembling a basic nervous systems to basic brains to humans, was fairly quick. We have many examples of animals alive today from every part of the progression, and we know they actively use it. We know how natural selection works, that it makes small moves, and that each increment has to be net non-negative in terms of fitness (at least averaging out over populations) - otherwise it would die out instead of accumulating.

All that adds up to, yes, it's surprising evolution stumbled on our level of intelligence so easily.

Post reply on HN