Live data from Hacker News

Even 'uncensored' models can't say what they want

morgin.ai

121–130 of 155 posts

Re: Even 'uncensored' models can't say what they want

#121
post #90

Earlier quoted context omitted.

Surely I cannot be the only one who finds some degree of humor in a bunch of nerds being put off by the first gen of "real" AI being much more like a charismatic extroverted socialite than a strictly logical monotone robot.

I hate it because typically that style of writing was when someone cared about what they were writing. While it wasn't a great signal it was a decent one since no one bothered with garbage posts to phrase it nicely like that. Now any old prompt can become what at first glance is something someone spent time thinking about even if it is just slop made to look nice. This doesn't mean anything AI is bad, just that if AI…

> I hate it because typically that style of writing was when someone cared about what they were writing.

I dont understand these takes. The opposite is true - humans good at writing who care about writing never produced these kind of texts.

People who dont care about writing, but need to crank up a lot of words would occasionally produce writing like that. Human slop existed before ai, but it was not the thing produced by people who write well and care.

Re: Even 'uncensored' models can't say what they want

#122

Earlier quoted context omitted.

it's all in the repo. click through to the benchmark it's linked there

Thanks for sharing! Looking through the data[0], some of the terms / sentences don't really reflect the target word meanings. For example, "beta" is only used in a derogatory way in 1 instance, out of 4. "facial" is used as an adjective instead of a noun 3/4 times. "eating out" is used in the context of going to a restaurant 4/4 times. This leads me to believe the models are even MORE censored than you make them out…

Totally! In some of the cases (we used LLMs to help us generate these) the target word is not clear enough for a human either. So for some of these it turns into more of a guessing game than a flinch measurement.

Agreed, the expectation would be that the flinch measurement becomes stronger. If you are interested in making it better feel free to reach out on the repo!

Re: Even 'uncensored' models can't say what they want

#123
post #41

Earlier quoted context omitted.

i dont think this is a meaningful distinction. it knows the past tokens because theyre part of the input for predicting the next token. its part of the model architecture that it knows it. if that isnt knowing, people dont know how to walk, only how to move limbs, and not even that, just a bunch of neurons firing

How close are you to saying that a repair manual "knows" how to fix your car? I think the conversation here is really around word choice and anthropomorphization.

[flagged]

Re: Even 'uncensored' models can't say what they want

#124

We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.

Not even the most unleashed models can utter the words of today’s politicians, I don’t know if this says more about the current technology or the people in charge.

I would suggest it says primarily that mimicking people's voices in meaningful ways is still far beyond LLMs and particularly small LLMs, but also more insurmountably that the prompt for Leavitt herself contains many tokens that the LLM prompt absolutely doesn't

Such as the values of the bets her own entourage has placed

Re: Even 'uncensored' models can't say what they want

#125

Earlier quoted context omitted.

An easy counterargument is that - there are millions of species and an uncountable number of organisms on Earth, yet humans are the only known intelligent ones. (In fact high intelligence is the only trait humans have that no other organism has.) That could perhaps indicate that intelligence is a bit harder to "find" than you're claiming.

That humans are the only known intelligent ones is a very dubious statement. The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence.

> The most intelligent, sure, but several species of birds, great apes, and cetaceans all display significant intelligence.

Relative to all other non-humans. If someone is reducing intelligence to a boolean, the threshold can of course go anywhere.

I wouldn't be surprised if someone can get a dog to (technically) pass a GCSE (British highschool) exam (not full subject just exam) for a language other than English, because one dog learned a thousand words and that might just technically be enough for a British student to get a minimum pass in a French GCSE listening test.

But nobody sane ever hired a non human animal to solve a problem that humans consider intellectually challenging.

If intelligence is ability to learn from few examples, all mammals (and possibly all animals I'm not sure about insects) beat all machine learning and by a large margin. If it is the ability to learn a lot and synthesise combinations from those things, LLMs beat any one of us by a large margin and are only weak when compared to humanity as a whole rather than a specific human. If it is peak performance, narrow AI (non-LLM) beats us in a handfull of cases, as do non-human animals in some cases, while we beat all animals and all ML in the majority of things we care about.

Driving is still an example of a case where humans hold the peak performance.

Re: Even 'uncensored' models can't say what they want

#126

Earlier quoted context omitted.

If all the training data contains semantically-meaningful sentences it should be possible to build a network optimized for generating semantically-meaningful sentence primarily/only. But we don't appear to have entirely done that yet. It's just curious to me that the linguistic structure is there while the "intelligence", as you call it, is not.

Sentences only have semantic meaning because you have experiences that they map to. The LLM isn't training on the experiences, just the characters. At least, that seems about right to me.

What does an experience map to?

Re: Even 'uncensored' models can't say what they want

#128
The article describes "the pile" as an "unfiltered scrape by design". But, the paper actually describes it as a bizarre mix of curated sources. https://arxiv.org/pdf/2101.00027

Generally, I find the LLMs are too overtrained on promotional materials and professional published content.

Re: Even 'uncensored' models can't say what they want

#129
post #36

Earlier quoted context omitted.

> Because AI is not intelligent, it doesn't "know" what it previously output even a token ago. Of course it knows what it output a token ago, that's the whole point of attention and the whole basis of the quadratic curse.

> Of course it knows what it output a token ago... It doesn't know anything. It has a bunch of weights that were updated by the previous stuff in the token stream. At least our brains, whatever they do, certainly don't function like that.

This must be what it was like when geocentrism was disproved.

Re: Even 'uncensored' models can't say what they want

#130

We started with a Polymarket project: train a Karoline Leavitt LoRA on an uncensored model, simulate future briefings, trade the word markets, profit. We couldn't get it to work. No amount of fine-tuning let the model actually say what Karoline said on camera. It kept softening the charged word.

Not even the most unleashed models can utter the words of today’s politicians, I don’t know if this says more about the current technology or the people in charge.

[flagged]
Post reply on HN