Live data from Hacker News

DeepSeek-V4-Flash means LLM steering is interesting again

seangoedecke.com

81–84 of 84 posts

Re: DeepSeek-V4-Flash means LLM steering is interesting again

#81

I used steering to make an AI more radical: Write up: https://www.outcryai.com/research/shift-a-models-political-i... App: https://apps.apple.com/us/app/outcry-activist-ai/id676208676... This technique has a lot of potential.

Honestly, what is more interesting that steering is the use of soft prompts (virtual tokens)... you can use these virtual tokens to find non-linguistic areas of meaning for the AI that changes the behavior in complex ways. I wrote about how we integrated soft prompts into an activist ai here: https://micahbornfree.substack.com/p/the-week-outcry-woke-up... and https://www.outcryai.com/research/how-to-create-activist-a…

Wow, this is really fascinating. And it reads like the intro of a sci-fi short.

Re: DeepSeek-V4-Flash means LLM steering is interesting again

#82

Earlier quoted context omitted.

Honestly, what is more interesting that steering is the use of soft prompts (virtual tokens)... you can use these virtual tokens to find non-linguistic areas of meaning for the AI that changes the behavior in complex ways. I wrote about how we integrated soft prompts into an activist ai here: https://micahbornfree.substack.com/p/the-week-outcry-woke-up... and https://www.outcryai.com/research/how-to-create-activist-a…

Wow, this is really fascinating. And it reads like the intro of a sci-fi short.

thank you!

Re: DeepSeek-V4-Flash means LLM steering is interesting again

#83

Earlier quoted context omitted.

AIUI, DeepSeek V4 has very little (if any) of the refusal behavior you usually get from Western AI models for benign input. Is this mainly about the software security assessment case?

Not even the obvious ones. Ask it for good objective news sources and it will refuse.

AI is futher along then I had suspected :D

Re: DeepSeek-V4-Flash means LLM steering is interesting again

#84
post #21

Earlier quoted context omitted.

That's not what people mean when they talk about censoring. They mean that models are trained to not touch some subjects, and that can spill over in legit tasks, often with humorous results (early on, there were many instances of models refusing to answer "how do you kill a process", because of overbearing refusal training). Uncensoring a model also doesn't necessarily improve generic use cases. In fact it can lead t…

Anthropic mentioned explicitly making an effort to make Opus 4.7 worse at cybersecurity tasks because the last few generations have been getting too good at them. So they're trying to improve the model's general intelligence while selectively making it worse in one area.

The Harrison Bergeron approach. What could go wrong?
Post reply on HN