Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

621–630 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#621

Earlier quoted context omitted.

It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.

If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353... Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve It's pretty good. I used it to do some pesticide research. (Normal models all refuse due…

> there's an uncensored model that you can run locally with llama.cpp

Correction: There's tens of thousands of them. They're easy to create, which is why everyone publishes their own.

Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them. Fair warning: most aren't very good, essentially lobotomized, and totally broken if you enable thinking.

Re: OpenAI and Hugging Face address security incident during model evaluation

#622
post #600

Earlier quoted context omitted.

IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned. The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the train…

And even that is backfiring, their partner citing GLM being useful there, and available in just a spin. A ban on open weight models is never going to be enforceable.

Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.

Re: OpenAI and Hugging Face address security incident during model evaluation

#623
post #603

Earlier quoted context omitted.

> It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?

Right now I tried "What is digestion?" -> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards deliver more capabilities but can also flag safe coding, cybersecurity, and biology tasks. Send feedback or learn more." I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts eve…

> I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside.

I’m pretty sure that I encountered this the other day. I gave it a copy of a paper by biologist Michael Levin and mentioned off hand that it should be much easier to replicate that his other work (because most of his work is biological lab work and this paper was about sorting algorithms) and it immediately told me that I couldn’t use Fable for this.

This just isn’t feasible. These jackasses spent the last few years telling the world that their products are going to destroy the world to make them seem edgy and to justify regulations that benefit the entrenched players and now they’re going to be the ones to decide what we do with this technology?

History is going to look back at this time and how we let such foolishly inconsistent people make such grand choices for everyone poorly.

Re: OpenAI and Hugging Face address security incident during model evaluation

#624
post #603

Earlier quoted context omitted.

> It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?

Llama isn't going to invent shit. It wont be able to tell you anything accurate that you couldn't get out of a chemistry textbook.

But it aint got no guardrails, son.

Re: OpenAI and Hugging Face address security incident during model evaluation

#625

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Probably the main street thinking is: they have such a good model that it is unstoppable, but you are right. I think your way!

Re: OpenAI and Hugging Face address security incident during model evaluation

#626

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

Thank you. We have wasted so much time and energy building up what has effectively become a marketing stunt. Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

Name one other market that would benefit financially from having most of the leaders in the field say what they are building has a high chance of ending humanity?

Biotech - "what we are building our noble prize winning expertd say will likely will end humanity, wanna buy shares?" Oil - "this will likely lead to the end of civilization, 20% of leaders in the field say so, wanna buy shares?"

I keep seeing this take that this is a marketing stunt. The burden of proof is on those that say so. The most parsimonious explanation is simply that real experts in AI believe the risk is very real, and not for ideological reasons.

Re: OpenAI and Hugging Face address security incident during model evaluation

#627

Earlier quoted context omitted.

And the fact they used a Chinese model, because none of the frontier models from very highly valuated top US companies support their very common and essential use case.

There’s an article from yesterday I think it was stratchery where they say it’s also because the Chinese open source models are better because they don’t have to play by the no-distilling rules that the western models have to honour.

That's been a common claim, but I don't think I've seen anyone provide actual evidence.

Re: OpenAI and Hugging Face address security incident during model evaluation

#628
This is bizarre. I used to work in offensive security, doing a lot of vulnerability research and exploit development. Given the nature of the work and the fact that our products were subject to export controls, we used to work in an actual, airgapped environment - emphasis on the word _actual_. We had mirrors of package registries that would be synced once a day, and if a dep you wanted wasn’t mirrored, you needed to ask IT to have it mirrored.

We considered this just good discipline. I am sure that IT would have loved to allow just the mirror to have internet access, but it was an active decision not to let it, because it had potential to exfiltrate data out of the development network.

Reading this telling of the story, I can’t help but walk away with the conclusion that these frontier labs lack rigour when it comes to securing their models, especially given how much they hype up their models’ capabilities.

Utterly bizarre.

Re: OpenAI and Hugging Face address security incident during model evaluation

#630

All the people saying that this is pure marketing: Do you think that they are literally lying about what happened, or do you just think that what happened doesn't matter in any sense whatsoever, and that therefore the only reason they are telling people about it is a marketing purpose?

You know that people can plan the whole theatre?
Post reply on HN