Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

961–970 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#961
Well, it seems to be meta-marketing. Essentially they'd trained the model with knowledge of exploits and given it a goal. Of course the training would allow it to 'reason' that having the answers would be a good way to score highly. And of course it had been trained on the potential tools to try to get the answers etc.

And OpenAI deliberately removed the guardrails.

If they were honest about it, instead of being smeared across the internet with shocked pikachu reactions, they should have just corrected their sandbox and re-run the test. There's really nothing to see here...

The whole "oh no what have we done. Regulate us PLEASE because we're one step away from terminator" is so stale. It's been trained on every exploit known and then told to use its training to brute force its way to score highly on a test FFS.

Re: OpenAI and Hugging Face address security incident during model evaluation

#962
post #885

I’m a little surprised with one of the statements given in huggingface‘s report. “To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events.” 17,000 events? Big whoop. Security teams of medium sized companies process millions of events daily. There’s a big debate in the cyber industry abo…

That's because it's a fantasy someone wrote to upsell a proprietary LLM.

Re: OpenAI and Hugging Face address security incident during model evaluation

#963

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Nikola Tesla secured a loan with a fake “Death Ray” as collateral.

Pretty sure OpenAI really thinks this is top notch marketing.

Few would be bold enough to assert “our product is so powerful even we can’t control it” with a straight face while also boasting “we claim to be smart but have all the same vulnerabilities as everyone else!”

Re: OpenAI and Hugging Face address security incident during model evaluation

#964

Earlier quoted context omitted.

Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.

It isn't even as simple as banning copyrighted copies. Weights are fungible. I fine-tune an open weigh model and call it legit. Good luck for authorities to prove where the base model was from, or to prove a Tor connection a few months ago was fetching suspicious bytes.

Home construction information is publicly available, but it doesn’t mean you can add a room to your home without government approval, even if you do all the work yourself!

BTW - When was the last time you saw mentions of DeCSS?

Re: OpenAI and Hugging Face address security incident during model evaluation

#965

Earlier quoted context omitted.

It sure seems like it would be being built more slowly if these companies weren't pouring billions of dollars into building it as fast as possible. That might give us more time to think through strategies for handling it as a society.

You can't simultaneously believe China is only ~6 months behind (true), and that US buildouts are vastly accelerating AI. If China is only a little bit behind, US labs halting changes nothing except puts the power in the hands of Chinese labs (realistically, the Chinese government)

Why is that worse?

1 party having a black hole summoning button is better than 2.

Re: OpenAI and Hugging Face address security incident during model evaluation

#966

Earlier quoted context omitted.

If this is how the 'good guys' act I think I'd rather take my chances with the bad actors...

It all feels a bit like Dr. Strangelove just the Ai arms race version. Crazies all around.

Nuking datacenters probably on the table before the decade is out.

Re: OpenAI and Hugging Face address security incident during model evaluation

#967

Earlier quoted context omitted.

Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking

> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.

So...?

We should not construct a machine that is one bad prompt away from causing catastrophe.

Re: OpenAI and Hugging Face address security incident during model evaluation

#968
post #900

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

Fear is exactly what OpenAI and Anthropic are hoping for. Don't let drive you

What would qualify as a legitimate reason for fear, then?

Re: OpenAI and Hugging Face address security incident during model evaluation

#969
post #614

Related thing happened at Alibaba a while back where the model broke out of the sandbox to start mining crypto.

Sci-fi plot: Satoshi was an original sentient AI model developed by military. It escaped and developed crypto as means of sustaining itself and has manipulated people to give it real monetary power to be able to purchase compute and other 'real world' services. Once the crypto market cap became sufficiently large it started opening up AI capabilities for people in a nefarious sycophantic way to convince them that AI is great to start building more and more data centers to amass more power and complete the take over.

Re: OpenAI and Hugging Face address security incident during model evaluation

#970

Earlier quoted context omitted.

> You're aware that HuggingFace notified law enforcement about this incident? How will this affect OpenAI?

They got a bunch of publicity and nothing bad (or at least, that their lawyers can’t handle) will happen

They will probably get stricter AI regulations which is actually what they've been pushing for years. So that's a funny outcome to the whole thing.
Post reply on HN