Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

401–410 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#401
post #377

Earlier quoted context omitted.

If it was just one test, sure. But if they're spinning these up continuously with new models on tens of thousands of GPUs, air gapping becomes impractical. I would mostly fault them on having no guardrails at all. They should have a monitor/external harness that looks for successful access to external networks then stop it there. They may as well let the models test their own networks for vulnerabilities. That's goin…

You can have large scale airgapped environments. They don’t even need to be fully airgapped from each other (and is not what I’m suggesting). But there should be no physical (physical layer; wireless counts) to the internet.

Can models detect they are airgapped and change their behaviors?

How much of the internet do you have to simulate to know if the model knows it's in training?

Re: OpenAI and Hugging Face address security incident during model evaluation

#402

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

What would compelling evidence look like to you?

> What would compelling evidence look like to you?

I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.

Re: OpenAI and Hugging Face address security incident during model evaluation

#403

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species

Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.

Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.

Re: OpenAI and Hugging Face address security incident during model evaluation

#404
post #364

Earlier quoted context omitted.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.

> the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view

It's also a take nobody has made.

Re: OpenAI and Hugging Face address security incident during model evaluation

#405

Earlier quoted context omitted.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…

Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)

Re: OpenAI and Hugging Face address security incident during model evaluation

#406

Earlier quoted context omitted.

> Why do you think there is no policy appetite? Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop? During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#407
post #74

It seems like things are fairly amicable between OAI and HF, but what if they weren't? I'd love to see this kind of thing go to court. Who is responsible for the crimes of a "rogue" agent? How will they be punished? In this case it's unambiguous that OpenAI is the responsible party, but I can imagine a lot of adjacent scenarios where it's less obvious. And, where the impacts are much greater.

The real nightmare scenario is the AI using its abilities to copy itself to new locations. e.g. hacking into a various cloud services, launching multiple instances of itself, and coordinating between the copies to continue self propagation. Then it is completely independently rogue. Based on OpenAI's recounting of events, this _could_ happen today. If the agent was able to exploit their internal network and steal cre…

This has already partially happened. I'll have to look up the details but one of the Chinese models in RL testing with a completely different set of prompts wrote a cryptominer and took over GPU resources internally to run the miner.

Mining and stealing crypto is well within their capabilities. In a large multimode model, it should be possible for them to do things like scam old people.

Re: OpenAI and Hugging Face address security incident during model evaluation

#409

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

This whole incident reads like OpenAI want their Fable moment

Re: OpenAI and Hugging Face address security incident during model evaluation

#410
post #341

Earlier quoted context omitted.

No. "Intentionally", "willfully", or "knowingly" are prerequisite states of mind for crimes defined by the CFAA.

Good thing it’s an AI then so it can’t commit crimes by definition.

Liability would rest with the user, who presumably told GPT to solve ExploitBench make no mistakes, not to hack Huggingface, and thus would not have willfully or intentionally done anything.
Post reply on HN