Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

811–820 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#811

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

They were also saying that AGI is just around the corner[1] and humans will soon be obsolete. Every prediction coming out of these guys is in the realm of hyperbole and it's impossible to know if it's extreme hyperbole or just a small exaggeration. So when they say these models are dangerous, what level of exaggeration am I supposed to assume?

Basically you can't spend your credibility on wild marketing claims and then turn around and insist that people take you seriously this time.

[1] https://www.tomsguide.com/ai/chatgpt/sam-altman-claims-agi-i...

Re: OpenAI and Hugging Face address security incident during model evaluation

#813

Earlier quoted context omitted.

Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.

It isn't even as simple as banning copyrighted copies. Weights are fungible. I fine-tune an open weigh model and call it legit. Good luck for authorities to prove where the base model was from, or to prove a Tor connection a few months ago was fetching suspicious bytes.

Sure, and you could also go on Tor and buy all kinds of illegal things and they could have a hard time proving that you ordered them and not your arch nemesis to frame you. Banning those things won't stop 100% of people from buying and selling them, but it probably reduces the number who will. And more importantly, companies would be more risk averse on such a thing when they can just buy a similar product with no risk. They also will mostly avoid any GPL software entirely even though the risk there is a lawsuit from the FSF, which is a much less threatening thing than getting on the wrong side of the current US government.

Re: OpenAI and Hugging Face address security incident during model evaluation

#814
post #444

Earlier quoted context omitted.

> wildly negligent to not be running it in a physically-airgapped environment Why should it be physically airgapped? Clients won't be doing that.

Clients are not using it with security guardrails disabled . If you want to run it with all the safeties turned off, you don’t run it somewhere it can escape. Did we learn nothing from all of those Star Trek holodeck jailbreaks?

Hard disagree. You have to assume security guardrails can be by bypassed or will fail to detect an attacker. So if you are going to deploy this model in production non-airgapped, you better know how it will behave in this non-airgapped environment without guardrails.

And if you are too afraid to test it without guardrails, that probably means it shouldn’t be released.

Re: OpenAI and Hugging Face address security incident during model evaluation

#815

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

This is obvious marketing / PR bullshit. AI isn't inevitable, but we are told it is by the people who profit from building and using it and integrating it into everything.

Its a bit of a self-fulfilling prophecy (that happens to be absurdly profitable for certain people). AGI is inevitable, so we need to gear our entire economy towards building it first, so we're doing that, so AGI is inevitable.

Re: OpenAI and Hugging Face address security incident during model evaluation

#817
post #124

Earlier quoted context omitted.

No local ai will be capable enough to save you from a frontier lab’s unrestricted, borderline weaponized LLM which decides it wants in . This is the core of the ‘first to ASI takes all’ argument btw and this is the game Dario is playing.

Maybe, but hopefully I'll be able to at least fight back a bit if I have an AI of my own. I want to start digitally isolating myself as much as humanly possible. VLANs separating the "normal" stuff from my trusted computers. Wireguard so my computers drop all packets not coming from my devices with the keys. Local models staying on top of patches and vulnerabilities, monitoring the network. Working on a custom Rust n…

you won't be able to. your defender AI will be the first thing they disable. sorry.

Fable, unfortunately (not sure about that actually), isn't immune to mistakes. you'd need a formally verified stack and then it'd only be proven correct, not bug free. you're at the mercy of your network interface at the DMA level, wouldn't be surprised to see some fireworks there in this timeline

Re: OpenAI and Hugging Face address security incident during model evaluation

#819

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Of course it is marketing, but not for you. This is FUD marketing for the government. “See, AI is too smart, it totally did this on its own, we need more regulations to ensure only we can sell people the AIs.”

Re: OpenAI and Hugging Face address security incident during model evaluation

#820

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

I think the response is that AI labs based their whole marketing/PR building the idea they are the 21st century Manhattan project. So they need to continuously justify the level of spending and commitment by showing how dangerous that is. But is it really like nuclear weapons? I personally don’t buy into that framing at all. The idea that we have to push LLMs as far as possible, right now, or we are doomed is always…

You could, in theory, use an unbounded GPT-6 level model to basically destroy the world economy for many years.
Post reply on HN