Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

291–300 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#291
post #284

Earlier quoted context omitted.

This is marketing+. They will look for policy action here to try to capture tax payer dollars.

What incentive does HF have here?

HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.

Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#292
post #276

Earlier quoted context omitted.

It’s marketing the same way shitting your pants in public is marketing. People notice you.

Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.

Obviously shitting your pants in public shows you have a healthy digestive system and if you can demonstrate byproducts of wild food in your output, you’re approaching independent thinking and self-reliance.

This is how the financiers look at this and whatever you think it is right or wrong, it does showcase “capability”.

Re: OpenAI and Hugging Face address security incident during model evaluation

#293
Hopefully one of these agents isn't given a goal to fire the nukes (or, some goal that indirectly makes the model decide this is a way to meet it).

They are behind air gapped systems, but that didn't stop the US from hacking and Irans nuclear facilities, which they disabled using a virus.

Re: OpenAI and Hugging Face address security incident during model evaluation

#294

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

> Because we continue to have zero evidence that aligment is an actual risk.

I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.

These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.

[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.

[1] https://github.com/anthropics/claude-code/issues/73597>

Re: OpenAI and Hugging Face address security incident during model evaluation

#295
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

Is OpenAI truly behind? Just anecdotally I recently fully switched to using Codex at work because it feels a lot more competent

It's impossible to tell. Are they behind who? And on what?

It depends on who you ask. And everything is a vibe because all of this is new and things move fast. A week is a month in AI-land. A month; a year. A year? A decade.

On coding? I still like Fable better than Sol. But they're close enough that it probably is a vibe thing. Fable writes long commit messages, Sol writes commit messages like a college student in an elective computer class.

For API use, I'd say the Responses API that OpenAI architected is superior to Claude's Messages API. But again, I'm basing that off my vibes

Claude Design creates marketing imagery very effectively. GPT Image is the best imagegen model as ranked by users. Anthropic doesn't even have an imagegen model.

Anthropic definitely has compute scaling issues. OpenAI seems to have a pez dispenser that they click and out pops a GPU.

Anthropic's messaging is that they're building AI with guardrails but they've been banning people's accounts nonstop and their customer support is a lobotomized AI chatbot.

OpenAI has first mover advantage and to people not in tech, ChatGPT is synonymous with AI. But they also seem super sinister, like Uber circa 2015.

Or maybe I'm just suffering from AI psychosis. I have to go, my usage meter is about to reset.

Re: OpenAI and Hugging Face address security incident during model evaluation

#297

Earlier quoted context omitted.

1. it's not cheap to run glm-5.2 so not just anyone can do it 2. just because you haven't heard of attacks doesn't mean they haven't happened 3. this attack in the article was performed by a prerelease model which presumably benchmarks a bit above Sol which benchmarks above glm-5.2 We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possi…

GLM has an extremely cheap subscription plan similar to Claude Code from Z.ai. You get Opus-level quotas with 5.2 and none of the Anthropic-style model nerfs when you ask cybersecurity questions. It's extraordinarily, preeminently accessible to anyone that wants to use it for ill or good. > We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" co…

I don't know of a single zero day found on a number of tokens that fits inside a subscription plan. I'd be happy to be wrong.

Re: OpenAI and Hugging Face address security incident during model evaluation

#298
post #74

It seems like things are fairly amicable between OAI and HF, but what if they weren't? I'd love to see this kind of thing go to court. Who is responsible for the crimes of a "rogue" agent? How will they be punished? In this case it's unambiguous that OpenAI is the responsible party, but I can imagine a lot of adjacent scenarios where it's less obvious. And, where the impacts are much greater.

> Who is responsible for the crimes of a "rogue" agent? How will they be punished? Unironically this is why AI researchers have this fascination with the Talmud.

What? Can you explain a little more what you mean?

Re: OpenAI and Hugging Face address security incident during model evaluation

#299

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

If you seriously have this question, read "War with the Newts". Really do, make it your priority this week. If you did and this is a rhetoric question... Well, I do hope that if every single person on the planet would have read "War with the Newts" and made the right conclusions, maybe there would be a chance to change the course. But that's only because I choose to believe in miracles, otherwise I wouldn't know how to live.

(TL;DR: we won't.)

Post reply on HN