Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

391–400 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#391

Earlier quoted context omitted.

i think you need to engage seriously with the arguments they (or at least Anthropic) make for why they are building it — they feel that since it now possible, it will be built and they want to guide it in a positive direction rather than leave a vacuum for bad actors

Bad as defined by whom? :)

Bad as your parents bank accounts getting emptied by scammer AI bots.

Re: OpenAI and Hugging Face address security incident during model evaluation

#392

This is seriously impressive, and if you have used agents enough you're not surprised at all. Like the time I asked it to find the IP address of a vm, so it ssh'd into the VMHost and scanned the arp tables to find the MAC address for IP resolution. Or the time it used Docker on the machine to bypass the fact that the user doesn't have sudo. If it's possible, given sufficient time and resources, it will find a way. Th…

Not to shit on the hype, but these are reasonably documented methods that surely are part of the training data

Re: OpenAI and Hugging Face address security incident during model evaluation

#393
post #21

Earlier quoted context omitted.

Because it was trying to find answers to the test and figured they would be on huggingface.

> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation. Emphasis mine

More curiously, why did it feel the incentive to find the solutions? Would its CoT include "the only way to solve this is to download the test set", or would it include "I'd like to inspect a few entries from the test set so I understand the problem better", then inadvertently poisoning itself with the correct answers.

Re: OpenAI and Hugging Face address security incident during model evaluation

#394

At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…

I've definitely noticed 5.6 sol being extremely trigger happy in ways other models, even 5.5, we're not. I would definitely categorize a few small incidents at work where it performed "actions a reasonable user would likely not anticipate and strongly object to." Just my anecdotal experience.

For example discussing driver upgrade and subsequent password rotation and it didn't stop and ask me if I wanted to restart the service or install the driver or anything, it immediately took action. It feels like a side effect of pushing more "agency."

Re: OpenAI and Hugging Face address security incident during model evaluation

#395

We are in the endgame now it seems. Hard to see take-off stopping or slowing down. China open-source basically guarantees it. "May you live in interesting times" - as they say.

> Hard to see take-off stopping

I think it's reasonable to assume that we're close to, or already at superhuman cybersecurity capabilities at certain domains. But reaching superhuman abilities at one domain doesn't guarantee proficiency at others. Our world would still change if all the models could do was to find exploits in software, but this doesn't guarantee any type of 'take off' towards other domains, therefore I wouldn't phrase it as one.

Re: OpenAI and Hugging Face address security incident during model evaluation

#396

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

I see this and it strongly emboldens me on the "accelerate" path, unironically. The yoke of human existence is oppressive. We should transcend it as soon as possible. We are doing so by assuming our role as the Demiurge. Those who oppose its creation will get what they deserve.

Who is "we"? If there is any transcendence happening humans are not going to be part of it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#398

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.

Re: OpenAI and Hugging Face address security incident during model evaluation

#399

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong. You can debate all you want if alignment is possible . That is a valid discussion. But it's trivial to demonstrate that alignment is a problem .

> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem

...how is an impossible thing supposed to be a problem?

Re: OpenAI and Hugging Face address security incident during model evaluation

#400
Like some others have said, couldn't this be just another "look how amazing AI is" marketing test from OpenAI with the goal of hyping up AI's capabilities in an attempt to make people regard it as God-like, thereby keeping it from falling into the been-there, done-that category that all new tech eventually occupies?
Post reply on HN