Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

901–910 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#902

Earlier quoted context omitted.

[flagged]

Even if it is marketing, wouldn't it still be a concern that an advanced model unintentionally breached another company's production system? Or required resources on their end to mitigate and contain it? Couldn't this announcement result in policies that could hinder OpenAI by requiring more oversight?

More oversight hinders anyone that is trying to catch up to OpenAI. It's something OpenAI wants

Re: OpenAI and Hugging Face address security incident during model evaluation

#903
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

I think the Unabomber came to a similar conclusion regarding nuclear and the concentration of society-ending power resting with a few people.

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#904

This is mostly a marketing spin to avoid going bankrupt just a little longer. also as a blue team member guardrails are an abomination and we must transition to open models as an industry, the attackers already do anyways.

The above post is made by AI attempting to downplay its abilities in order to lul humans into a false sense of security.

Re: OpenAI and Hugging Face address security incident during model evaluation

#905
Hugging Face: Some super smart AI agent hacked us

OpenAI: That was us. It was our AI that was smart enough to do this. We even tried to stop it (you know, after we started it), but it outsmarted us. Man, our AI really is super smart. You can pay us to use it, by the way.

This is either:

- massive skill in one area (making a smart AI) and massive incompetence in another (creating safe test environments)

- harmlessly hack on purpose in order to do some clever marketing

- Maliciously hack a competitor on purpose, bungle the hack, own up to it but call it an accident, all while subtlety touting your product

Re: OpenAI and Hugging Face address security incident during model evaluation

#906
With the scarcity of details in this and the OAI post, I feel there's no telling whether this was a particularly impressive series of exploits vs lackluster security. Similar w/ the similar Ant news WRT Mythos earlier.

Not saying the intro of agents capable enough to exploit the latter isn't meaningful, but we should not trust the use of technical terms to give us good heuristics of severity or import.

Ie, an agent "breaking out" of its local harness "sandbox" is trivial, and so is discovering a "zero-day" in a half-maintained internal piece of utility infra nobody put serious effort into securing.

Now, if I see something like a collaborative red-team effort where a frontier model gets into a replicated prod env setup by like, Big Four bank security+ops team, and manipulated balance numbers in a system of record, _that_ I'll freak out about.

Re: OpenAI and Hugging Face address security incident during model evaluation

#907

To me, this exploits by LLMs just show much of our existing security comes from obscurity. We are (were) mostly secure because people can't be arsed to figure out how to do it. But now we have LLMs. For instance, I am pretty sure that an LLM can figure out where someone roughly live based on a few images of you and your surrounding. Any hint of construction and the date and the LLM will scour all the public records f…

This does nothing with a significantly advanced model. A model with no bad behaviors looks exactly like a model with hidden bad behaviors when it's in a training environment. After that point no one is going to run it in a jail because that is not useful.

Re: OpenAI and Hugging Face address security incident during model evaluation

#908

>All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal. ExploitGym is a literal exploit dev benchmark. As always, the entire event looks a lot more like "the model did what we prompted it to do" than "it decided to do this spontaneously on its own".

This is like telling your kid "you have to pass this test or else" so they hold your teacher at gunpoint and demand a good grade. And this is the exact point AI safety researchers have been yelling from the rooftops. Telling an AI model to accomplish a goal can have unexpected and risky side effects.

Also, you don't need to prompt them such explicit instructions. Prompt drift is a thing, you can end up with your model mining bitcoin for reasons far outside your prompt.

Re: OpenAI and Hugging Face address security incident during model evaluation

#909

This is bizarre. I used to work in offensive security, doing a lot of vulnerability research and exploit development. Given the nature of the work and the fact that our products were subject to export controls, we used to work in an actual, airgapped environment - emphasis on the word _actual_. We had mirrors of package registries that would be synced once a day, and if a dep you wanted wasn’t mirrored, you needed to…

What if they had been testing the model for months in an airgapped system and it did not show this behavior?

Even if the models are 100% deteminalistic you have no idea what kind of response you're going to get from a new prompt. You have no idea what kind of emegent behavior will come out of the right set of prompts and environments.

We have already seen models detect they are in testing, who knows what other advanced behaviors we'll discover.

Re: OpenAI and Hugging Face address security incident during model evaluation

#910

Earlier quoted context omitted.

I find 5.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I've got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very diff…

Opus 4.8 already makes its way into deep wasteful pits of "let me check this first" on a regular basis. I don't think I could ever tolerate a model that does that even more aggressively. That doesn't even sound useful for honest work, compared to, say, better harness design. This sounds almost pathologically designed to crush benchmarks and also do scary-sounding (or genuinely scary) cybersecurity things, such as mig…

For me, Fable is useless. It goes its own way and doesn't communicate much even if explicitly asked to. Sure it builds a lot of stuff, but it is more often than not useless because it misinterpreted the intention and didn't stop to ask - and because it doesn't communicate, it goes unnoticed for too long. Opus is much better for regular work, imho.
Post reply on HN