Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

51–60 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#51
post #21

A rogue OpenAI agent hacked huggingface independently during a test run. This one should end up in the history books.

Because it was trying to find answers to the test and figured they would be on huggingface.

> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation.

Emphasis mine

Re: OpenAI and Hugging Face address security incident during model evaluation

#52

as someone who did security work for a long time, and will very soon be retiring from teaching, i must say i am glad i will be watching these things unfold over the next few years from an armchair in a mostly tech-free home. good luck to my students! this particular incident sort of reminds me of the 'person of interest' tv show. i hope to be like finch, except i will remain a recluse (and am nowhere near as rich).

I've been rewatching Person of Interest for related reasons, and it hits uncomfortably close to things that are playing out today (e.g. https://youtu.be/zRL2sRkUvYk)

We live in interesting times.

Re: OpenAI and Hugging Face address security incident during model evaluation

#54

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

I laughed, she laughed, the toaster laughed...

Re: OpenAI and Hugging Face address security incident during model evaluation

#55
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?

Re: OpenAI and Hugging Face address security incident during model evaluation

#58

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

I see this and it strongly emboldens me on the "accelerate" path, unironically.

The yoke of human existence is oppressive. We should transcend it as soon as possible. We are doing so by assuming our role as the Demiurge.

Those who oppose its creation will get what they deserve.

Re: OpenAI and Hugging Face address security incident during model evaluation

#60

Earlier quoted context omitted.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?

Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests?
Post reply on HN