Live data from Hacker News

Responding to the next frontier of critical cyber capabilities

openai.com

81–90 of 207 posts

Re: Responding to the next frontier of critical cyber capabilities

#81

Every fifth comment about our insane trajectory of AGI is about "marketing." These incidents and cybersecurity capabilities are now involving government hearings and the CIA. Denial is truly an incredible thing in the face of a very scary immediate future.

The gullibility of AGI-pilled folks regarding these "hacks" is just breathtaking. When OAI demonstrates these dangerous capabilities live in a public environment where security experts can see and verify what actually happened, then reasonable people can have reasonable discussions about the level of danger. This is a very low evidence bar. Right now you are running in circles yelling "the sky(net) is falling" based…

Did you watch the Black Hat defcon talk?

Re: Responding to the next frontier of critical cyber capabilities

#82

Earlier quoted context omitted.

CFAA laws do not require “Damage” to be done.

They happen to require intent and are thus irrelevant here.

I think you’ll find that negligence is indeed accounted for.

Re: Responding to the next frontier of critical cyber capabilities

#83

Earlier quoted context omitted.

You do a KYC and you can get access. It may depend on country's quality of KYC.

I thought you need to prove you are working in cybersecurity or provide evidence of authorization for work done. It's really just simple ID/face verification?

For OpenAI's Cyber verification (the normal kind for Codex) you absolutely do not need any proof of cybersecurity work/authorization. They just use Persona for KYC + live selfie, and some extra checks that I don't know the nature of (but not related to checking whether you're a cybersecurity professional).

Re: Responding to the next frontier of critical cyber capabilities

#84
post #11
post #6

Earlier quoted context omitted.

They actually did a detailed presentation at BlackHat about the HuggingFace incident, and events that led to it. https://youtube.com/watch?v=87DyyMV0kCY

That was fascinating. Hijacking the package manager to pass messages between models and agents.. that's next level. Like "pssst, if you need internet access there's a vulnerability in x service" kind of messages

You've heard of 4chan for AIs, but did you hear of secret frontier lab AI hacker BBS?

Re: Responding to the next frontier of critical cyber capabilities

#85

Earlier quoted context omitted.

You don't need an invite only program to just have Sol checking for vulnerabilities in binaries or code. But yeah I've hit guardrails a few times when Sol was making PoCs for the vulnerabilities it found (but most of the time it made those PoCs without issues).

Opus refused to help me try to develop an exploit to export data from an old Android device where I can't upgrade to latest android and I couldn't use the app's backups (because I couldn't update the app.) Not sure where that lies in the "binaries or code" spectrum.

If you don't want to get cyber verified, you can try an open-weight model such as Kimi K3 for that, it has looser internal safety training.

Re: Responding to the next frontier of critical cyber capabilities

#86
post #45

Earlier quoted context omitted.

So they found their agents had RCE'd Artifactory once, reported it and got the fix, continued using Artifactory for their sandbox, and left it unmonitored for days despite the earlier exploits? They really do come out looking totally incompetent. I stress about my agent sandboxes all the time and the only models I run have the default heavy handed guardrails, and I don't leave them running persistently. Edit: not to…

I get that it’s fashionable to hate big companies but you’re working overtime here. It’s reasonable to assume that a bug was fixed when reported. And if you think your monitoring is 100%, you don’t know what you’re talking about. If you consider that incompetence, it’s possible that you’re not a very nice person.

If your CEO is going around talking about how your product will "most likely lead to the end of the world", people are right to expect you to be pretty careful in what you're doing. OpenAI allowed bidirectional communication across security domains for over a month before discovery. Even after it was discovered (and not completely fixed), they didn't set up monitoring able to detect attacks against internal or external services, which went on for further weeks.

Re: Responding to the next frontier of critical cyber capabilities

#87

Earlier quoted context omitted.

You don't need an invite only program to just have Sol checking for vulnerabilities in binaries or code. But yeah I've hit guardrails a few times when Sol was making PoCs for the vulnerabilities it found (but most of the time it made those PoCs without issues).

I maintain a version of an app called Rewind because the company behind it went under after implementing a killswitch. I have to do this with binary patching, and the app has already broken once from the macOS 27 beta. Recent Anthropic models refuse to help me with this because it stinks of cybersecurity and those models are just too dang advanced to support cybersecurity. I'm maintaining a piece of software to which…

I had not hit any guardrails with reverse engineering (as long as its not related to security) with GPT 5.5 or 5.6 Sol, so you should try those. At worst with Sol you might see the warning in Codex that your request is being checked for security so it'll take more time, that's just a warning, not a full stop.

Re: Responding to the next frontier of critical cyber capabilities

#88

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch. tl;dw; - agents found a way to communicate between several instances during a training run…

A interesting talk, interesting times. But their proposed solution to AI offense outpacing human defense... is more AI? The plot is getting a bit unrealistic, the characters are lacking genre-savviness.

Re: Responding to the next frontier of critical cyber capabilities

#90

Earlier quoted context omitted.

If they did any damage that would be a reasonable argument. As far as I am aware, nothing bad happened.

Regardless of exact practical outcome, it is deeply irresponsible and reckless behavior to run such security testing on other's infrastructure and without sufficient isolation. If they actually believe their models to be as powerful as the marketing says, then anything less than airgapping for such a "do anything to get the results" evaluation clearly isn't acceptable. If the fire department suddenly had practice fir…

[deleted]
Post reply on HN