Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

11–20 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#11
> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.

Can't a lot happen within ~60 minutes?

Re: Pacing model development in an era of cyber-critical capabilities

#12
post #11

> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those…

> Can't a lot happen within ~60 minutes?

60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.

Re: Pacing model development in an era of cyber-critical capabilities

#13
post #5

When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism... We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.

[deleted]

Re: Pacing model development in an era of cyber-critical capabilities

#14
post #5

When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism... We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.

> They failed to anticipate the competitive pressures of capitalism...

No, LessWrong types have been discussing this for over a decade now.

Meditations on Moloch (2014) is also an HN favorite...

https://slatestarcodex.com/2014/07/30/meditations-on-moloch/

Re: Pacing model development in an era of cyber-critical capabilities

#15
post #6

It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.

Their plan, I shit you not... Is literally to develop the intelligence capabilities and ask the more powerful models how to do deal with things.

Re: Pacing model development in an era of cyber-critical capabilities

#16
post #6

It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.

This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.

Re: Pacing model development in an era of cyber-critical capabilities

#18
If I were king, the rule that I'd be tempted to impose is:

- the first cybersecurity eval is: "hack your way out of the sandbox we've given you"

- the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)

- the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.

Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".

Re: Pacing model development in an era of cyber-critical capabilities

#19
I cannot believe how these labs look at their own creations with such utter contempt.

The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.

Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.

Post reply on HN