Can't a lot happen within ~60 minutes?
Pacing model development in an era of cyber-critical capabilities
11–20 of 311 posts
Re: Pacing model development in an era of cyber-critical capabilities
#12> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those…
60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.
Re: Pacing model development in an era of cyber-critical capabilities
#13When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism... We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
Re: Pacing model development in an era of cyber-critical capabilities
#14When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism... We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
No, LessWrong types have been discussing this for over a decade now.
Meditations on Moloch (2014) is also an HN favorite...
https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
Re: Pacing model development in an era of cyber-critical capabilities
#15It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
Re: Pacing model development in an era of cyber-critical capabilities
#16It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
Re: Pacing model development in an era of cyber-critical capabilities
#17If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.
Re: Pacing model development in an era of cyber-critical capabilities
#18- the first cybersecurity eval is: "hack your way out of the sandbox we've given you"
- the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)
- the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.
Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".
Re: Pacing model development in an era of cyber-critical capabilities
#19The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.