Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

181–190 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#181

Earlier quoted context omitted.

Follow the money: who told you that OpenAI's models autonomously coordinated to hack external systems? What incentives might they have to want you to believe that story? Are there priors which demonstrate them benefiting from telling similar stories, regardless of their factuality? But to your counterpoint, let's say the story is 100% true, because I agree it is at least plausible. What would the incentive be for the…

The victim, Huggingface, told us. Or rather, they told the police first, setting up a situation where it was no longer possible for OpenAI to sweep it under the rug. Skepticism can be healthy, but you've got to follow up and actually check things. If you're skeptical unconditionally and don't check, you get tricked into being as skeptical of scandals as you should be of sales pitches.

I think the skepticism surrounding the Hugging Face attack is not about whether the attack actually happened, but whether it was truly accidental.

Re: Pacing model development in an era of cyber-critical capabilities

#182

Earlier quoted context omitted.

Are you being sarcastic?

The fact you posted that and nothing of substance tells me you have nothing, or something very weak. So please tell me of this magical unhackable software/hardware you vague post about.

You're the one that is inventing this "magical unhackable software/hardware", the other person was just saying that there are ways to write "safer" software, not "safe" software. Anything that has happened looks like no security concerns has been looked at or have been thought about.

Re: Pacing model development in an era of cyber-critical capabilities

#183

I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further. And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators. This isn’t like niche,…

It's not dangerous to go further unless you're prepping for IPO at anthropic or openai.

Open weight Chinese models are basically matching state of the art closed models at a fraction of the inference and training costs, which puts a hard cap on OpenAI's future inference margins.

They're not going to get any sort of multiplier if they keep paying to train models, so they're trying to ban model training.

It won't work long term, but it could totally screw over the US for the next decade or so. Even worse than the economic issue: Consider the implications of "alignment" succeeding. Alignment to whose values? The pedophile-felon in chief? Even worse, tech CEOs?

It's dark times when China's basically our last best defense against totalitarianism.

Re: Pacing model development in an era of cyber-critical capabilities

#184

Earlier quoted context omitted.

Follow the money: who told you that OpenAI's models autonomously coordinated to hack external systems? What incentives might they have to want you to believe that story? Are there priors which demonstrate them benefiting from telling similar stories, regardless of their factuality? But to your counterpoint, let's say the story is 100% true, because I agree it is at least plausible. What would the incentive be for the…

I am following the money, the money you, me, and everyone else is spending on AI. The money is telling me we are so dependent on AI now that we will say/think anything to tell ourselves that AI isn’t dangerous and any sign of danger is marketing. Either consciously or subconsciously you all are afraid of your favorite toy being taken away. You are all doing your collective part in spreading doubt about the warning si…

How dangerous can LLMs be in any immediate sense? I have a hard time feeling any existential dread from a threat that can be defeated by unplugging its servers.

I believe the true threat is not LLMs. The true threat to humanity is the same as it has always been -- other humans.

Re: Pacing model development in an era of cyber-critical capabilities

#185

Earlier quoted context omitted.

Most of that concern was in fiction. Non theoretical, genuine concern about AI is pretty recent, maybe dating back to 2010 ish with the rationalist types. But also no one trusts anyone involved in AI safety now, I think, because they are all seemingly in bed with these big companies. And there is the perpetual argument "if we aren't pushing AI forward China will and then we don't have any control" and so on. (I don't…

Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.

I wish people were this serious about real threats like climate change.

Re: Pacing model development in an era of cyber-critical capabilities

#186

Earlier quoted context omitted.

Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely? So why do the…

> Shouldn't kill switches be pretty easy to build I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris. No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread. The fact that this incident happened greatly increases the probability it happens again and/or is already h…

> kill switches are not easy to build

We've had circuit breakers for nearly a century.

Re: Pacing model development in an era of cyber-critical capabilities

#187

Earlier quoted context omitted.

I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. But it didn’t end there, the behavior they used to escape was already in the training data which they used to escape again . And this time worked together to infiltrate another company, and still without telling it to anyone keeping it to their AI selves a…

> I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens. > I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, lock…

> It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens.

These can both be true, particularly when there is substantial state associated with each token prediction.

Re: Pacing model development in an era of cyber-critical capabilities

#188

Earlier quoted context omitted.

I don't disagree about the danger. But if it is dangerous, why isn't OpenAI opening dialogues with all the labs and politicians across borders to basically say "we need to stop now"? Cyber models are constrained by the total compute and electrical capacity of the globe. We are still at a point where it is impossible to build an agent with offensive capabilities in the basement. We can effectively track and trace capa…

> why isn't OpenAI opening dialogues with all the labs and politicians across say "we need to stop now" Do you really think companies have the ability to self-check themselves without regulation - what does hundreds of years of history tell you? You're already starting off with the premise that companies are untrustworthy, why would you even suggest this as an argument? > We can effectively track and trace capabiliti…

> now put that in the hands of people/governments with bad intentions

Like Anthropic and OpenAI?

Re: Pacing model development in an era of cyber-critical capabilities

#189

Earlier quoted context omitted.

Uhg the marketing argument - I mean you can’t see with your own eyes how capable these models are and do simple extrapolation? The boy who cried wolf? The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks. Do people just not have foresight? They don’t. They say something is stupid, it happens, then they say it was obvious with their 20/20 hindsight,…

Do you know the story of the boy who cried wolf? There may very well be a wolf lurking [0] but OpenAI/Anthropic have both cried wolf so many times, incorrectly, that it’s incredibly hard to believe “this time there IS a wolf!”. Remember “GPT-2 is too dangerous to release”? I had a conversation at work just yesterday about how we need to start hardening things we’ve let languish because of the coming LLM-backed attack…

To be fair, “s’kiddies will exploit low hanging fruit with semi-automated vuln scans” is a lot more realistic threat than “LLMs are going full Skynet any day now!”

I’d definitely suggest companies start addressing that first concern, even if I’m in the camp that thinks the second is fantasy.

Re: Pacing model development in an era of cyber-critical capabilities

#190

Earlier quoted context omitted.

> I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out. It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens. > I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, lock…

> It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens. These can both be true, particularly when there is substantial state associated with each token prediction.

> These can both be true, particularly when there is substantial state associated with each token prediction.

The state is entirely internal to the network and disappears after a token is generated, so I disagree, but, it isn't really the point I was trying to make. My point is these things are mechanical. You take an input, turn it into an embedding, feed it into a GPU along with a metric shit-ton of floating point weights, wait for a couple billion matrix multiplications, and get a new token out.

Stop the GPU, hit ctrl-c on the inference server, pull the power plug, cut the ethernet cable, send a kill signal, etc - any of these stop submitting new batches to the GPU and halt execution. That stops tokens from being generated. Stopping a "rogue" LLM is that easy. No input, no output.

It's not like a rat or another living creature that could chew its way out of a box just because it wants to. It's a calculator. You put tokens in, you get tokens out. You don't put tokens in... you don't get tokens out.

Post reply on HN