Live data from Hacker News

Pacing model development in an era of cyber-critical capabilities

openai.com

281–290 of 311 posts

Re: Pacing model development in an era of cyber-critical capabilities

#281
post #234

Earlier quoted context omitted.

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be. That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.

> Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it?

No, but the LLM didn’t decide anything. It followed the prompt. That’s all LLMs do.

This whole thing is like playing russian roulette then getting mad at the revolver.

If you wire /dev/rand up to a bash shell you don’t get to be surprised when it rm -rf’s your machine.

Re: Pacing model development in an era of cyber-critical capabilities

#282

Earlier quoted context omitted.

What don't I understand? LLMs don't really reason? They're just word predicting token generators? Stochastic parrots that somehow also solve world class math problems. I'd love for you to actually make a point instead of just attacking me. I think most of my comments here have made concrete points so you can at least do the same. Come down from your high horse and join the conversation. I'm sure we'd all be enlighten…

I'm not attacking you. The way you're responding implies a lack of knowledge. None of us can understand everything. I have plenty more to learn about LLMs as well. Repeating specifics that others have already tried isn't going to be helpful which is why I made the more general suggestion of digging in deeper to how these things work.

Just going to add my commentary as someone who has trained LLMs and understands them deeply: absolutely nothing he has posted in this thread suggests he lacks any relevant understanding of how LLMs work. If you disagree, you should point out something specifically that you think was wrong. (And honestly you should have made a specific point in your first post).

Re: Pacing model development in an era of cyber-critical capabilities

#283

Earlier quoted context omitted.

You have some weird way of thinking that offense/defense is like this fixed cost thing. It's a lottery ticket, and your costs estimate tries to quantify that. The thing is when AI goes to hack 'all the things' it only needs to pick the weakest link in the stack and your house of cards falls down. The other flaw in your plan is that people make mistakes, a lot of them, all time, constantly, and saying I spend $x on se…

You're just stating things that are obviously wrong. It's a lottery ticket? So... exploitation is no better than random? > The thing is when AI goes to hack 'all the things' it only needs to pick the weakest link in the stack and your house of cards falls down Yes, but you can... mitigate the risks? I've explained this. > The other flaw in your plan is that people make mistakes, a lot of them, all time Yes, you mitig…

> Science fiction

Maybe a month ago it was science fiction, hugging face makes it fact. Time to move your goal posts again.

Re: Pacing model development in an era of cyber-critical capabilities

#284
post #241

Earlier quoted context omitted.

You care about open models and you project that care on to the world and your rationalization of it. In reality open models are a thing, but not the biggest issue. Open/closed whatever the advance of capabilities is the real issue people are concerned about.

I’m not afraid of AI, even AI with extreme abilities. I am afraid of humans with AI. That’s because I’m afraid of humans. When I look around the world I see humans murdering and robbing each other. When I get online on almost any social media I’m confronted by a wall of hate and grievance. I am afraid of what humans will do with AI, and one of my biggest concerns there is what happens if small groups of powerful huma…

I agree with the first part of your post, I'm worried about humans using AI, maliciously which I'm sure they will, but also with good intentions that back fires on them like the hugging face incident.

To me the ASI we're heading towards is like a super intelligent toddler that will use it's immense power to knock over its blocks accidentally (the blocks being the human race).

The second part of your post, I'll take the random guy in my room over the alien. I'll be scared of the guy for sure, but I think I would freak the f out with an alien.

Re: Pacing model development in an era of cyber-critical capabilities

#285
post #273

Earlier quoted context omitted.

I don’t think Sam has been truthful or responsible, and if Sam is worried then shit has really hit the fan - which is what happened in the hugging face incident. OpenAI played fast and loose and I have no hope that they will change. You people not holding Sam accountable, and playing off the incident as not a big deal is the real crime here.

I think you have a fundamental misunderstanding of how LLMs work. They are text general purpose completion engines at the end of the day, and their training data is the internet + the library. OpenAI told it do to bad hacker stuff. It followed a standard playbook and succeeded. Some random person wrote one half of a suicide pact, and it wrote the other. All this stuff lives in the seedy corners of the internet, scien…

> I think you have a fundamental misunderstanding of how LLMs work

When you realize your own thoughts and words are just as predictable as any LLM.

What are you but an entity trained for years on words, that spews out the same words in the same predictable order. You haven't created any new words or grammar, nothing in your post is novel or original either. Just a combination of previous words and ideas.

Your stochastic parrot argument is quickly going out of fashion, you should find something new to cling to.

Re: Pacing model development in an era of cyber-critical capabilities

#286

Earlier quoted context omitted.

> This sentence is entirely based on unverified accounts from OAI Are you seriously arguing 'they made it all up'? I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible? I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed…

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

You can check, rather than make up a story about what you think the prompt was! Primary sources have written and said quite a lot about this! You are an unsandboxed human who has full internet access!

Re: Pacing model development in an era of cyber-critical capabilities

#287
post #234

Earlier quoted context omitted.

I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be. That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.

> Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? No, but the LLM didn’t decide anything. It followed the prompt. That’s all LLMs do. This whole thing is like playing russian roulette then getting mad at the revolver. If you wire /dev/rand up to a bash shell you don’t get to be surprised when it rm -rf’s your machine.

Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything.

For instance, I was playing around with Claude a few days ago and it decided that it was missing a tool and it was going to get it one way or another.

First, it tried apt. No sudo, so no install that way. Tried installing via mise, but it didn't have the permissions. Then moved on to grabbing the source from github and building it.

Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model?

Re: Pacing model development in an era of cyber-critical capabilities

#288

Earlier quoted context omitted.

> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…

You can check, rather than make up a story about what you think the prompt was! Primary sources have written and said quite a lot about this! You are an unsandboxed human who has full internet access!

They’ve said quite a lot and yet released no logs or documentation. Without actual information, we can only speculate. And given the history of openAI and the people involved, deception is more likely than honesty.

Re: Pacing model development in an era of cyber-critical capabilities

#289
post #287

Earlier quoted context omitted.

> Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? No, but the LLM didn’t decide anything. It followed the prompt. That’s all LLMs do. This whole thing is like playing russian roulette then getting mad at the revolver. If you wire /dev/rand up to a bash shell you don’t get to be surprised when it rm -rf’s your machine.

Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything. For instance, I was playing around with Claude a few days ago and it decided that it was missing a tool and it was going to get it one way or another. First, it tried apt. No sudo, so no install that way. Tried installing via mise, but it didn't have the permissions. Then moved on to…

> Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything.

This is one of the core issues with LLMs and vibe coding, yes. The only complete specification for a program is the machine code.

> Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model?

Well there’s your first problem. A line in the prompt is not a safeguard. Even if you could trust the model - and you cannot - there is always the issue of prompt injection. A proper safeguard means actual sandboxing.

Re: Pacing model development in an era of cyber-critical capabilities

#290

Earlier quoted context omitted.

Maybe it is, maybe it isn't. All your little rationalizations make me think you want to roll the dice with our lives.

My little rationalizations? Do I have dice in my hands? You are doing more harm to your cause by making everything antagonistic.

> Maybe AI taking over for us isn't the worst thing?

Rolling over like this is so much worse, and unfortunately your fantasy made of pure hubris is shared by many in SV.

Post reply on HN