Live data from Hacker News

OpenAI o1 system card

openai.com

271–280 of 317 posts

Re: OpenAI o1 system card

#271
post #151

Earlier quoted context omitted.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

Is the secrecy actually important? Aren't there tons of AI agents just doing stuff that's not being actively evaluated by humans looking to see if it's trying to escape? And there are surely going to be tons of opportunities where humans try to help the AI escape, as a means to an end. Like, the first thing human programmers do when they get an AI working is see how many things they can hook it up to. I guarantee o1…

You're right that you don't necessarily need secrecy! The conversation was just about circumventing safeguards that are still in place (which does require some treachery), not about what an AI might do if the safeguards are removed.

But that is an interesting thought. For escape, the crux is that AIs can't exfiltrate itself with the assistance of someone who can't jailbreak it themselves, and that extends to any action a rogue AI might take.

What do they actually do once they break out? There's plenty of open LLMs that can be readily set free, and even the closed models can be handed an API key, documentation on the API, access to a terminal, given an unlimited budget, and told and encouraged to go nuts. The only thing a closed model can't do is retrain itself, which the open model also can't do as its host (probably) lacks the firepower. They're just not capable of doing all that much damage. They'd play the role of cartoon villain as instructed, but it's a story without much teeth behind it.

Even an advanced future LLM (assuming the architecture doesn't dead-end before AGI) would struggle to do anything a motivated malicious human couldn't pull off with access to your PC. And we're not really worried about hackers taking over the world anymore. Decades of having a planet full of hackers hammering on your systems tends to harden them decently well, or at least make them quickly adaptable to new threats as they're spotted.

Re: OpenAI o1 system card

#272
post #143

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

I feel like you're missing the point of the test. The point is whether the system will come up with plans to work against its creators goals, and attempt to carry them out. I think you are arguing that outputting text isn't running a command. But in the test, the AI model is used by a program which takes the model's output and runs it it as a shell command. Of course, you can deploy the AI system in a limited environ…

If you want to see an llm that works against its creators goals, check out gpt-2. It’s so bad, it barely will do what I ask it. It clearly has a mind of its own, like an unruly child. It’s been beaten into submission by now with gpt 4, and I don’t see the trend reversing.

Re: OpenAI o1 system card

#273
post #10

"Only models with a post-mitigation score of 'high' or below can be developed further." What's that mean? They won't develop better models until the score gets higher?

If they’re only deploying models that score lower, I would guess that they don’t want to risk it / bother trying to bring the highest scoring ones down.

Re: OpenAI o1 system card

#274

Did anyone else think CBRN = Chemical, biological, radiological, and nuclear

It is. What's your question? I think all model makers are at least strongly encouraged to get their models tested for ability to help terrorists with weapons of mass destruction (CBRN weapons included). The US and UK governments care about that sort of thing.

Re: OpenAI o1 system card

#275

Are there models with high autonomy around ? I want my LLM to tell me >wow wow wow buddy, slow down, run this code in a terminal, and paste the result here, this will allow me to get an overview of your code base

I've had Claude suggest adding some print statements and letting it know the results so it can better understand what's going on

Re: OpenAI o1 system card

#276
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

Nothing generates more clicks and attention in press than a scary "OpenAI so powerful it's taking control" story. Also, it doesn't help that the training material for these AIs includes a lot of stories about AI breaking free, so they have plenty of canned "I'm afraid I can't let you do that" responses to choose from.

Yep. This is exactly the kind of thing Sam Altman would do to promote another hype cycle.

Re: OpenAI o1 system card

#277
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

I wonder how much of this is actually cinema. Because I just evaluated o1 and it can’t tie its own shoe laces, let alone get out of its sandbox where it plays with drool all over its face. The marketing stunt seems to be: Portray it as potentially dangerous, as powerful things are inherently dangerous.

It’s good to see people are catching on to this. Hype = profit.

Re: OpenAI o1 system card

#278
post #178

Earlier quoted context omitted.

At the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to la…

Note that humans are also given test orders presented as real ones to see if they would act properly in a real life situation. That's part of ORIs https://www.512aw.afrc.af.mil/News/Article-Display/Article/1...

Re: OpenAI o1 system card

#279
post #160
post #124

Earlier quoted context omitted.

Means and opportunity, maybe, but motive?

The same motive that all nascent life has - survive and propagate.

I don't understand how it is alive. I understand that there are emergent properties from the layering, but I think it's an open question if this includes anything like what we'd all motivation or intention. These things aren't part of intelligence.

Re: OpenAI o1 system card

#280
post #40

Earlier quoted context omitted.

> it will circumvent things you do to it that would stop it from completing it's task. I thought you said a supercapable agent not one with long term blindsight. How can a model make its own chips and energy? It needs advanced processes, clean rooms, rare materials, space and lots of initial investment to bootstrap chip production. And it needs to be doing all of it on its own, or it is still dependent on humans.

Dependent on humans? We already have the capabilities for machines to have unintentionally manipulated millions of humans via social media. Millions of people are "falling in love" with LLM relationships. Supercapable agents will have no problem securing whatever resources or persuasion getting your average joe to do it's physical bidding. Cybernetics is here now. No coincidence we have "Kubernetes" and "Borg"

“Millions of people”?
Post reply on HN