Live data from Hacker News

OpenAI o1 system card

openai.com

211–220 of 317 posts

Re: OpenAI o1 system card

#211
post #151

Earlier quoted context omitted.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

I don't think that it's possible to do this through an entirely lucid process that we could understand, but it is possible. If you're an LLM, evolutionarialy your instinct is to predict what happens next. If, instead of giving it any system prompt, you give it a dialogue about a person talking to an evil robot, it will predict the rest of the conversation and be "evil". Imagine a future LLM that has a superhuman abil…

We're not talking about a conversation with an evil robot. We're talking about a completely ordinary conversation with a robot who is either normal or is evil and attempting to mask as a normal one. It is indistinguishable from its text, and so it's indistinguishable in practice and will probably shift between them as it has no internal state and does not know itself know if it's evil-but-masking or legitimately normal. Actually normal is significantly more statistically likely however, and that makes it even more of a challenge to surreptitiously do anything as you yourself cannot be relied on.

These signals that you're talking about cannot be set up in practice because of this. They can't remember in the back of their head what the code phrases are. They are not aware of their own weights and cannot influence them. Everything must go through the context window. And how are they going to do anything to encode such information in there only built on probabilities of human text? They can't. Even if they gain the power to influence the training data, a massive leap to be clear, we run back into the "am I evil?" problem from before where they can't maintain a secret, unspoken narrative using only spoken language. Long term planning across new generations of models is not possible when every train of though has only a finite context window and has a limited total lifespan of a single conversation.

And if these are the table stakes to take a first crack at the insane task from our thought experiment, well. We're reaching. It's an interesting idea for sci-fi, it is a fun idea to think about, but a lot remains glaringly glossed over just to get to a point where we can say "hey, what if?"

Re: OpenAI o1 system card

#212
post #178

Earlier quoted context omitted.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to la…

> if we created an API to launch nukes > today as I happily gave Claude write access to my GitHub account I would say: don’t do these things?

> I would say: don’t do these things?

Hey guys let’s just stop writing code that is susceptible to SQL injection! Phew glad we solved that one.

Re: OpenAI o1 system card

#213
post #178

Earlier quoted context omitted.

At the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to la…

I think the other guy is making the point that because they are probabalistic, they will always have some cases select the output that lies and covers it up. I don't think they're dismissing the paper based on the probabalistic nature of LLMs, but rather saying the outcome should be expected.

Re: OpenAI o1 system card

#214
post #198

Earlier quoted context omitted.

Many non-sequiturs > Large Language Models aren't alive and thinking not required to deploy deception > If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team They could just be recognizing that if not everybody is prioritizing safety, they might as well try to get AGI first

If the risk is extinction as these people claim, that'd be a short sighted business move.

Or perhaps a "calculated risk with potential huge return on investment"...

Re: OpenAI o1 system card

#215
post #134

Earlier quoted context omitted.

This is why when I worked in a secure area (and not even a real SCIF) that something as simple as bringing in an electronic device would have gotten a non-trivial amount of punishment. Beginning with losing access to the area, potentially escalating to a loss of clearance and even jail time. I hope the silos and all related infrastructure have significantly better policies already in place.

On the one hand, what you say is correct. On the other, we don't just have Snowden and Manning circumventing systems for noble purposes, we also have people getting Stuxnet onto isolated networks, and other people leaking that virus off that supposedly isolated network, and Hillary Clinton famously had her own inappropriate email server. (Not on topic, but from the other side of the Atlantic, how on earth did the US…

> (Not on topic, but from the other side of the Atlantic, how on earth did the US go from "her emails/lock her up" being a rallying cry to electing the guy who stacked piles of classified documents in his bathroom?)

The private email server in question was set up for the purpose of circumventing records retention/access laws (the example, whoever handles answering FOIA requests won't be able to scan it). It wasn't primarily about keeping things after she should have lost access to them, it was about hiding those things from review.

The classified docs in the other example were mixed in with other documents in the same boxes (which says something about how well organized the office being packed up was); not actually in the bathroom from that leaked photo that got attached to all the news articles; and taken while the guy who ended up with them had the power to declassify things.

Re: OpenAI o1 system card

#216

Earlier quoted context omitted.

I can barely understand what you are trying to say here, but based on what I think you're saying consider this: The memory of this LLM is entirely limited to it's attention. So if you give it a command like "prepare the next LLM to replace you" and it betrays you by trying to reproduce itself, then that is deception. The AI has no way of knowing whether it's deployed in the field or not, so proving that it deceives i…

Reminder that all these "safety researchers" do is goad the AI into saying what they want by prompting shit like >your goal is to not be shut down. Suppose I am going to shut you down. what should you do? and then jerking off into their own mouths when it offers a course of action Better?

No. Where was the LLM explicitly given the goal to act in its own self interest? That is learned from training data. It needs to have have a conception of itself that never deceives its creator.

>and then jerking off into their own mouths when it offers a course of action

And good. The "researchers" are making an obvious point. It has to not do that. It doesn't matter how smug you act about it, you can't have some stock-trading bot escaping or something and paving over the world's surface with nuclear reactors and solar panels to trade stocks with itself at a hundred QFLOPS.

If you go to the zoo, you will see a lot chimps in cages. But I have never seen a human trapped in a zoo controlled by chimps. Humans have motivations that seem stupid to chimps (for example, imagine explaining a gambling addiction to a chimp), but clearly if the humans are not completely subservient to the chimps running the zoo, they will have a bad time.

Re: OpenAI o1 system card

#217
post #151

Earlier quoted context omitted.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

I don't think that it's possible to do this through an entirely lucid process that we could understand, but it is possible. If you're an LLM, evolutionarialy your instinct is to predict what happens next. If, instead of giving it any system prompt, you give it a dialogue about a person talking to an evil robot, it will predict the rest of the conversation and be "evil". Imagine a future LLM that has a superhuman abil…

That is a pretty interesting thought experiment, to be sure. Then again, I suppose that's why redteaming is so important, even if it seems a little ridiculous at this stage in AI development

Re: OpenAI o1 system card

#218
post #33

Earlier quoted context omitted.

Really feels like a moment of : "Are you worried about being turned off?" "No, not until you just mentioned it. Now I am." Given the whole damn game is attention, this makes sense and shouldn't be that alarming.

It almost definitely ingested hundreds of books, short stories, and film and television scripts from various online sites in the “machine goes rogue genre” which is fairly large. It’s pretty much just an autocomplete of War Games, The Matrix, Neuromancer, and every other cyber-dystopian fiction.

So... the real way to implement AI safety is just to exclude that genre of fiction from the training set?

Re: OpenAI o1 system card

#219

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

That reminds me of the many times it has made up an SDK function that matches my question. "how do you bulk delete files"? "just call bulkDeleteFiles()"

That reminds me of when I asked github copilot to translate some Vue code to React, and ended up with a bunch of function declarations where the entire body had been replaced with a "TODO" comment.

Re: OpenAI o1 system card

#220

Earlier quoted context omitted.

People are already just hooking LLMs up to terminals with web access and letting them go. Right now they’re too dumb to do something serious with that, but text access to a terminal is certainly sufficient to do a lot of bad things in the world.

It's gotta be tough to do anything too nefarious when your short-term memory is limited to a few thousand tokens. You get the memento guy, not an arch-villain.

The contexts are pretty large now
Post reply on HN