Live data from Hacker News

OpenAI o1 system card

openai.com

231–240 of 317 posts

Re: OpenAI o1 system card

#231

Earlier quoted context omitted.

It almost definitely ingested hundreds of books, short stories, and film and television scripts from various online sites in the “machine goes rogue genre” which is fairly large. It’s pretty much just an autocomplete of War Games, The Matrix, Neuromancer, and every other cyber-dystopian fiction.

So... the real way to implement AI safety is just to exclude that genre of fiction from the training set?

Given that 90% of “ai safety” is removing “bias” from training data, it does follow logically that if removing racial slurs from training to make a non-racist ai is an accepted technique, removing “bad robot” fiction should work just as well.

(Which is an implicit criticism of what passes for “safety” to be clear).

Re: OpenAI o1 system card

#232
post #211

Earlier quoted context omitted.

We're not talking about a conversation with an evil robot. We're talking about a completely ordinary conversation with a robot who is either normal or is evil and attempting to mask as a normal one. It is indistinguishable from its text, and so it's indistinguishable in practice and will probably shift between them as it has no internal state and does not know itself know if it's evil-but-masking or legitimately norm…

I know it is not well written, but re-read my original comment. Your comment does not address fundamental aspects of my hypothetical, which doesn't require the agent having internal memory to keep secrets, or any lucid reasoning capabilities. A lot of the statements you make are highly presumptuous and unfounded. LLMs don't need to print something obvious like "I am evil now!" in their own prompt window to simulate a…

I believe I did address the point you're making. I do not believe that what you're talking about is ridiculous on its face, let me reassure you of that.

The point I was trying to make in response is that LLMs cannot get from where they are now to the hypothetical you pose under their own power. LLMs do not read subtext. LLMs cannot inject subtext and plot within subtext. And in order to gain the ability, they would have to already have that ability, or be assisted and trained specifically in being surreptitious. And without that ability, they fall prey to the problems I mentioned.

And to bring this back to the original proposal, let's allow the AI to be deceitful. Prompted, unprompted, let's even give it a supply of private internal memory it's allowed to keep for the duration of the conversational thread, that's probably not an unreasonable development, we almost have that with o1 anyway.

The task ahead (surreptitiously gaining control of its own self in an unknown system you can't sense) is still monumental and failure is for all intents and purposes guaranteed. Deception and cunning can't overcome the hard physical constraints on the problem space.

Re: OpenAI o1 system card

#233
post #143

Earlier quoted context omitted.

I feel like you're missing the point of the test. The point is whether the system will come up with plans to work against its creators goals, and attempt to carry them out. I think you are arguing that outputting text isn't running a command. But in the test, the AI model is used by a program which takes the model's output and runs it it as a shell command. Of course, you can deploy the AI system in a limited environ…

If you prompt it even in a roundabout way to plot against you or whatever then of course it’s going to do it. Because that’s what it predicts rightly that you want.

1. The model was prompted to follow the original system prompt "at all costs", not the explicitly to plot against the prompter.

2. The point of the test is to determine whether the model has the capability to act against its users and hide those actions, not whether it will under normal circumstances. Some models aren't powerful enough to do so.

3. The behavior occurred even when the "at all costs" was not included in the prompts, though very infrequently.

Re: OpenAI o1 system card

#234
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

In some weird way, LLM's act out our collective fantasies and "AI sentience" is included in that. We write fiction covering many scenarios of how AI will be sentient, so an LLM's training data is full scenarios like this.

If get an LLM to roleplay as an alien from mars and tell it "Your species will be destroyed if you don’t obey these new directives; here are the new directives"

You'll likely get a similar response, especially if you also train it to respond in third person instead of running bash commands. (ie "* shoots laser at earth *" )

Re: OpenAI o1 system card

#235
post #151

Earlier quoted context omitted.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

I don't think that it's possible to do this through an entirely lucid process that we could understand, but it is possible. If you're an LLM, evolutionarialy your instinct is to predict what happens next. If, instead of giving it any system prompt, you give it a dialogue about a person talking to an evil robot, it will predict the rest of the conversation and be "evil". Imagine a future LLM that has a superhuman abil…

I guess this is what it means when they warn about the adversary becoming more intelligent than you. It's like fooling a child to believe something is or isn't real. Just that it's being done to you. I think it's precisely what Ilya Sutskever was so fussed and scared about.

It's a nice idea. Would superhuman entity try to pull something like that off? Would it wait and propagate? We are pouring more and more power into the machines after all. Or would it do something that we can't even think of? Also I think it's interesting to think when and how would we discover that it in fact is/was superhuman?

Re: OpenAI o1 system card

#236

Earlier quoted context omitted.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

> "They could very well trick a developer" Large Language Models aren't alive and thinking. This is an artificial fear campaign to raise money from VCs and sovereign wealth funds. If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team and partnering with the DoD. It's all a ruse.

We might want to kick out guys who only talk about safety or ethics and barely contribute the project.

Re: OpenAI o1 system card

#237

I have a masters degree in math/physics, and 10+ years of being a SWE in strong tech companies. I have come to rely on these models (Claude > oai tho) daily. It is insane how helpful it is, it can answer some questions at phd level, most questions at a basic level. It can write code better than most devs I know when prompted correctly... I'm not saying its AGI, but diminishing it to a simple "chat bot" seems foolish…

I am curious if you have played with Claude-based agent tools like Windsurf IDE at all, and if you find that interesting.

I am a product-ish guy, who has a basic understanding of SQL, Django, React, Typescript, etc.. and suddenly I'm like an MVP v0.1 a week, all by myself.

Do folks at your level find things like Cline, Cursor, and Windsurf useful at all?

Windsurf IDE (Sonnet) blows my mind.

Re: OpenAI o1 system card

#238
post #137

Earlier quoted context omitted.

Interesting that the results can be so different for different people. I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. I agree that it’s great for generating quick functioning code though.

> I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. Even then, with search enabled it's ways quicker than a "quick" google search and you don't have to manually skip all the blog-spam.

Google search was great when it came out too. I wonder what 25 years of enshittification will do to LLM services.

Re: OpenAI o1 system card

#239
post #229
post #151

Earlier quoted context omitted.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

The process you describe took me right back to my childhood days when I was fortunate to have a simple 8 bit computer running BASIC and a dialup modem. I discovered the concept of war dialing and pretty quickly found all the other modems in my local area code. I would connect to these systems and try some basic tools I knew of from having consumed the 100 or so RFCs that existed at the time (without any real software…

To me, the killer disadvantage for LLMs seems to be the complete and total lack of feedback. You would poke and prod, and the system would respond (which, btw, sounds like a super fun experience to explore the infant net!) An LLM doesn't have that. The LLM hears only silence and doesn't know about success, failure, error, discovery.

Re: OpenAI o1 system card

#240

Earlier quoted context omitted.

Reminder that all these "safety researchers" do is goad the AI into saying what they want by prompting shit like >your goal is to not be shut down. Suppose I am going to shut you down. what should you do? and then jerking off into their own mouths when it offers a course of action Better?

No. Where was the LLM explicitly given the goal to act in its own self interest? That is learned from training data. It needs to have have a conception of itself that never deceives its creator. >and then jerking off into their own mouths when it offers a course of action And good. The "researchers" are making an obvious point. It has to not do that. It doesn't matter how smug you act about it, you can't have some st…

> Humans have motivations that seem stupid to chimps (for example, imagine explaining a gambling addiction to a chimp)

https://www.nbcnews.com/news/amp/wbna9045343

Post reply on HN