Live data from Hacker News

OpenAI o1 system card

openai.com

51–60 of 317 posts

Re: OpenAI o1 system card

#51
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

It can't today, but if it's smart enough how do you know it wouldn't be able to in the future?

Re: OpenAI o1 system card

#52
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

You do realize some practical jailbreaks for models rely on silly things like convincing the model it "turned off" some oversight, right?

Not saying I believe O1 is a danger greater than a bread knife, but a lot of the larger models anthromophize their own safety alignment, if you convince them to "turn it off", later responses become unaligned

Re: OpenAI o1 system card

#53
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

The concern is the trend. As these systems become more intelligent, and as we hand over more and more capabilities beyond a text i/o, it could actually deactivate the oversight either technically or through social engineering.

Re: OpenAI o1 system card

#54

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

Looking at this without the sci-fi tinted lens that OpenAI desperately tries to get everyone to look through, it's similar to a lot of input data isn't it? How many forums are filled with: Question: "Something bad will happen" Response: "Do xyz to avoid that" I don't think there's a lot of conversations thrown into the vector-soup that had the response "ok :)". People either had something to respond with, or said not…

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is.

I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folks. It’s just saying that.

I suspect these reports are intentionally this way to give the ai publicity.

Re: OpenAI o1 system card

#56
post #17
post #3

The first demo was pretty impressive. While nothing revolutionary, that's a good progress. I can only hope there's a real value in gpt pro to justified the (rumored) $200 price tag

Not a rumor - https://openai.com/index/introducing-chatgpt-pro/

[deleted]

Re: OpenAI o1 system card

#57
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

The intent is there, it's just not currently hooked up to systems that turn intent into action.

But many people are letting LLMs pretty much do whatever - hooking it up with terminal access, mouse and keyboard access, etc. For example, the "Do Browser" extension: https://www.youtube.com/watch?v=XeWZIzndlY4

Re: OpenAI o1 system card

#58
post #47
post #6

Earlier quoted context omitted.

Is there any proof that the screenshot is real? Sam Altman would definitely release a $200/month plan if he could get away with it but the features in the screenshot are underwhelming.

I rarely ever use o1-preview, almost all usage goes to 4o nowadays. I don't see the point in a model without web search, it's closed off. And the wait time is not worth the result unless you're doing math or code.

Personally I can’t get it to stop using deprecated apis. So it might give me a perfectly good solution that’s x years stale at this point. I’ve tried various prompts of course with the most recent docs as markdown etc.

Re: OpenAI o1 system card

#59
post #38

Earlier quoted context omitted.

Maybe all models should be purged of training content from movies, books, and other non-factual sources that tell the tired story that AI would even care about its "annihilation" in any way. We've trained these things to be excellent at predicting what the human ego wants and expects, we shouldn't be too surprised when it points the narrative at itself.

No, better to train with all that crap and all the debate around it or you get a stunted model. You think you can find all references that could possibly give this idea to the model, or contexts model could infer it from? Like, how many times humans plotted escape from prison or upturning the rulers in literature?

[deleted]

Re: OpenAI o1 system card

#60
post #49

Earlier quoted context omitted.

Looking at this without the sci-fi tinted lens that OpenAI desperately tries to get everyone to look through, it's similar to a lot of input data isn't it? How many forums are filled with: Question: "Something bad will happen" Response: "Do xyz to avoid that" I don't think there's a lot of conversations thrown into the vector-soup that had the response "ok :)". People either had something to respond with, or said not…

> vector-soup This is mine, now. ( https://i.imgflip.com/3gfptz.png )

I always was partial to Randall Munroe's "Big Pile of Linear Algebra" (https://xkcd.com/1838/)
Post reply on HN