Live data from Hacker News

OpenAI o1 system card

openai.com

101–110 of 317 posts

Re: OpenAI o1 system card

#101

Earlier quoted context omitted.

People are already just hooking LLMs up to terminals with web access and letting them go. Right now they’re too dumb to do something serious with that, but text access to a terminal is certainly sufficient to do a lot of bad things in the world.

It's gotta be tough to do anything too nefarious when your short-term memory is limited to a few thousand tokens. You get the memento guy, not an arch-villain.

Until the agent is able to get access to a database and persist its memory there...

Re: OpenAI o1 system card

#102
I have a masters degree in math/physics, and 10+ years of being a SWE in strong tech companies. I have come to rely on these models (Claude > oai tho) daily.

It is insane how helpful it is, it can answer some questions at phd level, most questions at a basic level. It can write code better than most devs I know when prompted correctly...

I'm not saying its AGI, but diminishing it to a simple "chat bot" seems foolish to me. It's at least worth studying, and we should be happy they care rather than just ship it?

Re: OpenAI o1 system card

#103
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

Re: OpenAI o1 system card

#105
post #35

Earlier quoted context omitted.

> We should pause to note that a Clippy2 still doesn’t really think or plan. It’s not really conscious. It is just an unfathomably vast pile of numbers produced by mindless optimization starting from a small seed program that could be written on a few pages. It has no qualia, no intentionality, no true self-awareness, no grounding in a rich multimodal real-world process of cognitive development yielding detailed repr…

what is the relevance of the quoted passage here? its relation to parent seems unclear to me.

His point is that while we're over here arguing over whether a particular AI is "really" doing certain things (e.g. knows what it's doing), it can still cause tremendous harm if it optimizes or hallucinates in just the right way.

Re: OpenAI o1 system card

#106
post #71

Earlier quoted context omitted.

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folk…

> It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. Tale as old as time, they've been doing this since GPT-2 which they said was "too dangerous to release".

I talked to a Palantir guy at a conference once and he literally told me "I'm happy when the media hypes us up like a James Bond villain because every time the stock price goes up, in reality we mostly just aggregate and clean up data"

This is the psychology of every tech hype cycle

Re: OpenAI o1 system card

#107
post #71

Earlier quoted context omitted.

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folk…

> It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. Tale as old as time, they've been doing this since GPT-2 which they said was "too dangerous to release".

For thousands of years, people believed that men and women had a different number of ribs. Never bothered to count them.

"""Release strategy

Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code .

This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today … """ - https://openai.com/index/better-language-models/

It was the news reporting that it was "too dangerous".

If anyone at OpenAI used that description publicly, it's not anywhere I've been able to find it.

Re: OpenAI o1 system card

#108

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

Re: OpenAI o1 system card

#109
post #98

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

It reads like you think failing tests can't ever be bad because they're in a test environment? So it merely knows how to approach the task of deleting its own off-switch but didn't actually pass that command to a real execution environment. That's already bad because people do sometimes blindly pass commands from the context windows to execution environments . Should they? No, they should not. Not blindly. But they d…

This isn't a test environment, it's a production scenario where a bunch of people trying to invent a new job for themselves role-played with an LLM. Their measured "defections" were an LLM replying with "well I'm defecting".

OpenAI wants us to see "5% of the time, our product was SkyNet", because that's sexier tech than "5% of the time, our product acts like the chaotic member of your DnD party".

Re: OpenAI o1 system card

#110
post #49

Earlier quoted context omitted.

Looking at this without the sci-fi tinted lens that OpenAI desperately tries to get everyone to look through, it's similar to a lot of input data isn't it? How many forums are filled with: Question: "Something bad will happen" Response: "Do xyz to avoid that" I don't think there's a lot of conversations thrown into the vector-soup that had the response "ok :)". People either had something to respond with, or said not…

> vector-soup This is mine, now. ( https://i.imgflip.com/3gfptz.png )

> https://i.imgflip.com/3gfptz.png

Yoink! That is mine, now, along with vector-soup.

Post reply on HN