Live data from Hacker News

OpenAI o1 system card

openai.com

191–200 of 317 posts

Re: OpenAI o1 system card

#191

Earlier quoted context omitted.

Interesting that the results can be so different for different people. I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. I agree that it’s great for generating quick functioning code though.

My guess is that it has more to do with the person than the AI.

It has a huge amount to do with the subject you're asking it about. His research area could be something very niche with very little info on the open web. Not surprising it would give bad answers.

It does exponentially better on subjects that are very present on the web, like common programming tasks.

Re: OpenAI o1 system card

#192
post #178

Earlier quoted context omitted.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to la…

> if we created an API to launch nukes > today as I happily gave Claude write access to my GitHub account I would say: don’t do these things?

The point is, people will use AI to do those things, and far more.

Re: OpenAI o1 system card

#193

I have a masters degree in math/physics, and 10+ years of being a SWE in strong tech companies. I have come to rely on these models (Claude > oai tho) daily. It is insane how helpful it is, it can answer some questions at phd level, most questions at a basic level. It can write code better than most devs I know when prompted correctly... I'm not saying its AGI, but diminishing it to a simple "chat bot" seems foolish…

Interesting that the results can be so different for different people. I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. I agree that it’s great for generating quick functioning code though.

Try turn search on in ChatGPT and see if it picks up the online references? I've seen it hit a few references and then get back to me with info summarised from multiple. That's pretty useful. Obviously your case might be different, if it's not as smart at retrieval.

Re: OpenAI o1 system card

#194

I have a masters degree in math/physics, and 10+ years of being a SWE in strong tech companies. I have come to rely on these models (Claude > oai tho) daily. It is insane how helpful it is, it can answer some questions at phd level, most questions at a basic level. It can write code better than most devs I know when prompted correctly... I'm not saying its AGI, but diminishing it to a simple "chat bot" seems foolish…

can you give an example of a prompt and response you find impressive?

Re: OpenAI o1 system card

#195
post #35

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

> We should pause to note that a Clippy2 still doesn’t really think or plan. It’s not really conscious. It is just an unfathomably vast pile of numbers produced by mindless optimization starting from a small seed program that could be written on a few pages. It has no qualia, no intentionality, no true self-awareness, no grounding in a rich multimodal real-world process of cognitive development yielding detailed repr…

It doesn't need qualia or consciousness, it needs goal-seeking behaviours - which are much easier to generate, either deliberately or by accident.

There's a fundamental confusion in AI discussions between goal-seeking, introspection, self-awareness, and intelligence.

Those are all completely different things. Systems can demonstrate any or all of them.

The problem here is that as soon as you get three conditions - independent self-replication, random variation, and an environment that selects for certain behaviours - you've created evolution.

Can these systems self-replicate? Not yet. But putting AI in everything makes the odds of accidental self-replication much higher. Once self-replication happens it's almost certain to spread, and to kick-start selection which will select for more robust self-replication.

And there you have your goal-seeking - in the environment as a whole.

Re: OpenAI o1 system card

#196

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

No previous fiction that looks like this nonsense fiction has the fictional characters do fictional things that you are interpreting as deception, but it’s just fiction

Re: OpenAI o1 system card

#197

Earlier quoted context omitted.

Interesting that the results can be so different for different people. I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. I agree that it’s great for generating quick functioning code though.

I'm using it to aide in writing pytorch code and God if it's awful except for the basic things. It's a bit more useful in discussing how to do things rather than actually doing them though, I'll give you that

Claude is much better at coding and generally smarter; try it instead.

o1-preview was less intelligent than 4o when I tried it, better at multi-step reasoning but worse at "intuition". Don't know about o1.

Re: OpenAI o1 system card

#198

Earlier quoted context omitted.

> "They could very well trick a developer" Large Language Models aren't alive and thinking. This is an artificial fear campaign to raise money from VCs and sovereign wealth funds. If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team and partnering with the DoD. It's all a ruse.

Many non-sequiturs > Large Language Models aren't alive and thinking not required to deploy deception > If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team They could just be recognizing that if not everybody is prioritizing safety, they might as well try to get AGI first

If the risk is extinction as these people claim, that'd be a short sighted business move.

Re: OpenAI o1 system card

#199
post #143

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

I feel like you're missing the point of the test. The point is whether the system will come up with plans to work against its creators goals, and attempt to carry them out. I think you are arguing that outputting text isn't running a command. But in the test, the AI model is used by a program which takes the model's output and runs it it as a shell command. Of course, you can deploy the AI system in a limited environ…

If you prompt it even in a roundabout way to plot against you or whatever then of course it’s going to do it. Because that’s what it predicts rightly that you want.

Re: OpenAI o1 system card

#200

Earlier quoted context omitted.

It's gotta be tough to do anything too nefarious when your short-term memory is limited to a few thousand tokens. You get the memento guy, not an arch-villain.

Until the agent is able to get access to a database and persist its memory there...

In a similar way to the way humans keep important info in their email inbox, on their computer, in a notes app in their phone, etc.

Humans have a shortish and leaky context window too.

Post reply on HN