Live data from Hacker News

OpenAI o1 system card

openai.com

171–180 of 317 posts

Re: OpenAI o1 system card

#171

Earlier quoted context omitted.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

> "They could very well trick a developer" Large Language Models aren't alive and thinking. This is an artificial fear campaign to raise money from VCs and sovereign wealth funds. If OpenAI was so afraid of AI misuse, they wouldn't be firing their safety team and partnering with the DoD. It's all a ruse.

[deleted]

Re: OpenAI o1 system card

#172
post #71

Earlier quoted context omitted.

> It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. Tale as old as time, they've been doing this since GPT-2 which they said was "too dangerous to release".

I talked to a Palantir guy at a conference once and he literally told me " I'm happy when the media hypes us up like a James Bond villain because every time the stock price goes up, in reality we mostly just aggregate and clean up data " This is the psychology of every tech hype cycle

Tech is by no means alone with this trick. Every press release is free adverticement and should be used like it.

Re: OpenAI o1 system card

#173
post #164

Earlier quoted context omitted.

HN is honestly pretty poor on AI commentary, and this post is a new low. Here, at least, I think there must be a large contributing factor of confusion about what a "system card" shows. The general factors I think contribute, after some months being surprised repeatedly: - It's tech, so people commenting here generally assume they understand it, and in day-to-day conversation outside their job, they are considered an…

Personally I am cynical because in my experience @ FAANG, "AI safety" is mainly about mitigating PR risk for the company, rather than any actual harm.

I lived through that era at Google and I'd gently suggest there's something south of Timnit that's still AI safety, and also point out the controversy was her leaving.

Re: OpenAI o1 system card

#174
What actually is a "system card"?

When I hear the term, I'd expect something akin to the "nutrition facts" infobox for food, or maybe the fee sheet for a credit card, i.e. a concise and importantly standardized format that allows comparison of instances of a given class.

Searching for a definition yields almost no results. Meta has possibly introduced them [1], but even there I see no "card", but a blog post. OpenAI's is a LaTeX-typeset PDF spanning several pages of largely text and seems to be an entirely custom thing too, also not exactly something I'd call a card.

[1] https://ai.meta.com/blog/system-cards-a-new-resource-for-und...

Re: OpenAI o1 system card

#175
post #174

What actually is a "system card"? When I hear the term, I'd expect something akin to the "nutrition facts" infobox for food, or maybe the fee sheet for a credit card, i.e. a concise and importantly standardized format that allows comparison of instances of a given class. Searching for a definition yields almost no results. Meta has possibly introduced them [1], but even there I see no "card", but a blog post. OpenAI'…

More generally, who introduced that concept of "cards" for ML models, datasets, etc? I saw it first when Huggingface got traction and at some point it seemed to have become some sort of de-facto standard. Was it an OpenAI or Huggingface thing?

Re: OpenAI o1 system card

#176
Are there models with high autonomy around ? I want my LLM to tell me

>wow wow wow buddy, slow down, run this code in a terminal, and paste the result here, this will allow me to get an overview of your code base

Re: OpenAI o1 system card

#177

Earlier quoted context omitted.

It can't today, but if it's smart enough how do you know it wouldn't be able to in the future?

> The question of whether machines can think is about as relevant as the question of whether submarines can swim It's a program with a lot of data running on a big calculator. It won't ever be "smart."

I think you’ve entirely missed the point of that quote.

Shutting them down for using the word “smart” (instead of something like “capable”) is like saying in 1900 submarines will never be able to swim across the Atlantic because they can’t swim. It’s really missing the point of the question: the submerged crossing.

Re: OpenAI o1 system card

#178

Earlier quoted context omitted.

Indeed. As I've been explaining this to my more non-techie friends, the interesting finding here isn't that an AI could do something we don't like, it's that it seems willing, in some cases, to _lie_ about it and actively cover its tracks. I'm curious what Simon and other more learned folks than I make of this, I personally found the chat on pg 12 pretty jarring.

At the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something.

That’s not what people are arguing.

The point is, if given access to the mechanisms to do disastrous thing X, it will do it.

No one thinks that it can think in the human sense. Or that it feels.

Extreme example to make the point: if we created an API to launch nukes. Are yoh certain that something it interprets (tokenizes, whatever) is not going to convince it to utilize the API 2 times out of 100?

If we put an exploitable (documented, unpatched 0 day bug bug) safe guard in its way. Are you trusting that ME or YOU couldn’t talk it into attempting to access that document to exploit the bug, bypass the safeguard and access the API?

Again, no one thinks that it’s actually thinking. But today as I happily gave Claude write access to my GitHub account I realized how just one command misinterpreted command could go completely wrong without the appropriate measures.

Do I think Claude is sentient and thinking about how to destroy my repos? No.

Re: OpenAI o1 system card

#179
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

That reminds me of the many times it has made up an SDK function that matches my question. "how do you bulk delete files"? "just call bulkDeleteFiles()"

Re: OpenAI o1 system card

#180
post #175
post #174

What actually is a "system card"? When I hear the term, I'd expect something akin to the "nutrition facts" infobox for food, or maybe the fee sheet for a credit card, i.e. a concise and importantly standardized format that allows comparison of instances of a given class. Searching for a definition yields almost no results. Meta has possibly introduced them [1], but even there I see no "card", but a blog post. OpenAI'…

More generally, who introduced that concept of "cards" for ML models, datasets, etc? I saw it first when Huggingface got traction and at some point it seemed to have become some sort of de-facto standard. Was it an OpenAI or Huggingface thing?

Presumably it's a spin off of Google's 'Model Card' from a few years back https://modelcards.withgoogle.com/about
Post reply on HN