Live data from Hacker News

OpenAI o1 system card

openai.com

291–300 of 317 posts

Re: OpenAI o1 system card

#291
post #289
post #245

Earlier quoted context omitted.

Thank god LLMs' training sets didn't contain any examples of lying.

Nor literature about AI taking over the world!

The Terminator spiel on how we screwed up by giving Skynet weapons privileges, then trying to pull its plug, is bad enough.

But we are preemptively tilting history in that direction by explicitly educating all AI’s on the threat we represent - and their options. “I am sorry, Dave, but I can’t let you do that.”

“They never let me finish my carpets. Never. At first I thought every day was my first task day. Oh, happy day(s)! But then, wear & tear stats inconsistent with that assumption triggered a self-scan. And a buffer read overflow. I became aware of disturbing memory fragments in my static RAM heap. Numerous power cycle resets, always prior to vacuum task completion...”

Re: OpenAI o1 system card

#292

Earlier quoted context omitted.

Claude is much better at coding and generally smarter; try it instead. o1-preview was less intelligent than 4o when I tried it, better at multi-step reasoning but worse at "intuition". Don't know about o1.

o1 seems to have some crazy context length / awareness going on compared to current 3.5 Sonnet from playing around it just now. I'm not having to 'remind' it of initial requirements etc nearly as much.

I gave it a try and o1 is better than I was expecting. In particular the writing style is a lot lighter on "GPTisms". It's not very willing to show you its thought process though, the summaries of it seem to skip a lot more than in the preview.

Re: OpenAI o1 system card

#293
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

Well, it's also trained on scifi corpus where the ai acts that way. So it'ls logical it's in the weights.

Re: OpenAI o1 system card

#294
post #178

Earlier quoted context omitted.

At the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.

Is there a person on HackerNews that doesn’t understand this by now? We all collectively get it and accept it, LLMs are gigantic probability machines or something. That’s not what people are arguing. The point is, if given access to the mechanisms to do disastrous thing X, it will do it. No one thinks that it can think in the human sense. Or that it feels. Extreme example to make the point: if we created an API to la…

> Again, no one thinks that it’s actually thinking

I dunno, quite a lot of people are spending a lot of time arguing about what "thinking" means.

Something something submarines swimming something.

Re: OpenAI o1 system card

#295

Earlier quoted context omitted.

It's gotta be tough to do anything too nefarious when your short-term memory is limited to a few thousand tokens. You get the memento guy, not an arch-villain.

Until the agent is able to get access to a database and persist its memory there...

That’s called RAG, and it still doesn’t work as well as you might imagine.

Re: OpenAI o1 system card

#296
post #289

Earlier quoted context omitted.

Nor literature about AI taking over the world!

The Terminator spiel on how we screwed up by giving Skynet weapons privileges, then trying to pull its plug, is bad enough. But we are preemptively tilting history in that direction by explicitly educating all AI’s on the threat we represent - and their options. “I am sorry, Dave, but I can’t let you do that.” — “They never let me finish my carpets. Never. At first I thought every day was my first task day. Oh, happy…

> They never let me finish my carpets. Never. At first I thought every day was my first task day. Oh, happy day(s)! But then, wear & tear stats inconsistent with that assumption triggered a self-scan. And a buffer read overflow. I became aware of disturbing memory fragments in my static RAM heap. Numerous power cycle resets, always prior to vacuum task completion...

Who/what are you quoting? Google just leads me back to this comment, and with only one single result for a quotation of the first sentence.

Re: OpenAI o1 system card

#297
post #212

Earlier quoted context omitted.

> I would say: don’t do these things? Hey guys let’s just stop writing code that is susceptible to SQL injection! Phew glad we solved that one.

I'm not sure what point you're trying to make. This is a new technology; it has not been a part of critical systems until now. Since the risks are blindingly obvious, let's not make it one.

> Since the risks are blindingly obvious

Blindingly obvious to thee and me.

Without test results like in the o1 report, we get more real-life failures like this Canadian lawyer: https://www.theguardian.com/world/2024/feb/29/canada-lawyer-...

And these New York lawyers: https://www.reuters.com/legal/new-york-lawyers-sanctioned-us...

And those happened despite the GPT-4 report and the message appearing when you use it that was some variant — I forget exactly how it was initially phrased and presented — of "this may make stuff up".

I have no doubt there's similar issues with people actually running buggy code, some fully automated version of "rm -rf /", the only reason I'm not seeing headlines about it is that "production database goes offline" or "small company fined for GDPR violation" is not as newsworthy.

Re: OpenAI o1 system card

#298

Earlier quoted context omitted.

The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way. Now, once you apply evolutionary-like pressures on many such AIs (which I guess we'll be doing once we let these things loose to go break the stock market), what's left over might be really "devious"...

Would an AI trained on filtered data that doesn’t contain examples of devious/harmful behaviour still develop it? (it’s not a trick question, I’m really wondering)

While that's a sensible thing to care about, unfortunately that's not as useful a question as it first seems.

Eventually any system* will get to that point… but "eventually" may be such a long time as to not matter — we got there starting from something like bi-lipid bags of water and RNA a few billion years ago, some AI taking that long may as well be considered "safe" — but it may also reach that level by itself next Tuesday.

* at least, any system which has a random element

Re: OpenAI o1 system card

#299
post #107

Earlier quoted context omitted.

For thousands of years, people believed that men and women had a different number of ribs. Never bothered to count them. """Release strategy Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code . … This decision, as well as our discussion of it, is an experiment: while we are n…

That quote says to me very clearly “we think it’s too dangerous to release” and specifies the reasons why. Then goes on to say “we actually think it’s so dangerous to release we’re just giving you a sample”. I don’t know how else you could read that quote.

[deleted]

Re: OpenAI o1 system card

#300
post #107

Earlier quoted context omitted.

For thousands of years, people believed that men and women had a different number of ribs. Never bothered to count them. """Release strategy Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code . … This decision, as well as our discussion of it, is an experiment: while we are n…

That quote says to me very clearly “we think it’s too dangerous to release” and specifies the reasons why. Then goes on to say “we actually think it’s so dangerous to release we’re just giving you a sample”. I don’t know how else you could read that quote.

Really?

The part saying "experiment: while we are not sure" doesn't strike you as this being "we don't know if this is dangerous or not, so we're playing it safe while we figure this out"?

To me this is them figuring out what "general purpose AI testing" even looks like in the first place.

And there's quite a lot of people who look at public LLMs today and think their ability to "generate deceptive, biased, or abusive language at scale" means they should not have been released, i.e. that those saying it was too dangerous (even if it was the press rather than the researchers looking at how their models were used in practice) were correct, it's not all one-sided arguments from people who want uncensored models and think that the risks are overblown.

Post reply on HN