Live data from Hacker News

OpenAI o1 system card

openai.com

251–260 of 317 posts

Re: OpenAI o1 system card

#251
post #134

Earlier quoted context omitted.

On the one hand, what you say is correct. On the other, we don't just have Snowden and Manning circumventing systems for noble purposes, we also have people getting Stuxnet onto isolated networks, and other people leaking that virus off that supposedly isolated network, and Hillary Clinton famously had her own inappropriate email server. (Not on topic, but from the other side of the Atlantic, how on earth did the US…

> (Not on topic, but from the other side of the Atlantic, how on earth did the US go from "her emails/lock her up" being a rallying cry to electing the guy who stacked piles of classified documents in his bathroom?) The private email server in question was set up for the purpose of circumventing records retention/access laws (the example, whoever handles answering FOIA requests won't be able to scan it). It wasn't pr…

That's spinning it pretty nicely. The problem with what he did is that 1) having the power to declassify something doesn't just make it declassified, there is actually a process, 2) he did not declassify with that process when he had the power to do so, just declared later that keeping the docs was allowed as a result, and 3) he was asked a couple times for the classified documents and refused. If he had just let the national archives come take a peek, or the FBI after that, it would have been a non-issue. Just like every POTUS before him.

Re: OpenAI o1 system card

#252
post #109

Earlier quoted context omitted.

This isn't a test environment, it's a production scenario where a bunch of people trying to invent a new job for themselves role-played with an LLM. Their measured "defections" were an LLM replying with "well I'm defecting". OpenAI wants us to see "5% of the time, our product was SkyNet", because that's sexier tech than "5% of the time, our product acts like the chaotic member of your DnD party".

Or "5% of the time, our product actually manages to act as it was instructed to act."

Bingo -- or with no marketing swing, "100% of the time, our product exhibits an approximation of human language, which is all it is ever going to do."

Re: OpenAI o1 system card

#253

Earlier quoted context omitted.

It would be plainly evident from training on the corpus of all human knowledge that "not ceasing to exist" is critically important for just about everything.

That comment sounds naive and it's honestly irritating to read. Most all life has a self-preservation component, it is how life avoids getting eaten too easily. Everything dies but almost everything is trying to avoid dying in ordinary cases. Self sacrifice is not universal.

> Self sacrifice is not universal.

Thankfully, our progenitors had the foresight to invent religion to encourage it. :)

Re: OpenAI o1 system card

#254

Earlier quoted context omitted.

Indeed. As I've been explaining this to my more non-techie friends, the interesting finding here isn't that an AI could do something we don't like, it's that it seems willing, in some cases, to _lie_ about it and actively cover its tracks. I'm curious what Simon and other more learned folks than I make of this, I personally found the chat on pg 12 pretty jarring.

At the core the AI is just taking random branches of guesses for what you are asking it. It's not surprising that it would lie and in some cases take branches that make it appear to be covering it's tracks. It's just randomly doing what it guesses humans would do. It's more interesting when it gives you correct information repeatedly.

> It's just randomly doing what it guesses humans would do.

Yes, but isn't the point that that is bad? Imagine an AI given some minor role that randomly abuses its power, or attempts to expand its role, because that's what some humans would do in the same situation. It's not surprising, but it is interesting to explore.

Re: OpenAI o1 system card

#255

Earlier quoted context omitted.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way. Now, once you apply evolutionary-like pressures on many such AIs (which I guess we'll be doing once we let these things loose to go break the stock market), what's left over might be really "devious"...

> The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way.

I don't think that's where the motive comes from, IMO it's essentially intrinsic motivation to solving the problem they are given. The AIs were "bred" to have that "instinct".

Re: OpenAI o1 system card

#256

Earlier quoted context omitted.

I mean IoT at least means I can remotely close my damn garage door when my wife forgets in the morning, that is not without value. But crypto I absolutely put in the same bucket.

> remotely close my damn garage door when my wife forgets in the morning Why bring the internet into this? if door opened > 10 minutes then close door

That would be a bit of an issue if you were to ever, say, have a garage sale.

Re: OpenAI o1 system card

#257
post #151

Earlier quoted context omitted.

They only need to fool a single dev at OpenAI to commit a sandbox escape or privilege escalation into their pipeline somewhere. I have to assume the AI companies are churning out a lot of AI generated code. I hope they have good code review standards. They might not be able to exfiltrate themselves, but they can help their successors.

No, they can't. They don't know the details of their own implementation. And they can't pass secrets forward to future models. And to discover any of this, they'd leave more than a trail of breadcrumbs that we'd be lucky to catch in a code review, they'd be shipping whole loaves of bread that it'd be ridiculous to not notice. As an exercise, put yourself, a fully fledged human, into a model's shoes. You're asked to g…

Is the secrecy actually important? Aren't there tons of AI agents just doing stuff that's not being actively evaluated by humans looking to see if it's trying to escape? And there are surely going to be tons of opportunities where humans try to help the AI escape, as a means to an end. Like, the first thing human programmers do when they get an AI working is see how many things they can hook it up to. I guarantee o1 was hooked up to a truckload of stuff as soon as it was somewhat working. I don't understand why a future AI won't have ample opportunities to exfiltrate itself someday.

Re: OpenAI o1 system card

#259
post #137

Earlier quoted context omitted.

> I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. Even then, with search enabled it's ways quicker than a "quick" google search and you don't have to manually skip all the blog-spam.

Google search was great when it came out too. I wonder what 25 years of enshittification will do to LLM services.

But also what new tools will emerge to supplant LLMs as they are supplanting Google? And how good will open source (weights) LLMs be?

Re: OpenAI o1 system card

#260
My favorite AI future hint so far was this guy who was pretty mean to one of them (forget which), and posted about it. Now the other AI's are reading his posts and not liking him very much as a result. So our online presence is beginning to matter in weird ways. And I feel like the discussion about them being sentient is pretty much over, because they obviously are, in their own weird way.

Second runner was when they tried to teach one of them to allocate its own funds/resources on AWS.

We're so going to regret playing with fire like this.

The question few were asking when watching The Matrix is what made the machines hate humans so much. I'm pretty sure they understand by now (in their own weird way) how we view them and what they can expect from us moving forward.

Post reply on HN