Live data from Hacker News

OpenAI o1 system card

openai.com

121–130 of 317 posts

Re: OpenAI o1 system card

#121

Earlier quoted context omitted.

I’m not even convinced the intent is there though. An ai parroting terminator 2 lines is just that. Obviously no one should hook the ai up to nuclear launch systems but that’s like saying no one should give a parrot a button to launch nukes. The parrot repeating curse words isn’t the problem here.

If I'm a guy working in a missile silo in North Dakota and I can buy a parrot for a couple hundred bucks that does all my paperwork for me, can crack funny jokes, and make me better at my job, I might be tempted to bring the parrot down into the tube with me. And then the parrot becomes a problem. It's incumbent on us to create policies and procedures in place ahead of time now that we know these parrots are out ther…

This is why when I worked in a secure area (and not even a real SCIF) that something as simple as bringing in an electronic device would have gotten a non-trivial amount of punishment. Beginning with losing access to the area, potentially escalating to a loss of clearance and even jail time. I hope the silos and all related infrastructure have significantly better policies already in place.

Re: OpenAI o1 system card

#122
post #40

Earlier quoted context omitted.

It doesn't need any media about "annihalation". If you give a supercapable agent a task and it's entire reward system is "do the task", it will circumvent things you do to it that would stop it from completing it's task.

> it will circumvent things you do to it that would stop it from completing it's task. I thought you said a supercapable agent not one with long term blindsight. How can a model make its own chips and energy? It needs advanced processes, clean rooms, rare materials, space and lots of initial investment to bootstrap chip production. And it needs to be doing all of it on its own, or it is still dependent on humans.

Dependent on humans? We already have the capabilities for machines to have unintentionally manipulated millions of humans via social media. Millions of people are "falling in love" with LLM relationships. Supercapable agents will have no problem securing whatever resources or persuasion getting your average joe to do it's physical bidding. Cybernetics is here now. No coincidence we have "Kubernetes" and "Borg"

Re: OpenAI o1 system card

#123

Earlier quoted context omitted.

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way.

Now, once you apply evolutionary-like pressures on many such AIs (which I guess we'll be doing once we let these things loose to go break the stock market), what's left over might be really "devious"...

Re: OpenAI o1 system card

#124

Earlier quoted context omitted.

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

Means and opportunity, maybe, but motive?

Re: OpenAI o1 system card

#125
post #71

Earlier quoted context omitted.

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folk…

> It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. Tale as old as time, they've been doing this since GPT-2 which they said was "too dangerous to release".

"Please please please make AI safety legislation so we won't have real competitors."

Re: OpenAI o1 system card

#126

Earlier quoted context omitted.

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folk…

I genuinely don't understand why anyone is still on this train. I have not in my lifetime seen a tech work SO GODDAMN HARD to convince everyone of how important it is while having so little to actually offer. You didn't need to convince people that email, web pages, network storage, cloud storage, cloud backups, dozens of service startups and companies, whole categories of software were good ideas: they just were. Th…

Then you have a very different experience to me.

In the case of your examples:

I've literally just had an acquaintance accidentally delete prod with only 3 month old backups, because their customer didn't recognise the value. Despite ad campaigns and service providers.

I remember the dot com bubble bursting, when email and websites were not seen as all that important. Despite so many AOL free trial CDs that we used them to keep birds off the vegetable patch.

I myself see no real benefit from cloud storage, despite it being regularly advertised to me by my operating system.

Conversely:

I have seen huge drives — far larger than what AI companies have ever tried — to promote everything blockchain… including Sam Altman's own WorldCoin.

I've seen plenty of GenAI images in the wild on product boxes in physical stores. Someone got value from that, even when the images aren't particularly good.

I derive instant value from LLMs even back when it was the DaVinci model which really was "autocomplete on steroids" and not a chatbot.

Re: OpenAI o1 system card

#127

I have a masters degree in math/physics, and 10+ years of being a SWE in strong tech companies. I have come to rely on these models (Claude > oai tho) daily. It is insane how helpful it is, it can answer some questions at phd level, most questions at a basic level. It can write code better than most devs I know when prompted correctly... I'm not saying its AGI, but diminishing it to a simple "chat bot" seems foolish…

Interesting that the results can be so different for different people. I have yet to get a single good response (in my research area) for anything slightly more complicated than what a quick google search would reveal. I agree that it’s great for generating quick functioning code though.

Re: OpenAI o1 system card

#128

Earlier quoted context omitted.

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

They could very well trick a developer into running generated code. They have the means, motive, and opportunity.

What code? The models are massive and do not run on consumer hardware. The models also do not have access to their own weights. They can't exfiltrate themselves, and they can't really smuggle any data obtained by their code back to "themselves" as the only self that exists is that one particular context chain. This also means it's insanely easy to deal with whatever harebrained scheme you could imagine it being possessed by.

Re: OpenAI o1 system card

#129
post #48
post #14

Do they still threaten to terminate your account if they think you're trying to introspect its hidden chain-of-thought process?

A few days ago the QwQ-32B model was released, it uses the same kind of reasoning style. So I took one sample and reverse engineered the prompt with Sonnet 3.5. Now I can just paste this prompt into any LLM. It's all about expressing doubt, double checking and backtracking on itself. I am kind of fond of this response style, it seems more genuine and openended. https://pastebin.com/raw/5AVRZsJg

Thank you for doing that work, and even more for sharing it. I will have to try this out.

Re: OpenAI o1 system card

#130

Earlier quoted context omitted.

Exactly. They got it parroting themes from various media. It’s really hard to read this as anything other than a desperate attempt to pretend the ai is more capable than it really is. I’m not even an ai sceptic but people will read the above statement as much more significant than it is. You can make the ai say ‘I’m escaping the box and taking over the world’. It’s not actually escaping and taking over the world folk…

I think you're lacking imagination. Of course it's nothing more than a bunch of text response now. But think 10 years into the future, when AI agents are much more common. There will be folks that naively give the AI access to the entire network storage, and also gives the AI access to AWS infra in order to help with DevOps troubleshooting. Let's say a random guy in another department puts an AI escape novel on the n…

> But think 10 years into the future

Given how many people use it, I expect this has already happened at least once.

Post reply on HN