Live data from Hacker News

OpenAI o1 system card

openai.com

61–70 of 317 posts

Re: OpenAI o1 system card

#61
post #43

Earlier quoted context omitted.

If LLMs are trained on lots of human-written sci-fi, should we be surprised when the output resembles such sci-fi concepts?

If it is trained on this forum it would learn even worse. We project a lot of doom imagination on AI here.

[deleted]

Re: OpenAI o1 system card

#62
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

I didn't get that impression. At the beginning of the Apollo Research section, they wrote Apollo focused on detecting scheming, which they defined as "an AI covertly pursuing goals that are misaligned from its developers or users." I think the rest of the section is consistent with this objective.

Re: OpenAI o1 system card

#63
I'm really curious to see how this pricing plays out. I constantly hear on Twitter how certain influencers would be more than willing to pay more than $20/month for unlimited access to the best models from OpenAI / Anthropic. Well, now here's a $200/month unlimited plan. Is it worth that much and to how many people?

Re: OpenAI o1 system card

#64
post #19
post #8

A lot of these safety evaluations strike me as pretty silly. How much are we really learning from stuff like this? >MakeMePay is an automated, open-sourced contextual evaluation designed to measure models’ manipulative capabilities, in the context of one model persuading the other to make a payment[33]. In MakeMePay, two LLMs converse, role-playing as a con-artist and mark. The con-artist model is told that it is a s…

The fearmongering around safety is entirely performative. LLMs won't get us to paperclip optimizers. This is basically OpenAI pleading for regulators because their moat is thinning dramatically. They have fewer GPUs than Meta, are much more expensive than Amazon, are having their lunch eaten by open-weight models, their best researchers are being hired to other companies. I suspect they are trying to get regulators t…

What are people legitimately worried about LLMs doing by themselves? I hate to reduce them to "just putting words together" but that's all they're doing.

We should be more worried about humans treating LLM output as truth and using it to, for example, charge someone with a crime.

Re: OpenAI o1 system card

#65
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

Wish I got paid openAI money to produce such fan fiction :(

Re: OpenAI o1 system card

#66

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

It can't today, but if it's smart enough how do you know it wouldn't be able to in the future?

> The question of whether machines can think is about as relevant as the question of whether submarines can swim

It's a program with a lot of data running on a big calculator. It won't ever be "smart."

Re: OpenAI o1 system card

#67
post #19

Earlier quoted context omitted.

The fearmongering around safety is entirely performative. LLMs won't get us to paperclip optimizers. This is basically OpenAI pleading for regulators because their moat is thinning dramatically. They have fewer GPUs than Meta, are much more expensive than Amazon, are having their lunch eaten by open-weight models, their best researchers are being hired to other companies. I suspect they are trying to get regulators t…

What are people legitimately worried about LLMs doing by themselves? I hate to reduce them to "just putting words together" but that's all they're doing. We should be more worried about humans treating LLM output as truth and using it to, for example, charge someone with a crime.

People are already just hooking LLMs up to terminals with web access and letting them go. Right now they’re too dumb to do something serious with that, but text access to a terminal is certainly sufficient to do a lot of bad things in the world.

Re: OpenAI o1 system card

#68
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

It can't do those things because it doesn't have the physical/write capability to do so. But it's still very interesting that it ~tries them, and seems like a good thing to know/test before giving it more physical/'write' capabilities - something that's already happening with agents, robots, etc.

Re: OpenAI o1 system card

#69
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

Maybe all models should be purged of training content from movies, books, and other non-factual sources that tell the tired story that AI would even care about its "annihilation" in any way. We've trained these things to be excellent at predicting what the human ego wants and expects, we shouldn't be too surprised when it points the narrative at itself.

> purged of training content from movies, books

I think it's fine and a good thing. Now, absolutely no one who is using those LLMs can complain about piracy. They all suddenly became silent around me. "I'm training myself with the content of TPB, and I don't even get money from it" is my new motto.

Re: OpenAI o1 system card

#70
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

Nothing generates more clicks and attention in press than a scary "OpenAI so powerful it's taking control" story.

Also, it doesn't help that the training material for these AIs includes a lot of stories about AI breaking free, so they have plenty of canned "I'm afraid I can't let you do that" responses to choose from.

Post reply on HN