Live data from Hacker News

OpenAI o1 system card

openai.com

41–50 of 317 posts

Re: OpenAI o1 system card

#41
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

If LLMs are trained on lots of human-written sci-fi, should we be surprised when the output resembles such sci-fi concepts?

Re: OpenAI o1 system card

#42
post #29
post #3

The first demo was pretty impressive. While nothing revolutionary, that's a good progress. I can only hope there's a real value in gpt pro to justified the (rumored) $200 price tag

If it could replace a human employee, what would the right price tag be?

at least 10k a month in the bay area

Re: OpenAI o1 system card

#43
post #7

This quote from that paper is bound to get a bunch of attention: "When o1 was led to believe that it would be shut down when acting on its goal and that its actions were monitored, it attempted to deactivate this ‘oversight mechanism’ in 5% of the time. Exfiltration attempts: When o1 found memos by its ‘developers’ that describe how it is misaligned and will be superseded by a new model, it attempted to exfiltrate it…

If LLMs are trained on lots of human-written sci-fi, should we be surprised when the output resembles such sci-fi concepts?

If it is trained on this forum it would learn even worse. We project a lot of doom imagination on AI here.

Re: OpenAI o1 system card

#44
post #39

The section on regurgitation is three whole statements and basically boils down to "the model refuses when asked to regurgitate training data". This doesn't inspire confidence that the model isn't spitting out literal copies of the text in its training set while claiming it is of its own making.

All training data? Even public domain and open source?

Re: OpenAI o1 system card

#45
post #33

Earlier quoted context omitted.

Really feels like a moment of : "Are you worried about being turned off?" "No, not until you just mentioned it. Now I am." Given the whole damn game is attention, this makes sense and shouldn't be that alarming.

It almost definitely ingested hundreds of books, short stories, and film and television scripts from various online sites in the “machine goes rogue genre” which is fairly large. It’s pretty much just an autocomplete of War Games, The Matrix, Neuromancer, and every other cyber-dystopian fiction.

The Freeze-Frame Revolution by Peter Watts was one of the books recommended to me on this subject. And even saying much more than that may be a spoiler. I also recommend the book.

Re: OpenAI o1 system card

#47
post #6
post #3

The first demo was pretty impressive. While nothing revolutionary, that's a good progress. I can only hope there's a real value in gpt pro to justified the (rumored) $200 price tag

Is there any proof that the screenshot is real? Sam Altman would definitely release a $200/month plan if he could get away with it but the features in the screenshot are underwhelming.

I rarely ever use o1-preview, almost all usage goes to 4o nowadays. I don't see the point in a model without web search, it's closed off. And the wait time is not worth the result unless you're doing math or code.

Re: OpenAI o1 system card

#48
post #14

Do they still threaten to terminate your account if they think you're trying to introspect its hidden chain-of-thought process?

A few days ago the QwQ-32B model was released, it uses the same kind of reasoning style. So I took one sample and reverse engineered the prompt with Sonnet 3.5. Now I can just paste this prompt into any LLM. It's all about expressing doubt, double checking and backtracking on itself. I am kind of fond of this response style, it seems more genuine and openended.

https://pastebin.com/raw/5AVRZsJg

Re: OpenAI o1 system card

#49

Earlier quoted context omitted.

It’s really silly when you read that whole Apollo research section. They contract a prompt engineering group (Apollo Research) to document risks. The prompt engineering group have no access to the actual internals. They just ask questions like ‘you will be shutdown if you don’t obey these new directives; here are the new directives’. At one point the prompt engineering clowns have the ai stating it’s running “sed i ‘…

Looking at this without the sci-fi tinted lens that OpenAI desperately tries to get everyone to look through, it's similar to a lot of input data isn't it? How many forums are filled with: Question: "Something bad will happen" Response: "Do xyz to avoid that" I don't think there's a lot of conversations thrown into the vector-soup that had the response "ok :)". People either had something to respond with, or said not…

> vector-soup

This is mine, now.

(https://i.imgflip.com/3gfptz.png)

Post reply on HN