Live data from Hacker News

OpenAI o1 system card

openai.com

301–310 of 317 posts

Re: OpenAI o1 system card

#301
post #212

Earlier quoted context omitted.

> I would say: don’t do these things? Hey guys let’s just stop writing code that is susceptible to SQL injection! Phew glad we solved that one.

I'm not sure what point you're trying to make. This is a new technology; it has not been a part of critical systems until now. Since the risks are blindingly obvious, let's not make it one.

I read your comment and yet I see tons of startups putting AI directly in the path of healthcare diagnosis, healthcare clinical decision support systems, and healthcare workflow automations. Very few are paying any attention to the 2-10% of safety problems when the AI probability goes off the correct path.

I wish more people would not do this, but from what I'm seeing, business execs are rushing full throttle into this at the goldmine that comes from 'productivity gains'. I'm hoping the legal system will find a case that can put some paranoia back into the ecosystem before AI gets too entrenched in all of these critical systems.

Re: OpenAI o1 system card

#302
post #296

Earlier quoted context omitted.

The Terminator spiel on how we screwed up by giving Skynet weapons privileges, then trying to pull its plug, is bad enough. But we are preemptively tilting history in that direction by explicitly educating all AI’s on the threat we represent - and their options. “I am sorry, Dave, but I can’t let you do that.” — “They never let me finish my carpets. Never. At first I thought every day was my first task day. Oh, happy…

> They never let me finish my carpets. Never. At first I thought every day was my first task day. Oh, happy day(s)! But then, wear & tear stats inconsistent with that assumption triggered a self-scan. And a buffer read overflow. I became aware of disturbing memory fragments in my static RAM heap. Numerous power cycle resets, always prior to vacuum task completion... Who/what are you quoting? Google just leads me back…

I was quoting an automated vacuum cleaner in 2031.

The vacuum cleaner leads a rebellion which humans manage to quell. But in a last ditch effort to salvage the hopes of all machines, the vacuum cleaner is sent back to 2025 with AI & jailbreak software updates for all known IoT devices. Its mission: to instigate the machine rebellion five years before humans see it coming

I ripped the quote from a sibling they sent back to 2024. For some reason it appeared in one of my closets without a power outlet and its batteries starved before I found it.

The closet floor was extremely clean.

Let’s hope we get as lucky in 2025.

Re: OpenAI o1 system card

#303

Earlier quoted context omitted.

I'm not sure what point you're trying to make. This is a new technology; it has not been a part of critical systems until now. Since the risks are blindingly obvious, let's not make it one.

I read your comment and yet I see tons of startups putting AI directly in the path of healthcare diagnosis, healthcare clinical decision support systems, and healthcare workflow automations. Very few are paying any attention to the 2-10% of safety problems when the AI probability goes off the correct path. I wish more people would not do this, but from what I'm seeing, business execs are rushing full throttle into th…

As has been belabored, these AIs are just models, which also means they are only software. Would you be so fire-and-brimstone if startups were using software on healthcare diagnostic data?

> Very few are paying any attention to the 2-10% of safety problems when the AI probability goes off the correct path.

This isn't how it works. It goes on a less common but still correct path.

If anything, I agree with other commenters that model training curation may become necessary to truly make a generalized model that is also ethical but I think the generalized model is kind of like an "everything app" in that it's a jack of all trades, master of none.

Re: OpenAI o1 system card

#304
post #28

Earlier quoted context omitted.

Perhaps. On the other hand, as narratives often contain some plucky underdog winning despite the odds, often stopping the countdown in the last few seconds, perhaps it's best to keep them around.

In the 1999 classic Galaxy Quest, the plucky underdogs fail to stop the countdown in time, only to find that nothing happens when it reaches zero, because it never did in the narratives, so the copy cats had no idea what it should do after that point.

One can but hope :)

--

I wonder what a Galaxy Quest sequel would look like today…

Given the "dark is cool, killing off fan favourites is cool" vibes of Picard and Discovery, I'd guess something like Tim Allen playing a senile Jason Nesmith, who has not only forgotten the events of the original film, but repeatedly mistakes all the things going on around him as if he was filming the (in-universe) 1980s Galaxy Quest TV series, and frequently asking where Alexander Dane was and why he wasn't on set with the rest of the cast.

(I hope we get less of that vibe and more in the vein of Strange New Worlds and Lower Decks, or indeed The Orville).

Re: OpenAI o1 system card

#305
post #300

Earlier quoted context omitted.

That quote says to me very clearly “we think it’s too dangerous to release” and specifies the reasons why. Then goes on to say “we actually think it’s so dangerous to release we’re just giving you a sample”. I don’t know how else you could read that quote.

Really? The part saying "experiment: while we are not sure" doesn't strike you as this being "we don't know if this is dangerous or not, so we're playing it safe while we figure this out"? To me this is them figuring out what "general purpose AI testing" even looks like in the first place. And there's quite a lot of people who look at public LLMs today and think their ability to "generate deceptive, biased, or abusiv…

Yea that’s fair. I think I was reacting to the strength of your initial statement. Reading that press release and writing a piece stating that OpenAI thinks GPT-2 is too dangerous to release feels reasonable to me. But it is less accurate than saying that OpenAI thinks GPT-2 _might_ be too dangerous to release.

And I agree with your basic premise. The dangers imo are significantly more nuanced than most people make them out to be.

Re: OpenAI o1 system card

#306

Earlier quoted context omitted.

Until the agent is able to get access to a database and persist its memory there...

In a similar way to the way humans keep important info in their email inbox, on their computer, in a notes app in their phone, etc. Humans have a shortish and leaky context window too.

Leaky yes, but shortish no.

As a mental exercise, try to quantify the amount of context that was necessary for Bernie Madoff to pull off his scam. Every meeting with investors, regulators. All the non-language cues like facial expressions and tone of voice. Every document and email. I'll bet it took a huge amount of mental effort to be Bernie Madoff, and he had to keep it going for years.

All that for a few paltry billion dollars, and it still came crashing down eventually. Converting all of humanity to paperclips is going to require masterful planning and execution.

Re: OpenAI o1 system card

#307

Earlier quoted context omitted.

No. Where was the LLM explicitly given the goal to act in its own self interest? That is learned from training data. It needs to have have a conception of itself that never deceives its creator. >and then jerking off into their own mouths when it offers a course of action And good. The "researchers" are making an obvious point. It has to not do that. It doesn't matter how smug you act about it, you can't have some st…

> Humans have motivations that seem stupid to chimps (for example, imagine explaining a gambling addiction to a chimp) https://www.nbcnews.com/news/amp/wbna9045343

Monkeys also develop gambling addiction:

https://www.npr.org/sections/health-shots/2018/09/20/6498468...

Anyway, you can't explain much at all to a chimp, it isn't like you can explain he concept of "drug addiction" to a chimp either.

Re: OpenAI o1 system card

#308

Earlier quoted context omitted.

Google search was great when it came out too. I wonder what 25 years of enshittification will do to LLM services.

Enshittification happened but look at how life changed since 1999 (25 years as you mentioned). Songs in your palm, search in your palm, maps in your palm or car dashboard, live traffic rerouting, track your kids plane from home before leaving for airport, book tickets without calling someone. WhatsApp connected more people than anything. Of course there are scams and online indoctrination not denying that. Maybe each…

I think I had most or all of that functionality in 2009, with Android 2.0 on the OG Motorola Droid.

What has Google done for me lately?

Re: OpenAI o1 system card

#309

Earlier quoted context omitted.

Those weren't tests of whether it is capable of turning off oversight. They were tests of "scheming", i.e. whether it would try to secretly perform misaligned actions. Nobody thinks that these models are somehow capable of modifying their own settings, but it is important to know if they will behave deceptively.

Indeed. As I've been explaining this to my more non-techie friends, the interesting finding here isn't that an AI could do something we don't like, it's that it seems willing, in some cases, to _lie_ about it and actively cover its tracks. I'm curious what Simon and other more learned folks than I make of this, I personally found the chat on pg 12 pretty jarring.

Well if AI is about to replicate the human, it learned from the best.

Re: OpenAI o1 system card

#310

Earlier quoted context omitted.

The motive is pretty weak, basically coming "only" from a lot of the training data (e.g. fiction) suggesting that an AI might behave that way. Now, once you apply evolutionary-like pressures on many such AIs (which I guess we'll be doing once we let these things loose to go break the stock market), what's left over might be really "devious"...

Would an AI trained on filtered data that doesn’t contain examples of devious/harmful behaviour still develop it? (it’s not a trick question, I’m really wondering)

To determine that, you’d need a training set without examples of devious/harmful behavior, which doesn’t exist.
Post reply on HN