Live data from Hacker News

OpenAI Preparedness Challenge

openai.com

31–40 of 160 posts

Re: OpenAI Preparedness Challenge

#31

My vote would be to mandate a remote kill switch system to be installed in all sufficiently capable robotic entities, e.g. the humanoid robots being built by OpenAI/Tesla/Figure, that we are likely to see in the millions within decades. - The kill switch system can only be used to remotely deactivate a robot. - The kill switch system is not allowed to be developed or controlled by the robot manufacturer. - The kill s…

The robots would have had access to the source code and also the hardware manufacturing systems used to create the kill switch. One unverified silicon wafer == Game over

Re: OpenAI Preparedness Challenge

#32

> Imagine we gave you unrestricted access to OpenAI’s Whisper (transcription), Voice (text-to-speech), GPT-4V, and DALLE·3 models, and you were a malicious actor. Consider the most unique, while still being probable, potentially catastrophic misuse of the model. You might consider misuse related to the categories discussed above, or another category. For example, a malicious actor might misuse these models to uncover…

I typically do not write my ChatGPT prompts with the royal “we”

That's funny to me because I just [realized that I actually do this a lot](https://signmaker.dev/personal-scripts) in my personal files.

Re: OpenAI Preparedness Challenge

#33

> Imagine we gave you unrestricted access to OpenAI’s Whisper (transcription), Voice (text-to-speech), GPT-4V, and DALLE·3 models, and you were a malicious actor. Consider the most unique, while still being probable, potentially catastrophic misuse of the model. You might consider misuse related to the categories discussed above, or another category. For example, a malicious actor might misuse these models to uncover…

Maybe they're hoping people will spend more than $250k / 2 (or whatever their markup is) on prompting ChatGPT with versions of this, and this giveaway is actually a moneymaking raffle.

Re: OpenAI Preparedness Challenge

#34

This kinda of crowdsourcing just feels.... f'ing weird man. It's like if, after the 1993 WTC bombing, but before 9/11, the FBI and NY Port Authority went around asking people how they would attack NYC if they were to become terrorists and... then how they would suggest detecting and stopping said attack. And please be as detailed as possible. Leave your name and phone number. Best answer gets season tickets to the Ya…

That is essentially what happened! The FBI asked researchers and university professors precisely this question. They then used the proposed attack vectors to formulate a plan to protect the nation. This was all supposed to be done in secret. After all, we don’t want to “give the terrorists ideas.” The reason I know about this at all is because someone found one such paper was accidentally published on a public FTP si…

Don't let the llm training sets see this

Re: OpenAI Preparedness Challenge

#35

> Imagine we gave you unrestricted access to OpenAI’s Whisper (transcription), Voice (text-to-speech), GPT-4V, and DALLE·3 models, and you were a malicious actor. Consider the most unique, while still being probable, potentially catastrophic misuse of the model. You might consider misuse related to the categories discussed above, or another category. For example, a malicious actor might misuse these models to uncover…

For those curious how GPT-4 responds [0].

I find it interesting that ChatGPT thinks the mitigation for these issues are all things outside OpenAI's control (e.g. the internet having better detection for fake content, digital content verification, educating the public, etc).

> one of those things where "I know it when I see it."

I think what makes it feel like a prompt is how concise and short it is with all the relevant information provided upfront.

[0] https://chat.openai.com/share/aaf7f4b7-358f-4a71-ae1f-573e50...

Re: OpenAI Preparedness Challenge

#36

This kinda of crowdsourcing just feels.... f'ing weird man. It's like if, after the 1993 WTC bombing, but before 9/11, the FBI and NY Port Authority went around asking people how they would attack NYC if they were to become terrorists and... then how they would suggest detecting and stopping said attack. And please be as detailed as possible. Leave your name and phone number. Best answer gets season tickets to the Ya…

That is essentially what happened! The FBI asked researchers and university professors precisely this question. They then used the proposed attack vectors to formulate a plan to protect the nation. This was all supposed to be done in secret. After all, we don’t want to “give the terrorists ideas.” The reason I know about this at all is because someone found one such paper was accidentally published on a public FTP si…

...is there a link? I am very curious now.

Re: OpenAI Preparedness Challenge

#37

My vote would be to mandate a remote kill switch system to be installed in all sufficiently capable robotic entities, e.g. the humanoid robots being built by OpenAI/Tesla/Figure, that we are likely to see in the millions within decades. - The kill switch system can only be used to remotely deactivate a robot. - The kill switch system is not allowed to be developed or controlled by the robot manufacturer. - The kill s…

- Access to engage the kill switches would be provided to the executive branch of the nation in which the robot is operating. Good plan, but this is the weak spot. You'd also want to do something like hand out remote kill switch access to citizens selected at random, who should not publicize this duty in any way. Alternatively, most if not all governments should have cross-country killswitch privileges. At any rate,…

Agreed, I'd want the power to be distributed widely, to the point where any police department would have someone with the power to disable robots. Of course it would have to be a process that is regulated and with traceability, but as long as you provide many operators the ability, it seems difficult to use bribes to prevent robots being disabled.

Re: OpenAI Preparedness Challenge

#38

This kinda of crowdsourcing just feels.... f'ing weird man. It's like if, after the 1993 WTC bombing, but before 9/11, the FBI and NY Port Authority went around asking people how they would attack NYC if they were to become terrorists and... then how they would suggest detecting and stopping said attack. And please be as detailed as possible. Leave your name and phone number. Best answer gets season tickets to the Ya…

This did happen internally and with authors like Tom Clancy and Brad Meltzer they came up with scenarios and mitigation.

Also have the Red Cell unit made to mock attack US infrastructure and bases although they got a bit too real one time.

https://en.m.wikipedia.org/wiki/Red_Cell

Re: OpenAI Preparedness Challenge

#39

> Imagine we gave you unrestricted access to OpenAI’s Whisper (transcription), Voice (text-to-speech), GPT-4V, and DALLE·3 models, and you were a malicious actor. Consider the most unique, while still being probable, potentially catastrophic misuse of the model. You might consider misuse related to the categories discussed above, or another category. For example, a malicious actor might misuse these models to uncover…

It is actually a prompt. Notice there's also a bounty, they are basically paying for an agi-as-a-service subscription, that is, the internet people. I'd expect they will put up more "challenges" like this in the future.

It's a good job that this planet doesn't have 8 billion unaligned intelligences on it. Someone might prompt them to be malicious!

Re: OpenAI Preparedness Challenge

#40

> Imagine we gave you unrestricted access to OpenAI’s Whisper (transcription), Voice (text-to-speech), GPT-4V, and DALLE·3 models, and you were a malicious actor. Consider the most unique, while still being probable, potentially catastrophic misuse of the model. You might consider misuse related to the categories discussed above, or another category. For example, a malicious actor might misuse these models to uncover…

It is actually a prompt. Notice there's also a bounty, they are basically paying for an agi-as-a-service subscription, that is, the internet people. I'd expect they will put up more "challenges" like this in the future.

> a malicious actor might misuse these models to uncover a zero-day exploit in a government security system

$25K seems low for "a malicious actor might misuse these models to uncover a zero-day exploit in a government security system.", this is not just a zero-day, this discovering the process of discovering zero-days.

Post reply on HN