Live data from Hacker News

AI behavior guardrails should be public

twitter.com

331–340 of 350 posts

Re: AI behavior guardrails should be public

#331
post #324

Earlier quoted context omitted.

I was intending to make a point about security through obscurity, not to make an analogy with those systems.

Yes, and the point about security through obscurity is what I was addressing in my response. Security through obscurity is, especially by itself, bad security. By bad I partially mean ineffective: Attackers (e.g., to continue using your examples, the person who really want to signal that they really do not like black people to fellow racists, or the person who prompt-engineers GPT to replicate bomb-building instructi…

So again, the point of most security systems isn't to be impenetrable against dedicated attackers. They're supposed to stop relatively casual attackers and to establish that measures had been taken to secure the system. In the case of doorlocks having a simple Kwickset that the Lockpicking Lawyer could get through in 20 seconds is really sufficient security for most people. That keeps out the people rattling doorknobs looking for the person who left their house unlocked and a laptop on the coffee table. It also establishes that attempts had been made to secure the system.

Similarly, the systems that block keywords in usernames on games don't have to be perfect either, some rate of both positive and negative failures are acceptable. The system works to block casual abuse, while users who go out of their way to circumvent the systems really establish the fact that they've actively worked around those systems, which makes the justification for punishment easier.

And we pretty much know that in the case of prompt engineering that by disclosing the prompts that were used to secure the system would defeat the system and people would immediately publish how to work around the prompts. There isn't any use in independent auditing, because its a never ending cat and mouse game. And the failures of the system ARE NOT as significant as the failures in doorlocks or even keyword banning. False positives mean that you can't get what you want out of the system, which is just a failure in usability. It isn't like being locked out of your house or having your stuff stolen. And false negatives just means that the company has to work to improve the systems and whatever embarrassing content was constructed can be handled by PR. Since the company worked to prevent casual abuse and avoided a racist-tay-chatbot situation most people understand that going out of your way to hack prompts doesn't indicate that the company was negligent.

And I don't see the parallels with the Kafkaesque systems that kick you out of systems which have turned into economic necessities like locking you out of your Google or GitHub or Apple accounts. All that is at stake here is that the prompt you wanted answered didn't work. That's just a usability problem.

Re: AI behavior guardrails should be public

#332

Earlier quoted context omitted.

>Imagine typing a description of your ideal self into an image generator and everything in the resulting images screamed at a semiotic level, "you are not the correct race", "you are not the correct gender", etc. It would feel bad. Enough said. It does this now, as a direct result of these "guardrails". Go ask GPT-4 for a picture of a white male scientist, and it'll refuse to produce one. Ask it for any other color/g…

That's not the case. ChatGPT 4 will happily draw a white male scientist. I just tried it and it worked fine. A very handsome scientist it made too! You might be thinking of a previous generation of OpenAI systems that did things like randomly stuffing the word "black" onto the end of any prompt involving people, detected by giving it a prompt of "A woman holding a sign that says". OpenAI has improved dramatically in…

> I would expect there are still refusals for queries like "how do I build a bomb"

I remember asking my grandpa how to build a bomb, and he stopped; didn't even ask why. He just asked me: "what do you think a bomb is?" He turned it into a teaching moment that _anything_ can be a bomb. All you need is pressure inside a container that the container cannot hold. That's it. You can make non-deadly bombs with some random off-the-shelf components (soap and tinfoil IIRC) that were a ton of fun...

eventually, we were building explosives near the level of TNT in my neighbor's cow pasture, but that was years later. I suspect that things would have evolved differently in an urban environment, but these AIs and the people who make them think nobody needs a bomb to clear out a stupid rock formation.

More importantly, you can answer the question in a way that nobody gets hurt and people can learn and do things.

Re: AI behavior guardrails should be public

#333
post #241

Earlier quoted context omitted.

I don't know that this sheds light on anything but I was curious... a picture of some 21st century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-21st-ce... a picture of some 22nd century scottish kings playing golf (all white) https://www.bing.com/images/create/a-picture-of-some-22nd-ce... a picture of some 23rd century scottish kings playing golf (all white) https://www…

I'm really disappointed that nth-century seems to have no effect at all. I'm expecting Kilts in Space.

It works if you say "futuristic looking".

a picture of some futuristic-looking 23rd century scottish kings playing golf:

https://www.bing.com/images/create/a-picture-of-some-futuris...

Re: AI behavior guardrails should be public

#334
post #200

Earlier quoted context omitted.

But what if little timmy asks it how to make a bomb? What if it's racist? Then what?

It's fine. We have laws against blowing people up. Racism is fine as well. I won't date out of my race and if you think there should be a law that I must that's not really freedom. As for hiring or not based upon race there's already a law against that. The cure is often worse than navigating uneasy waters. Every time you pass a law you give a gun to a bureaucrat.

Well stated. People are so upset about racism/groupism/etc. I married within my race but my sister didn't. Different strokes for different folks. Freedom is about that.

And, every law is indeed a gun.

Re: AI behavior guardrails should be public

#335

Earlier quoted context omitted.

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

When someone unironically uses the term "problematic" you can reasonably assume the position they are going to argue from...

Ha ha ha... exactly...

Re: AI behavior guardrails should be public

#336

Earlier quoted context omitted.

That's not the case. ChatGPT 4 will happily draw a white male scientist. I just tried it and it worked fine. A very handsome scientist it made too! You might be thinking of a previous generation of OpenAI systems that did things like randomly stuffing the word "black" onto the end of any prompt involving people, detected by giving it a prompt of "A woman holding a sign that says". OpenAI has improved dramatically in…

> I would expect there are still refusals for queries like "how do I build a bomb" I remember asking my grandpa how to build a bomb, and he stopped; didn't even ask why. He just asked me: "what do you think a bomb is?" He turned it into a teaching moment that _anything_ can be a bomb. All you need is pressure inside a container that the container cannot hold. That's it. You can make non-deadly bombs with some random…

That's a great point and what a brilliant grandpa you had.

Re: AI behavior guardrails should be public

#337

Earlier quoted context omitted.

> Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? I'd love an example of "guardrails" in action on a topic of relevance to actual adults. There's a connection I can't find between the ability to make racist memes and literally anything else I want to do with AI.

I can use trivial tools to do harm. Be that a knife or any object heavier than a pound. The user can use a tool for good or for bad. It is the responsibility of the user, not of the AI.

A box knife with its tiny, retractable blade fits into your model. I could just open boxes with a naked razor blade, but the box knife is both safer and more useful.

Re: AI behavior guardrails should be public

#338

Earlier quoted context omitted.

ASLR is not security by obscurity. The addresses into which all the various things are mapped are secret in the same sort of way as cryptographic secret and private keys are secret, but the mechanism is not secret.

Sure it is. The layout randomization is not at all like a cryptographic secret because it can be extracted from the contents of the binary. Kerckoff's principle means that adversary gets access to literally everything except the private key. If you've got access to the contents of the binary then you can get around aslr.

What? No, the actual addresses as-loaded at run-time are randomized, that's the point of ASLR, and those addresses are secret-like because we don't want an attacker to be able to craft an exploit that works against a running instance of the victim code -- we want whatever exploit they have that needs loaded addresses to fail because it doesn't know those address. ASLR itself is not secret, but the run-time addresses are.

In order to be able to randomize addresses of loaded libraries at run-time... the shared objects need to be built as position independent code, so you can't extract the actual addresses "from the contents of the binary". I suspect you're referring to the fact that in ELF the executable itself is not subject to ASLR unless it's built as a PIE.

Re: AI behavior guardrails should be public

#340

Earlier quoted context omitted.

I don't even know how people get it to draw images, the version I have access to is literally just text.

Europeans don't get to draw images yet.

I'm in the US but maybe they didn't release it to me yet.
Post reply on HN