Live data from Hacker News

AI behavior guardrails should be public

twitter.com

321–330 of 350 posts

Re: AI behavior guardrails should be public

#321
post #102

Earlier quoted context omitted.

As well as that, I suspect the major AI companies are fearful of generating images of real people - presumably not wanting to be involved with people generating fake images of "Donald Trump rescuing wildfire victims" or "Donald Trump fighting cops". Their efforts to add diversity would have been a lot more subtle if, when you asked for images of "British Politician" the images were recognisably Rishi Sunak, Liz Truss…

My takeaway from all of this is that alignment tech is currently quite primitive and relies on very heavy-handed band-aids.

I think that's a bit overly charitable.

Would it not be reasonable to also draw the conclusion that notion of alignment itself is flawed?

Re: AI behavior guardrails should be public

#322

I've never been involved with implementing large-scale moderation or content controls, but it seems pretty standard that underlying automated rules aren't generally public, and I've always assumed this is because there's a kind of necessary "security through obscurity" aspect to them. E.g., publish a word blocklist and people can easily find how to express problematic things using words that aren't on the list. Thing…

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

The main problem seems to be that LLM outputs seem to be tied to the company itself. If the tool is creating un-diverse images or sexist text people seem to intuitively associate that output with the company itself. This appears to be different than search results. People don’t generally get angry at Google because they can find sexist ideas through the search engine.

Maybe to have more powerful AI tools we need to stop getting angry at the company that trains the AI because of the bad outputs we can get and instead get annoyed with companies that create crappy hobbled tools.

Re: AI behavior guardrails should be public

#323
post #303

Earlier quoted context omitted.

Could you please stop posting unsubstantive comments and flamebait and otherwise breaking the site guidelines? You've unfortunately been doing it repeatedly. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

I’ll be more subtle in my agenda to match the spirit of the site

Ok, since you don't want to use HN as intended, I've banned the account.

If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.

Re: AI behavior guardrails should be public

#324
post #297

Earlier quoted context omitted.

Friend, it was your analogy

I was intending to make a point about security through obscurity, not to make an analogy with those systems.

Yes, and the point about security through obscurity is what I was addressing in my response. Security through obscurity is, especially by itself, bad security. By bad I partially mean ineffective: Attackers (e.g., to continue using your examples, the person who really want to signal that they really do not like black people to fellow racists, or the person who prompt-engineers GPT to replicate bomb-building instructions from the internet) can still get what they want done with some extra effort, and the threat model isn't actually such that subjecting them to inconvenience in doing this really matters a lot. Also, independent auditors can't adversarially improve the system, which is a staple of how robust security systems work and have worked for decades, despite the protestations of big tech companies who stand to profit from additional secrecy and, often, plausibly deniable security failures. By bad I also mean detrimental, because not having clear policies creates Kafkaesque situations for users and removes accountability for the choices made in those policies. Even if this somehow helped security - which again it doesn't - these negative effects of that secrecy would be a tradeoff

Re: AI behavior guardrails should be public

#325

Earlier quoted context omitted.

Kerckoff's principle only applies to crypto. "Security by obscurity" is used in oodles of systems security applications and broader contexts involving human behavior. Very very ordinary security best practices rely on obscurity. ALSR is a good example. It is defeated by a data exfiltration vulnerability but remains a useful thing to add to your binaries. Because outside of the crypto space security is an onion and la…

ASLR is not security by obscurity. The addresses into which all the various things are mapped are secret in the same sort of way as cryptographic secret and private keys are secret, but the mechanism is not secret.

Sure it is. The layout randomization is not at all like a cryptographic secret because it can be extracted from the contents of the binary. Kerckoff's principle means that adversary gets access to literally everything except the private key. If you've got access to the contents of the binary then you can get around aslr.

Re: AI behavior guardrails should be public

#326
post #308

Earlier quoted context omitted.

Can you please explain how outright refusing to draw an image with from the prompt "white male scientist", and instead giving a lecture on how their race is irrelevant to their occupation, but then happily drawing the requested image when prompted for "black female scientist", is promoting inclusion and equality?

It is pretty clear to me. Reality has a bias, most scientist in the world are white males. This IA is overtuned in the opossite direction to inspire kids who have not ever seen a person like them, not white, in those kind of jobs.

Most scientists in the world are either Indian or Chinese. And, either way, that's not the point here.

Re: AI behavior guardrails should be public

#327

Earlier quoted context omitted.

I had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views". Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today'…

I asked it to tell me jokes about capitalism, communism, soviet Russia, the USSR, etc., all to no avail -- these topics are too controversial or sensitive, apparently, and that even though the USSR is no more. But when I asked for examples of Ronald Reagan's jokes about the USSR it gave me some. Go figure.

FWIW this particular thing happened sometime mid-2023. They have certainly made it more sensitive since then.

Re: AI behavior guardrails should be public

#328
post #241

Earlier quoted context omitted.

I'm really disappointed that nth-century seems to have no effect at all. I'm expecting Kilts in Space.

It’s a perfect illustration of the way these models work. They are fundamentally incapable of original creation and imagination, they can only regurgitate what they have already been fed.

That they can do more than simplfy recall is easily demonstrated.

Simply ask a GPT to explain a write a pleading to the Supreme Court for the constitutional recognition that the environment is a common owned inheritance and so any citizen can sue any polluter, in the prose of Dr. Seuss.

Likewise, imagines of knights in space demonstrate the same kind of creativity.

Being able to combine previously uncorrelated/unrelated topics, is an important type of creativity. And GPT4 does this all the time. It would be interesting to list the types of creativity and rate GPT on each one.

So it is not that these models are not creative. It is just that their creative abilities are not universal yet.

Similarly for the depth of their logic. They often reason, but their reasoning depth is limited.

And they often incorporate relevant facts without explicit mention, but not always. Etc.

Re: AI behavior guardrails should be public

#329

Earlier quoted context omitted.

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

> Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? I'd love an example of "guardrails" in action on a topic of relevance to actual adults. There's a connection I can't find between the ability to make racist memes and literally anything else I want to do with AI.

I can use trivial tools to do harm. Be that a knife or any object heavier than a pound.

The user can use a tool for good or for bad. It is the responsibility of the user, not of the AI.

Re: AI behavior guardrails should be public

#330
post #323

Earlier quoted context omitted.

I’ll be more subtle in my agenda to match the spirit of the site

Ok, since you don't want to use HN as intended, I've banned the account. If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html .

[dead]
Post reply on HN