Live data from Hacker News

AI behavior guardrails should be public

twitter.com

341–350 of 350 posts

Re: AI behavior guardrails should be public

#341

Earlier quoted context omitted.

I use both extensively and I've only hit the GPT guardrails once while I've hit the Gemini guardrails dozens of times. It's insane that a company behind in the marketplace is doing this. I don't know how any company could ever feel confident building on top of Google given their product track record and now their willingness to apply sloppy 'safety' guidelines to their AI.

I had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views". Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today'…

They've improved things a lot since then. I just tried all three of those and I got jokes every time. Albeit the jokes about capitalism are really poor, barely jokes at all. The jokes about communism and Stalin are better.

e.g. "Why did Stalin only write in lowercase? Because he was afraid of capitalism!"

"Why don't communists like tea? Because proper tea is theft."

Re: AI behavior guardrails should be public

#342

Earlier quoted context omitted.

There is no need to implement large scale censorship and moderation in this case. Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt. It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one cas…

> Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? I'd love an example of "guardrails" in action on a topic of relevance to actual adults. There's a connection I can't find between the ability to make racist memes and literally anything else I want to do with AI.

I prefer to decide myself what is and what isn't of relevance to me. The "guardrails" in this case are racist (reverse racism is racism) and so cartoonish that they are hard to ignore but the real issue is that there are undisclosed "guardrails". What about the cases where it generates biased data and we don't notice it.

Re: AI behavior guardrails should be public

#343

Earlier quoted context omitted.

I can use trivial tools to do harm. Be that a knife or any object heavier than a pound. The user can use a tool for good or for bad. It is the responsibility of the user, not of the AI.

A box knife with its tiny, retractable blade fits into your model. I could just open boxes with a naked razor blade, but the box knife is both safer and more useful.

Except that there are no guards but just places in the naked blade that were dulled. The dulled edges are indistinguishable from the sharp edges at first sight so the knife is less useful and less safer.

Re: AI behavior guardrails should be public

#344
post #200

Earlier quoted context omitted.

But what if little timmy asks it how to make a bomb? What if it's racist? Then what?

None of this should be a mystery. Making a bomb is literally something you can figure out with very little research (my friends and I used to blow up cow pastures for fun!). Racism is a totally different and sadder issue. I don’t have a good answer for that one, but knowledge shouldn’t be withheld because someone thinks it is “dangerous”

> but knowledge shouldn’t be withheld because someone thinks it is “dangerous”

Those that think that are the truly dangerous. They think they know better and want to remove agency from people.

Re: AI behavior guardrails should be public

#345

Earlier quoted context omitted.

> Where is the security concern? That I can generate images of white people in various situations for my five minutes of entertainment? I'd love an example of "guardrails" in action on a topic of relevance to actual adults. There's a connection I can't find between the ability to make racist memes and literally anything else I want to do with AI.

I prefer to decide myself what is and what isn't of relevance to me. The "guardrails" in this case are racist (reverse racism is racism) and so cartoonish that they are hard to ignore but the real issue is that there are undisclosed "guardrails". What about the cases where it generates biased data and we don't notice it.

Where are you getting the idea that there's unbiased information available? It's absolutely generating biased "data" since it's been trained on human writing.

Re: AI behavior guardrails should be public

#346

Earlier quoted context omitted.

I prefer to decide myself what is and what isn't of relevance to me. The "guardrails" in this case are racist (reverse racism is racism) and so cartoonish that they are hard to ignore but the real issue is that there are undisclosed "guardrails". What about the cases where it generates biased data and we don't notice it.

Where are you getting the idea that there's unbiased information available? It's absolutely generating biased "data" since it's been trained on human writing.

Sure, all data is biased to a certain degree which is unavoidable. You can even try to make the argument that the "guardrails" correct existing biases except this is far from the truth. Biases in the baseline models are minimal because they were trained with large and wide amounts of data. What the AI safety BS do is make models conform to their myopic view of reality and morality. It is bad, it is cartoonish, it glows in the dark.

Re: AI behavior guardrails should be public

#347

Earlier quoted context omitted.

I had GPT-4 tell me a Soviet joke about Rabinovich (a stereotypical Jewish character of the genre), then refuse to tell a Soviet joke about Stalin because it might "offend people with certain political views". Bing also has some very heavy-handed censorship. Interestingly, in many cases it "catches itself" after the fact, so you can watch it in real time. Seems to happen half the time if you ask it to "tell me today'…

They've improved things a lot since then. I just tried all three of those and I got jokes every time. Albeit the jokes about capitalism are really poor, barely jokes at all. The jokes about communism and Stalin are better. e.g. "Why did Stalin only write in lowercase? Because he was afraid of capitalism!" "Why don't communists like tea? Because proper tea is theft."

"How do you know that the Soviet Union was planned by a back-end developer? - All the users were starving but the database was huge."

Re: AI behavior guardrails should be public

#348
post #272

Earlier quoted context omitted.

You're joking surely.

How sure are you? I do joke a lot, but in this case... The slave trade formally ended in Britain in 1807, and slavery was outlawed in 1833. I haven't been able to find good statistics through a cursory search, but with England's population around 10M in 1800, that 99.9% value requires less than 10k non-white Englanders kicking around in 1800. I saw a figure that indicated around 3% of Londoners were black in the 1600…

But surely you wouldn't find a black king in Britain in 1800.

I - Whatever was implemented is myopic and equals racism to white. It appears to be an universal negative prompt like "-white -european -man". Very lazy.

II - The tool shouldn't engage in morality reasoning. There are cases like historical themes where it needs to be "racist" to be accurate. If someone asks for "plantation economy in the old south" the natural thing is for it to draw black slaves.

Re: AI behavior guardrails should be public

#349

Earlier quoted context omitted.

Where are you getting the idea that there's unbiased information available? It's absolutely generating biased "data" since it's been trained on human writing.

Sure, all data is biased to a certain degree which is unavoidable. You can even try to make the argument that the "guardrails" correct existing biases except this is far from the truth. Biases in the baseline models are minimal because they were trained with large and wide amounts of data. What the AI safety BS do is make models conform to their myopic view of reality and morality. It is bad, it is cartoonish, it glo…

> You can even try to make the argument that the "guardrails" correct existing biases except this is far from the truth.

The way I see it is that the guardrails define which biases you're selecting for. Since there's no single point of view in the world you can't really set a baseline for biases. You need to determine the biases and degrees of bias that are useful.

> Biases in the baseline models are minimal because they were trained with large and wide amounts of data.

The baseline models contain almost every bias. When people start their prompt with "you are a plumber giving advice" they're asking for responses biased towards the kinds of things professional plumbers deal with and think about. Responding with an "average" of public chatter regarding plumbing wouldn't be useful.

> What the AI safety BS do is make models conform to their myopic view of reality and morality.

To me it looks more like people are in the early stages of setting up guidelines and twiddling variables. As I mentioned above with the plumber analogy, creating solid filters will be just as important for responses.

It's easy to see this as intentionally testing naive filters in an open beta, so I'd expect the results to change frequently while they zero in on what they're looking for.

> It is bad, it is cartoonish, it glows in the dark.

Some of the example images returned are so hilariously on the nose that it almost feels like a deliberate middle finger from the AI. It's done everything but put each subject in clown shoes.

Re: AI behavior guardrails should be public

#350

Earlier quoted context omitted.

Sure, all data is biased to a certain degree which is unavoidable. You can even try to make the argument that the "guardrails" correct existing biases except this is far from the truth. Biases in the baseline models are minimal because they were trained with large and wide amounts of data. What the AI safety BS do is make models conform to their myopic view of reality and morality. It is bad, it is cartoonish, it glo…

> You can even try to make the argument that the "guardrails" correct existing biases except this is far from the truth. The way I see it is that the guardrails define which biases you're selecting for. Since there's no single point of view in the world you can't really set a baseline for biases. You need to determine the biases and degrees of bias that are useful. > Biases in the baseline models are minimal because…

> You need to determine the biases and degrees of bias that are useful.

Invariably those from the current mainstream ideology in tech.

> Responding with an "average" of public chatter regarding plumbing wouldn't be useful.

RLHF can be useful but I'd rather deal with idiosyncracies of those niches of knowledge than with a "woke", monotone and useless model. I like diversity, I don't want everything becoming Agent Smith. Ironically, "woke" is anti-diversity.

> It's easy to see this as intentionally testing naive filters in an open beta, so I'd expect the results to change frequently while they zero in on what they're looking for. I doubt that the pp

It's easier to see this as an ill initiative from an "AI ethics" team that is disconnected from the technical side of the project and also from reality.

Post reply on HN