Live data from Hacker News

Ask HN: Have top AI research institutions just given up on the idea of safety?

news.ycombinator.com

51–60 of 99 posts

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#51
post #13

Are you asking about top AI research institutions or leading AI businesses? There’s tons of work in research communities.

Where can I find some of these researches? Any links or pointers are very much appreciated.

Everything I find by searching is marketing BS, or the same half-baked prompt injection protection that only works for cherry picked problems.

Really need some help here finding the right communities.

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#52

Isn't AI safety mostly a marketing thing? Like, we employ these safety people to make sure our chat bot does not turn into Skynet, implying the chat bot could turn into Skynet i.e. it's powerful and magic and please give us money. Maybe the text prediction programs are too familiar to people for the Skynet marketing to bite like it used to. Or maybe it was not just a marketing thing and the AI bros really did believe…

> we employ these safety people to make sure our chat bot does not turn into Skynet

i think it's mostly about not showing up in some NYT article titled "look what crazy thing i got this AI to say". There were a bunch of those early on and it really hurt the cause. Microsoft had some famous ones, even prior to chatgpt, where the AI got pretty testy in the chat.

https://en.wikipedia.org/wiki/Tay_(chatbot)

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#53
post #20

"safe" is such a subjective concept to begin with, have any of the model providers ever defined what they mean by "safe"? It doesn't mean much to me if a safe model is one that does not output the recipe for mustard gas, that information is trivially available elsewhere. Or, is a safe model one that doesn't come off as racist? Ok but i would classify that as unoffensive instead of safe but I admit definitions of word…

What if I tell the model to go commit fraud or crimes and it complies? What if users are having psychotic episodes driven by their interactions with the model? Just because safety is a hard and messy problem doesn't mean we should just wash our hands of it.

But think about how much money there is to be made by just ignoring it all!

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#54
Also an outsider, but my perspective is that "safety" has always been a nebulous term for a variety of concepts. No AI institution will ever give up on alignment because "the AI does what you want it to" is a pure functionality thing. On the other end of the scale there's a censorship aspect to it where models will refuse to provide wikipedia level information because it's "dangerous". The latter is very much subject to the whims of the labs, politicians, journalists, etc.

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#55
post #20

"safe" is such a subjective concept to begin with, have any of the model providers ever defined what they mean by "safe"? It doesn't mean much to me if a safe model is one that does not output the recipe for mustard gas, that information is trivially available elsewhere. Or, is a safe model one that doesn't come off as racist? Ok but i would classify that as unoffensive instead of safe but I admit definitions of word…

My preferred version of "safe" is "in its actions considers and mostly upholds usually unstated constraints like 'don't kill unless necessary', 'keep Earth inhabitable', 'avoid toppling society unless really well justified for the greater good', etc. The kind of framing that was prevalent pre-ChatGPT. Not terribly relevant for a chat software, but increasingly important as chat models turn into agents.

Of course once you have that framing, additional goals like "don't give people psychosis", "don't give step-by-step instructions on making explosives, even if wikipedia already tells you how to do it" or "don't harm our company's reputation by being racist" are conceptually similar.

On the other hand "don't make weapon systems" or "never harm anyone" might not be viable goals. Not only because they are difficult to impossible to define, but also because there is huge financial and political pressure not to limit your AI in that way (see Anthropic)

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#56

Humans can't develop safety until there is enough blood in the streets. Only issue with AI is that threshold may come at a point where its too far gone to recover. But humans can't put in seatbelts until we're losing 40k people per year in car crashes. Unfortunately its just how we're wired. Those that are careful are outcompeted by the brash and the fast-moving, until the relative value of moving fast is removed, th…

If we had rules like that in the past we never would have had the industrial revolution.

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#57

I was the author for the practitioners implementation section for the IEEE 7010 standard for assessing human impact from AI software https://standards.ieee.org/ieee/7010/7718/ I also worked closely with Jack Clark at OpenAI before he disappeared on all these issues as CTO back in 2018 There are literally zero “AI labs” that have ever cared about “safety” none of them have ever done anything tangible with any kind of…

Can you elaborate this part please?

> The concept itself doesn’t even make sense if you fully understand the intersectional scope of technology and society Societies demands are the things that are unsafe not the technologies themselves

Where can I learn more about it?

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#58
post #20

"safe" is such a subjective concept to begin with, have any of the model providers ever defined what they mean by "safe"? It doesn't mean much to me if a safe model is one that does not output the recipe for mustard gas, that information is trivially available elsewhere. Or, is a safe model one that doesn't come off as racist? Ok but i would classify that as unoffensive instead of safe but I admit definitions of word…

>Is a safe model one that refuses to produce code for a weapons system? Well.. does a PID controller count? I can use that to keep a gun pointed at a target or i can use that to prevent a baby rocker from falling over.

I've been using LLMs for some cyber-y tasks and this is exactly how it ends up going. You can't ask "hack this IP" (for some models), but more discrete tasks it'll have no such qualms.

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#59

I was the author for the practitioners implementation section for the IEEE 7010 standard for assessing human impact from AI software https://standards.ieee.org/ieee/7010/7718/ I also worked closely with Jack Clark at OpenAI before he disappeared on all these issues as CTO back in 2018 There are literally zero “AI labs” that have ever cared about “safety” none of them have ever done anything tangible with any kind of…

if would be super helpful if you could give the elevator pitch version of what a safe AI is.

Re: Ask HN: Have top AI research institutions just given up on the idea of safety?

#60

It’s just still so trivial to jailbreak even the latest Anthropic models (via api, and not talking about the silly ENI or Pliny breaks) I don’t understand where the safety teams are doing their work. Is it in the default chat-trained model?

It's more of a research program than a product feature. No-one knows how to fully prevent a model from responding based on what's in its base training data, which is what you're seeing with jailbreaks.

And going to one of the roots of the issue - the base training data - comes with its own set of unsolved challenges, not least of which is the unavoidable subjectivity of what is or isn't "safe".

Post reply on HN