Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

211–220 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#211

This is absolutely insane: Repro (de-identified): sample_dataset_group1.tsv - Geometry: Heatmap - X axis: frac_set set + condition (two columns → the "Add column" cross join) - Y axis: condition - Color: mean frac_set value, Sequential When the X axis is a cross join of two columns (the second added via "Add column"), the x-axis tick labels (frac_set_2, frac_set_3, frac_set_4, frac_set_5) render in a broken state, ro…

Here's one that was flagged for me: a question about a niche Reinforcement Learning paper from 2012

I've been reading the option-option model paper by David Silver. It appears that they achieved quite an effective result. Why hasn't there been more work on it since?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#212

Earlier quoted context omitted.

"Arms race" is the term used colloquially to describe the dynamic that emerges in "winner-take-all" markets. It seems that the frontier labs believe they're participants in a winner-take-all market. Therefore they're in "an arms race." Winner-take-all markets do not require that the winner literally destroys the losers, but only that the winner enjoys disproportionate returns compared to their actual superiority. Whe…

I don't know why you think I'm taking anything literally, cf. my first comment. I understand what a metaphorical arms race is. I don't think that Anthropic can forestall others' AI development by getting there first. It can't be literal destruction. It can't be economic destruction (some actors interested in it aren't motivated by money). What's left? I'm all ears. As far as naivete, wouldn't it be more naive to take…

> These are private companies, they are not vested with the necessary authority to destroy anything

You're pretty explicitly saying that dominating the competition is not the type of "destruction" necessary to qualify as an arms race.

> As far as naivete, wouldn't it be more naive to take their EA claims at face value, rather than the more realistic assumption that they like money?

Huh? Greed is – quite obviously – the major driving force behind the arms race. That is not a mitigation whatsoever.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#213
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

> If it was just plain monetary concerns and sabotage of competitors I'd almost be fine with it, but it seems they actively want to monopolize most of human progress in their enlightened hands

But that is “plain monetary concerns and sabotage of competitors”, they are just more ambitious than most people doing sabotage of competitors in the fields they hope to dominate by that tactic.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#214
I know this isn't going to be a popular take, but here goes anyway...

The complaints that Anthropic are routing your requests to a different model reminds me of an old Louis CK bit about airplane wifi. Clearly Anthropic was too aggressive with whatever guardrails they put in, but the response seems overly entitled to a model people didn't even know existed not that long ago.

https://youtube.com/watch?v=me4BZBsHwZs

Re: Anthropic apologizes for invisible Claude Fable guardrails

#215

Earlier quoted context omitted.

I don't know why you think I'm taking anything literally, cf. my first comment. I understand what a metaphorical arms race is. I don't think that Anthropic can forestall others' AI development by getting there first. It can't be literal destruction. It can't be economic destruction (some actors interested in it aren't motivated by money). What's left? I'm all ears. As far as naivete, wouldn't it be more naive to take…

> These are private companies, they are not vested with the necessary authority to destroy anything You're pretty explicitly saying that dominating the competition is not the type of "destruction" necessary to qualify as an arms race. > As far as naivete, wouldn't it be more naive to take their EA claims at face value, rather than the more realistic assumption that they like money? Huh? Greed is – quite obviously – t…

> I will charitably assume Anthropic does not intend to literally destroy anyone and merely wants to become an AGI monopoly.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#216
post #95

Earlier quoted context omitted.

Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…

[flagged]

> asking for domestic safety testing of frontier models only is not regulatory capture.

Yeah, asking for additional state-provided barriers to a market entry to a valuable market a provider already is one of a narrow few dominating only for firms that are a competitive threat is exactly regulatory capture.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#218
post #136

Earlier quoted context omitted.

Yes, that is basically the plan. It's based on the belief that unfettered AI would let anyone be a supervillain and destroy the world. There are enough would-be supervillains out there, but they rarely get far because they can't get teams of smart people to build doomsday machines for them. So the AI has to not let anyone do evil with it. Unfortunately, that won't feel very much like freedom.

It sounds like you might not agree with that belief. While I don't agree with their actions here, I do think there's sufficient reason to hold that belief. On some fronts (e.g. security, on which you've experienced more than me), I think there are surmountable challenges. But on other fronts (e.g. bio), a single errant actor could reasonably kill millions or billions of people with sufficiently powerful AI. We don't…

The model release cards for Opus have repeatedly and consistently stressed that the model doesn't have the fiddly know-how that's required to provide meaningful assistance in possibly dangerous subfields of biology. Mythos (Fable without the overly strict guardrails) has shown improvements in things like drug design, but even then the situation isn't really that different. This risk is ridiculously overblown, and the way to manage it sensibly is to introduce meaningful oversight for actors that seek to order the actual specialized materials involved (especially any synthetically generated genes/proteins/whatever).

Re: Anthropic apologizes for invisible Claude Fable guardrails

#220

Earlier quoted context omitted.

Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.

You'd be fine if the PRC gets to ASI first? That's an interesting opinion.

> You'd be fine if the PRC gets to ASI first?

How do rules that inhibit what AI can be sold on the US market (adding additional costs to trading in that market) do anything to inhibit a competing nation from reaching ASI first? Insofar as they inhibit anyone from reaching ASI, its firms whose primary commercial interest is selling AI services in the US market, not foreign threat actors except to the extent those two categories overlap.

Post reply on HN