Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

481–489 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#481

Earlier quoted context omitted.

> You can, in fact, opt out. You can opt out and do your damndest to stop what's happening, throw every cent you have at it, bend any ear that will listen, make use of the fact that your voice (as Anthropic leadership) has some meaningful weight. There are billions of people who have opted out of playing the game. Has the game stopped? Has any game stopped because the people not playing it decided that it ought to? O…

(I wrote a longer comment originally, but I think it would have fallen on deaf ears.) > The only variable of disagreement is around AI doom. The source of our disagreement seems to be your belief that somebody can either a) believe "AI doom" is inevitable , or b) not believe it's possible . This is an obvious false dichotomy that's stunting your ability to engage effectively with what I've written, and also stunting…

Then surely you can articulate how specifically – in a way that doesn't require large numbers of people acting against their own incentives – that AI doom is possible but not guaranteed.

I actually already know how you'd perceive Dario's opting out of the race because I already know how you perceive Dario's requests for regulation, which is the milder version of the same logic, and is vulnerable to the same cynical allegations of self-serving, which you've already expressed.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#482

Earlier quoted context omitted.

Those aren't values. Maybe goals or motivations but not values in any conceivable way, shape or form. This site is full of pod people I swear.

Maybe it would help if you shared your private personal definition of "value", since you're clearly not using the one from the dictionary...

Only a Yankee would need an explanation of what values are or why only caring about making money isn't one.

What a disgusting country filled with hollow degenerates. No wonder you keep voting for a senile grifting pedophile.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#483

Earlier quoted context omitted.

Maybe it would help if you shared your private personal definition of "value", since you're clearly not using the one from the dictionary...

Only a Yankee would need an explanation of what values are or why only caring about making money isn't one. What a disgusting country filled with hollow degenerates. No wonder you keep voting for a senile grifting pedophile.

Look, self-loathing is all the rage, but just because you don't understand what "values" mean doesn't mean you have to insult your entire country along with you.

As for making money, you're right, it's not one of your values. It takes a special case of main character syndrome to think it's not anyone's value.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#484

Earlier quoted context omitted.

No, it is not OK. But also, not what you quoted says. Not sure who you are arguing with really. There seems to be a few logical leaps in between each response. I also didn't say anything like that.

You asked, I answered. The direct and immediate effect of Amodei getting what he asks for in that essay will be to empower the Trump administration to approve model releases.

Well, that post certainly aged well.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#485

Earlier quoted context omitted.

Joking? Companies absolutely do get into arms races.

They do but they very very very don't have to. They can stop.

In reality there is basically no mechanism by which a large company (i.e. with a board) can do this so long as there is 1) significant profit motive and 2) no government intervention.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#486
post #473

Earlier quoted context omitted.

To quote notorious effective altruist Scott Alexander: > Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”. https://slatestarcodex.com/2017/04/07/yes-we-have-noticed-th...

To me it sounds a bit like this: "Look. I see that it doesn't work. I want it to work, so I will continue trying, even if it fundamentally cannot work. I am not interested in thinking about whether or not it can work. I am interested in showing to the world that I am well-intentioned and trying to do something , even if that something doesn't make sense".

I mean... Yes, any short snappy explanation is going to be easy to strawman by someone motivated to do so.

The longer non-snappy explanation is that "I will arbitrarily set numbers on things and call it impartial" obviously doesn't match EA's self-conception, that lots of EA cause areas are speculative and don't focus on numbers, that EAs that do focus on numbers do a lot of work to make sure the numbers aren't arbitrary, that EAs as a general rule don't claim to be impartial, and that awareness of Goodhart's law doesn't mean "never trying to objectively measure anything at all".

> I am interested in showing to the world that I am well-intentioned and trying to do something, even if that something doesn't make sense".

This is the kind of pre-conception that's essentially immune to reality. I hear the same thing about vegans (oh they say they care about animal suffering, but everybody knows about factory farms, they just want to feel superior to everybody else) or environmentalists (they say that climate change is a threat to humanity but really they just want to lecture us about our cars).

All I can say is that it doesn't match my experience, and that the effective altruists I've met spend quite a lot of time "thinking about whether or not it can work" and trying to learn from other people's mistakes.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#487
post #473

Earlier quoted context omitted.

To me it sounds a bit like this: "Look. I see that it doesn't work. I want it to work, so I will continue trying, even if it fundamentally cannot work. I am not interested in thinking about whether or not it can work. I am interested in showing to the world that I am well-intentioned and trying to do something , even if that something doesn't make sense".

I mean... Yes, any short snappy explanation is going to be easy to strawman by someone motivated to do so. The longer non-snappy explanation is that "I will arbitrarily set numbers on things and call it impartial" obviously doesn't match EA's self-conception, that lots of EA cause areas are speculative and don't focus on numbers, that EAs that do focus on numbers do a lot of work to make sure the numbers aren't arbit…

Those are fair points indeed. Let me try to elaborate on my opinion of EA:

> that awareness of Goodhart's law doesn't mean "never trying to objectively measure anything at all".

Goodhart's law doesn't say "never try to measure anything at all". It says "if you try to optimise for the metric, then your metric is doomed". What EA does is pretty much say "let's devise a metric and optimise for it". It does NOT say "let's measure something without influencing it at all". That is totally different.

Wikipedia says (happy to read your corrections if you think it is incorrect):

> Effective altruism (EA) is a [...] movement that advocates impartially calculating benefits and prioritizing causes to provide the greatest good. It is motivated by "using evidence and reason to figure out how to benefit others as much as possible, and taking action on that basis".

While I appreciate the idea of "trying to provide the greatest good" (difficult to go against that :-), my criticism is about the method.

* It is not very hard to convince oneself that if we stopped eating animals, then we would stop abusing chickens (did you know that tens of millions of chickens die during transport in trucks every year in England?) and emptying the oceans, and it would be objectively better in terms of animal suffering and for the biodiversity.

* It is not very hard to convince oneself that our CO2 emissions are literally going to get most of us killed, and that it would be globally better for us "humans who are currently alive" to do something about it. But there already, it's not entirely clear to me if the better outcome for life on Earth is to save the human species. Kind reminder that the human species is currently, measurably destroying all other species at a speed orders of magnitude faster than the extinction of the dinosaurs.

Effective altruism wants to do "the greatest good", but what is "good"? It may be good for a subset of humans to bomb another country and steal their oil, but obviously that would not be good for the subset of humans in the bombed country. It may be good for humans to find a clean magical energy, but that wouldn't change the current mass extinction for the other species (kind reminder that the current mass extinction has nothing to do with climate change, it is all about... well humans having easy access to energy and doing what humans do when they have cheap energy).

I feel like effective altruism says: "We can't define what the greatest good is, but we want to believe that anything is better than nothing. So we define a metric that we call 'impartial' (but that obviously isn't) and optimise for it, knowing that optimising for a metric defeats the purpose of that metric". Really it's rich people who want to do something good but don't want to bother getting informed and convincing themselves about what they want to do. "I'll give a ton of money and in return I get philanthropy points to share with my rich friends, but I don't want to have to think about what is being done with that money".

When someone invests a ton of money and energy into something they genuinely care about, they don't call themselves effective altruists, do they? They are just working for that cause. Effective altruism seems to be about rich people delegating the work of "doing something good" by donating some extra money, while they keep doing what made them rich in the first place (which almost always is something that is going against whatever I would consider the greatest good).

Re: Anthropic apologizes for invisible Claude Fable guardrails

#488

I moved off Claude Code 3 months ago. That decision keeps getting better and better as time goes on.

What model / runtime / harness and host have you settled on?

For now codex. Didn't manage to get others to work well. And fully aware that I'll have to move to another thing after OpenAI enshittifies this as well.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#489
post #487

Earlier quoted context omitted.

I mean... Yes, any short snappy explanation is going to be easy to strawman by someone motivated to do so. The longer non-snappy explanation is that "I will arbitrarily set numbers on things and call it impartial" obviously doesn't match EA's self-conception, that lots of EA cause areas are speculative and don't focus on numbers, that EAs that do focus on numbers do a lot of work to make sure the numbers aren't arbit…

Those are fair points indeed. Let me try to elaborate on my opinion of EA: > that awareness of Goodhart's law doesn't mean "never trying to objectively measure anything at all". Goodhart's law doesn't say "never try to measure anything at all". It says "if you try to optimise for the metric, then your metric is doomed". What EA does is pretty much say "let's devise a metric and optimise for it". It does NOT say "let'…

> Really it's rich people who want to do something good but don't want to bother getting informed and convincing themselves about what they want to do. "I'll give a ton of money and in return I get philanthropy points to share with my rich friends, but I don't want to have to think about what is being done with that money".

The people I have met at effective altruist conferences are not rich, though they lean upper-middle class. I've seen way more "enthusiastic broke student" types than millionaires.

> When someone invests a ton of money and energy into something they genuinely care about, they don't call themselves effective altruists, do they?

Well N=1, but I do.

(And also I've met tons of EA people who were not shy about investing all their energy in a cause they care about, even when all mainstream society tells them it's pointless.)

Post reply on HN