Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

471–480 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#471
post #320

Earlier quoted context omitted.

Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.

> the purpose of Excel is pretty clear cut and the scope is small. That has to be the understatement of the century.

I don’t think excel can give me the instructions for building a house or how to cook a particular meal or write my emails for me. The potential output of LLMs is quite obviously more broad than excel.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#472
post #329

Earlier quoted context omitted.

They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…

the output is definitely better. and i find it crazy how every time a new model comes out people trip over themselves to say how much worse it is than previous models, when in fact that is basically an impossibility. like, they've got the numbers, man-- you only release a new model when the numbers get gooder. the burden of proof is on the "didn't get better" side, not the "prove that it's better" side, because the a…

You're wrong in lots of ways.

Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...

This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/

The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150

Re: Anthropic apologizes for invisible Claude Fable guardrails

#473
post #415

Earlier quoted context omitted.

I may be naive, but I have the feeling that "I will arbitrarily set numbers on things and call it impartial" is... weird at best. I understand how one may wonder if there was a way to do that, but it feels insane to me that one would actually conclude that "yes, it is possible". We have examples everywhere showing that it is generally impossible to define a metric that correctly represents the underlying concept we w…

To quote notorious effective altruist Scott Alexander: > Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”. https://slatestarcodex.com/2017/04/07/yes-we-have-noticed-th...

To me it sounds a bit like this:

"Look. I see that it doesn't work. I want it to work, so I will continue trying, even if it fundamentally cannot work. I am not interested in thinking about whether or not it can work. I am interested in showing to the world that I am well-intentioned and trying to do something, even if that something doesn't make sense".

Re: Anthropic apologizes for invisible Claude Fable guardrails

#474

Earlier quoted context omitted.

Sometimes governments have to deal with the weapons made by their enemies and that gets them stuck in an arms race. Companies don't have to do that. If they're getting into actually dangerous territory, they can stop as soon as they want to.

Joking? Companies absolutely do get into arms races.

They do but they very very very don't have to. They can stop.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#475

Earlier quoted context omitted.

Nonsense. Everyone has values. "Make myself maximum money" is a value. "Amass maximum power over the world's information" is a value. It's clear Amodei certainly follows the latter, and I would soften the former somewhat for him; they did after all decline the Pentagon contract that would have made money but would have meant giving up some control of information.

Those aren't values. Maybe goals or motivations but not values in any conceivable way, shape or form. This site is full of pod people I swear.

Maybe it would help if you shared your private personal definition of "value", since you're clearly not using the one from the dictionary...

Re: Anthropic apologizes for invisible Claude Fable guardrails

#476
post #52

Earlier quoted context omitted.

It was in the announcement, too. I’m 99% sure they edited it after they changed their mind, because I knew about it from reading that, and never opened the model card.

On the earliest web archive snapshot I can find [0], I do not see any mention of the safeguard/sabotage under discussion [1]. And to be clear, this isn't the safeguard where the model is explicitly downgraded to Opus, but rather where the Fable/Mythos model's "effectiveness" is transparently "limited" via "prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)". [0]: https://web.archive.org/…

Indeed. I can’t find it, but I swear I knew after reading the announcement, without opening the model card. It stood out because it was so weird. Maybe they changed it before the first snapshot, maybe something else happened, I don’t know.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#477

Earlier quoted context omitted.

I’ve noticed that too many HN folks seem to think that cynicism makes them more intelligent. I think it must be some kind of insecurity, about not wanting to be seen as naive or something. It’s pretty sad though, I wonder how some of these people find any peace or joy in their lives.

Believing they have any interests other than theirs at heart is like believing the stripper is really in love with you. That's not cynicism, that's just common sense.

Do you feel really smart now?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#478

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…

mixing up bioweapons, malware, with hate speech (which is basically a censorship) shows how very basic people like Trump can win. Hopefully you won't wait to be censored before realizing that anything could be interpreted as "hate speech".

Re: Anthropic apologizes for invisible Claude Fable guardrails

#479

I'll defend Anthropic. They are clear about the reasons for guardrails: prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs. What is the critique against that? It seems pretty reasonable to me. You want AI-accelerated biological or radiological experiments running in your neighbors backyard? You want PRC-backed labs to continue to steal…

Having a chatbot that talks to you about synthetic biology or nuclear physics is just not the same as being equipped to develop biological weapons or atomic bombs.

None of this will happen in the "neighbors backyard." You are exaggerating the threats to "democracy" while simultaneously invoking democracy to limit freedom of information. The suggestion that somehow the bad guys will get nukes if we let people access information is just absurd.

Society at large is not concerned about whether someone asks the chatbot about organic chemistry. They are concerned that they will be de-facto forced to interact with some shitty automated system to get by in life, like having to pass an AI-powered ATS to get a job.

They are tired of the hype and tired of idiots like Amodei being elevated to heights of power and influence. They are concerned that the things they love are being devalued. But they don't give a fuck if I ask an AI about genetically modifying viruses. This is a pet issue among some of the AI safety crowd.

So, yes, I am 100% fine with PRC-backed labs distilling Anthropic's models. I do not care about Anthropic. They have demonstrated that they are not on my side, and that they are at best ambivalent about actually empowering their users. I'm not a fan of the PRC either, but their distance makes them far less of a threat to me than companies like Anthropic and my own government.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#480

Earlier quoted context omitted.

I agree 100%. Doing a worse job IS an error. It should be treated as such. Or at the very least make that behavior opt-in. The default should not be pretending like nothing happened and just quietly doing a worse job. Imagine your healthcare provider just sometimes decided not to read your test results very carefully and you risked death? Now realize that healthcare providers use Claude now and that scenario wasn't h…

Yes, but as with spam/phishing/abuse prevention, too much information about what does and doesn't trigger things can be very useful to attackers. An explicit error is something you can feed into another AI to find jailbreaks. I think it's a fundamentally impossible thing to fix, though. There's no 100% correct answer.

I understand completely, and respect the tough spot they're in. They have a choice between human safety and cybersecurity/business needs here. I don't envy that position.

That said, this thing is in real production use with war fighters, doctors, and financial experts. Just YOLO'ing to a dumber model midway through a multi-step process and pretending everything is fine is not a real or defensible option. Someone is going to die, and its going to be the fault of whoever decided to make this the default rather than opt-in.

Personally, I couldn't live with myself.

Post reply on HN