Earlier quoted context omitted.
Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.
> the purpose of Excel is pretty clear cut and the scope is small. That has to be the understatement of the century.
Anthropic apologizes for invisible Claude Fable guardrails
471–480 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#472Earlier quoted context omitted.
They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…
the output is definitely better. and i find it crazy how every time a new model comes out people trip over themselves to say how much worse it is than previous models, when in fact that is basically an impossibility. like, they've got the numbers, man-- you only release a new model when the numbers get gooder. the burden of proof is on the "didn't get better" side, not the "prove that it's better" side, because the a…
Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/
The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150
Re: Anthropic apologizes for invisible Claude Fable guardrails
#473Earlier quoted context omitted.
I may be naive, but I have the feeling that "I will arbitrarily set numbers on things and call it impartial" is... weird at best. I understand how one may wonder if there was a way to do that, but it feels insane to me that one would actually conclude that "yes, it is possible". We have examples everywhere showing that it is generally impossible to define a metric that correctly represents the underlying concept we w…
To quote notorious effective altruist Scott Alexander: > Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”. https://slatestarcodex.com/2017/04/07/yes-we-have-noticed-th...
"Look. I see that it doesn't work. I want it to work, so I will continue trying, even if it fundamentally cannot work. I am not interested in thinking about whether or not it can work. I am interested in showing to the world that I am well-intentioned and trying to do something, even if that something doesn't make sense".
Re: Anthropic apologizes for invisible Claude Fable guardrails
#474Earlier quoted context omitted.
Sometimes governments have to deal with the weapons made by their enemies and that gets them stuck in an arms race. Companies don't have to do that. If they're getting into actually dangerous territory, they can stop as soon as they want to.
Joking? Companies absolutely do get into arms races.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#475Earlier quoted context omitted.
Nonsense. Everyone has values. "Make myself maximum money" is a value. "Amass maximum power over the world's information" is a value. It's clear Amodei certainly follows the latter, and I would soften the former somewhat for him; they did after all decline the Pentagon contract that would have made money but would have meant giving up some control of information.
Those aren't values. Maybe goals or motivations but not values in any conceivable way, shape or form. This site is full of pod people I swear.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#476Earlier quoted context omitted.
It was in the announcement, too. I’m 99% sure they edited it after they changed their mind, because I knew about it from reading that, and never opened the model card.
On the earliest web archive snapshot I can find [0], I do not see any mention of the safeguard/sabotage under discussion [1]. And to be clear, this isn't the safeguard where the model is explicitly downgraded to Opus, but rather where the Fable/Mythos model's "effectiveness" is transparently "limited" via "prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)". [0]: https://web.archive.org/…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#477Earlier quoted context omitted.
I’ve noticed that too many HN folks seem to think that cynicism makes them more intelligent. I think it must be some kind of insecurity, about not wanting to be seen as naive or something. It’s pretty sad though, I wonder how some of these people find any peace or joy in their lives.
Believing they have any interests other than theirs at heart is like believing the stripper is really in love with you. That's not cynicism, that's just common sense.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#478I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#479I'll defend Anthropic. They are clear about the reasons for guardrails: prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs. What is the critique against that? It seems pretty reasonable to me. You want AI-accelerated biological or radiological experiments running in your neighbors backyard? You want PRC-backed labs to continue to steal…
None of this will happen in the "neighbors backyard." You are exaggerating the threats to "democracy" while simultaneously invoking democracy to limit freedom of information. The suggestion that somehow the bad guys will get nukes if we let people access information is just absurd.
Society at large is not concerned about whether someone asks the chatbot about organic chemistry. They are concerned that they will be de-facto forced to interact with some shitty automated system to get by in life, like having to pass an AI-powered ATS to get a job.
They are tired of the hype and tired of idiots like Amodei being elevated to heights of power and influence. They are concerned that the things they love are being devalued. But they don't give a fuck if I ask an AI about genetically modifying viruses. This is a pet issue among some of the AI safety crowd.
So, yes, I am 100% fine with PRC-backed labs distilling Anthropic's models. I do not care about Anthropic. They have demonstrated that they are not on my side, and that they are at best ambivalent about actually empowering their users. I'm not a fan of the PRC either, but their distance makes them far less of a threat to me than companies like Anthropic and my own government.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#480Earlier quoted context omitted.
I agree 100%. Doing a worse job IS an error. It should be treated as such. Or at the very least make that behavior opt-in. The default should not be pretending like nothing happened and just quietly doing a worse job. Imagine your healthcare provider just sometimes decided not to read your test results very carefully and you risked death? Now realize that healthcare providers use Claude now and that scenario wasn't h…
Yes, but as with spam/phishing/abuse prevention, too much information about what does and doesn't trigger things can be very useful to attackers. An explicit error is something you can feed into another AI to find jailbreaks. I think it's a fundamentally impossible thing to fix, though. There's no 100% correct answer.
That said, this thing is in real production use with war fighters, doctors, and financial experts. Just YOLO'ing to a dumber model midway through a multi-step process and pretending everything is fine is not a real or defensible option. Someone is going to die, and its going to be the fault of whoever decided to make this the default rather than opt-in.
Personally, I couldn't live with myself.