Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

221–230 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#221

Earlier quoted context omitted.

What backward logic is this? PRC doesn't give a fuck about how US regulates AI companies. Pushing more regulation would ensure that Chinese companies catch up sooner. If you think otherwise you need to think harder.

The original topic was Anthropic's guardrails, which were meant in part to stop China from using Anthropic's models to bootstrap their own. I take it the logic of the comment was that pulling attention to Anthropic's stance on regulation is switching to the topic. But for what it's worth, I also think that people are way to quick to assume that strong regulations would only help China and thereby hurt safety. There a…

I don't subscribe to the belief that regulations in the US will lead to China advancing further.

But I also don't buy into the "China bad" narrative that gets frequently spread in online circles and in political circles. Its the cold war all over again, but this time its China instead of the Soviet Union.

Regardless of that, the regulations being proposed by Anthropic recently are not focused on the current issues which is my problem with all the hype marketing around hypothetical AGI/ASI. What is being proposed to be put in place will further cement the current frontier labs in their marketing leading position, and work to block new entrants, and open source competitors. That is the problem.

The other problem is none of them are talking about the real, difficult issues we are experiencing right now in the present. We don't need to talk about a sci-fi future scenario to recognize that LLMs have already caused and are causing harm in the real world. "We should probably regulate future frontier models" does nothing to help the current issues.

Wake me up when Anthropic says "The government should immediately stop us from hoovering up data and selling it back to the public. They should immediately stop us and others from enabling misinformation at scale that is already negatively effecting our democratic process. They should immediately stop us from building out new data centers until we have a large scale switch to renewables in the country, shore up the grids, or force us to generate our own power only with renewables" so on and so forth. Notice how any time the labs propose regulations, its only for a future hypothetical super intelligent model. Its never about their current operational liabilities.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#222
I'll defend Anthropic.

They are clear about the reasons for guardrails: prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs.

What is the critique against that? It seems pretty reasonable to me. You want AI-accelerated biological or radiological experiments running in your neighbors backyard? You want PRC-backed labs to continue to steal Anthropic's models via distillation?

Mitigating the harms of dual-use tech is notoriously difficult and fraught with trade offs. What I would want to see is cautious rollout and quick response, which is EXACTLY what they're doing.

Instead, this thread is full of bad-faith arguments about Anthropic being dishonest, making a "useless" model, or "the power is going to their heads." You can't read Anthropic's System Cards and come away with any of these impressions. Quite the opposite, in fact. They are honest to a fault, acknowledging problems they discovered even when it hurts them.

If your harmless request was downgraded to Opus, you're billed for Opus. They were 100% clear about that. I'd much rather have a Mythos-class model that falls back to Opus 10% of the time than be capped to Opus 100% of the time. If that doesn't work for you, then make a suggestion for something better!

If you are a white-hat security engineer hitting guardrails, I don't think you have standing to complain. I really don't. Their Glasswing program actually got banks and the industrial sector to take action to fix security vulnerabilities. Do you realize how special that is? A huge portion of the economy runs on vulnerable code and has for decades, despite security experts testifying to Congress, begging business leaders, pleading for intervention-- with no results. But suddenly they're all enrolled in a program that will find *and fix* vulnerabilities! White-hat security people should be rejoicing. Instead some of them are throwing rocks. Unbelievable. Shameful.

Meanwhile, society is screaming at the AI labs to be more conscientious about potential harms of AI. Legislatures are passing laws limiting data center construction. There are protests. And you, the HN community, the vanguard of our profession, have the temerity to demand "NO GUARDRAILS!" "HOW DARE YOU TRY TO PROTECT DEMOCRACY!" "MY SOFTWARE PROJECT IS MORE IMPORTANT THAN KEEPING NUKES AWAY FROM THE BAD GUYS!"

Go ahead HN, downvote me. It'd be an honor.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#223

Earlier quoted context omitted.

Well it's difficult to argue against something that was never specifically stated. If someone is able to state specifically how this is an arms race in any other way than that it's a race at all then I'm happy to have that conversation.

"Arms race" is the term used colloquially to describe the dynamic that emerges in "winner-take-all" markets. It seems that the frontier labs believe they're participants in a winner-take-all market. Therefore they're in "an arms race." Winner-take-all markets do not require that the winner literally destroys the losers, but only that the winner enjoys disproportionate returns compared to their actual superiority. Whe…

Creative destruction is absolutely a thing in the market, but the way things are going it seems more likely that open source models will just destroy everything else as far as most users are concerned. The big proprietary labs will be effectively left with Fable, GPT-Pro and Gemini Deep Research - stuff that by all indications needs very large scale compute to even feasibly run. We'll probably find out that each has its own strengths, weaknesses and viable niches, so there's no reason to expect any of those models to utterly destroy the others. They can all survive as specialty services.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#224

I know this isn't going to be a popular take, but here goes anyway... The complaints that Anthropic are routing your requests to a different model reminds me of an old Louis CK bit about airplane wifi. Clearly Anthropic was too aggressive with whatever guardrails they put in, but the response seems overly entitled to a model people didn't even know existed not that long ago. https://youtube.com/watch?v=me4BZBsHwZs

If you charge me for X, but under the hood you are delivering Y IT'S FRAUD!

The filter that downgrades you to opus sucks, but at least you know and you are charged accordingly.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#225

Earlier quoted context omitted.

I cannot overstate how much I think this take is wrong. Please please reconsider, look at the rate of progress being made, and consider that even if you only think ASI 'may' never happen in your lifetime it should still be one of your #1 concerns. Honestly, that respect for 'copyright protections' has somehow become a leftist shibboleth is bizarre to me and indicative that something has become deeply warped in our di…

> I cannot overstate how much I think this take is wrong. Please please reconsider, look at the rate of progress being made, and consider that even if you only think ASI 'may' never happen in your lifetime it should still be one of your #1 concerns. Frankly, this appeal comes across as the same kind of impassioned plea that a missionary might make when begging the faithless to repent and come to Christ before it's to…

Maybe if you really squint. I'm asking them to reconsider their views because the cumulative result of many opinions is policy. And yes, I'm making moral claims. So perhaps that makes it religious? I don't really think so, but I recognize that comparing things to religion is an effective dismissal tactic on here.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#226

Earlier quoted context omitted.

Nothing, they are just trying to scare monger the public and prime the pump for a massive bailout when it crashes out because apparently China are the big bad meanies.

You'd be fine if the PRC gets to ASI first? That's an interesting opinion.

Yes, why wouldn't I be? How is that worse than China getting it second?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#227

Earlier quoted context omitted.

I cannot overstate how much I think this take is wrong. Please please reconsider, look at the rate of progress being made, and consider that even if you only think ASI 'may' never happen in your lifetime it should still be one of your #1 concerns. Honestly, that respect for 'copyright protections' has somehow become a leftist shibboleth is bizarre to me and indicative that something has become deeply warped in our di…

There's nothing warped about it at all. Like it or not, it is a real issue. It's also an issue of license washing GPL code to privatize it. It's full scale theft of collective human knowledge, being sold back to us in a for profit private product. Outside of that though, there are other issues right now that need addressed before we speculate about what might be possible with ASI in the future. If the potential for a…

The "global stop order" is just generally perceived as an impossible coordination problem. So instead we see a mix of labs voluntarily putting in guardrails and regulatory efforts (which are not only aimed at hypothetical super-AIs of the future). Of course labs are also in a competitive race. And I actually think that it does make sense that the richest companies in the most dominant positions would in a better position to worry about safety than a startup that is just trying to survive at all. And just in general, it seems reasonable that the fewer companies have access to dangerous tech the better. This isn't really about some highly speculative future tech either -- current models already pose lots of risks, and the pace of model improvement is something wildly unprecedented. Whether or not you call it ASI, the capabilities we will have two years from now are hard to even imagine properly. Also, I don't think the issues that you are highlighting are all ones that Anthropic would dismiss as second-tier. In particular, mass unemployment from AI is how we will deal with a massive devaluation of human labor is one of the most serious concerns. And about other issues, reasonable people may differ. I'm more worried about biorisk than environmental damage, for example, but clearly we should be keeping an eye on both. Serious risks and problems, just because they aren't already harming people today, are not just a distraction.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#228

Earlier quoted context omitted.

Claude Opus 4.6 and 4.8 find vulns in source code just fine and 4.6 will pentest without source for you given a proper harness WITHOUT jailbreaking. WITH jailbreaks, you can probably imagine what they are capable of. Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.

public safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.

Distillation is not a thing unless you actually have the model weights. What people misleadingly call distillation is just training on chat logs, which has always been routine practice in the industry. There's a reason why every model today talks like early releases of ChatGPT.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#229
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

First time? They've always been misanthropic, ironically. They seem to hate their users and think that their AI is so dangerous it'll destroy the world and not to be trusted, I mean Anthropic was literally started because people at OpenAI thought the latter was too forgiving on "safety."

Re: Anthropic apologizes for invisible Claude Fable guardrails

#230
post #95

Earlier quoted context omitted.

Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll co…

[flagged]

The flawed premise is thinking that AGI is a real risk, and that they care about it more than making money, that is why HN does think it's simply regulatory capture.
Post reply on HN