Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

281–290 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#281
post #95

Earlier quoted context omitted.

Why is this surprising or a problem?! It's a model demo, & their reasoning is reasonable and fair. Why all this drama.

Because most people in tech never took a philosophy course or an ethics course and think that tech is obviously a good for the world and that there are no downsides to advancing tech. So any efforts that try to apply ethics to it are overreaching, ignorant, and futile in the face of the good that is tech!

Or alternatively, it is plain and obvious that Anthropic is using ethics to justify business decisions.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#282
post #212

Earlier quoted context omitted.

So as not to be vague, and since I just pushed a version I'm starting to be vaguely happy with... https://tylereaves.github.io/uk-rail-map/ This is the result of probably a few hundred round trips. The really interesting part of the problem is keeping it both relatively true to real geometry, while greatly exaggerating it horizontally so you can actually see the individual running lines/sidings, like a signaling sche…

Fascinating. Can you explain why southern London is DC while northern London is AC?

Prior to 1948 when they were all nationalized into British Rail, there were various railroad companies operating across the country. One of these was the Southern Railway, which, well, operated in the South. They started electrifying very aggressively in the mid 1920's. At the time most of what little electrification there was was in London on the Underground.

Compared to AC, 3rd Rail DC is cheaper to install, especially as a retrofit (Overhead wires require bigger tunnels, and increased spacing around tracks for the masts). Downside is that it's not really great for speeds above about 60-70mph, as well as being a bit of a pedestrian hazard. (Ever the one about not peeing on the rails so you don't get shocked? That's 3rd rail DC.)

For the Southern, with it's mostly short routes with many stops, electricfiation was a pretty obvious win, and doing 3rd rail made sense because they could do it quickly and cheaply.

In contrast, the northern routes were electrified muuuch later, after steam had gone away. The main East Coast Mainline from London up to Newscastle and on to Edinburgh wasn't fully electrified until 1991. By the '60s and '70s, with train speeds increasing to 80mph and up, overhead AC was the clear winner.

If you look closely there are a few exceptions - the Merseyrail network in Liverpool is DC. Built 1970s, but using some existing underwater tunnels, and slow speed commuter. Then running ESE from London you have the high speed AC lines leading to the Channel tunnel. Well spotted, the trend generally is quite distinct.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#283

The guardrails are pretty tight. It is even refusing to decode morse code: https://x.com/Schappi/status/2064839631137546503?s=20 The prompt was: please translate .. ..-. / -.-- --- ..- / -.-. .- -. / .-. . .- -.. / - .... .. ... --..-- / - --- ..- -.-. .... / --. .-. .- ... ...

Yeah, this shouldn't have been released yet.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#284
News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o...

> “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.”

Sounds like the widespread condemnation worked.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#285

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

Or if GPU companies detected you were trying to train a model and injected intentional numerical errors.

Nvidia already did something similar with Lite Hash Rate (LHR), limiting performance on purpose just when running mining apps...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#286
post #185

The question is: If biological, computer security, and ML research are so bad, why do they even train on the relevant data? The only answer that makes sense is they wanted the model to be competent and usable in these fields, just not by you , which is why they had to bolt on a badly functioning crippling device after the fact.

Or they wanted the model to be good at these things, for the companies that legitimately need access to these capabilities.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#287

Earlier quoted context omitted.

Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?

If a for profit company does a thing that could be motivated by profit or altruism, which of those 2 motivations do you think is most likely?

When they've repeatedly made decisions against their for profit nature, it changes the calculus a bit.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#288

Let's all vote with our wallets and collectively boycott misAnthropic or at least their feeble fable safety theater. Whining on social media only goes so far, especially when they're concealing their anticompetitive strategies under the veil of safety.

Agreed. I've already cancelled my subs, and everyone else needs to do the same, including boycott it for their companies, otherwise nothing will ever change. You can't reason with psychopaths. The only recourse is to hit them where it hurts - their wallet. Still though, the world would be a better place if open-source crushes Anthropic and they fade away into obscurity until the end of time. We don't need or want companies and people like this at the helm of humanities progress.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#289
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

The "tradeoff" warning implies they stand by their thinking and don't think there was anything qualitatively wrong with it which, if nothing else, is helpful so potential customers can know how they think. I think the core lesson is if you want reliable infrastructure to build into an application you should use a different provider. (edit: I'm not specifically an Anthropic hater, but having just spent some time adding complexity to an app to deal with the existing refusal behavior in Sonnet... I understand why they might want this in an end user chatbot but for an API it's really not acceptable)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#290

Earlier quoted context omitted.

Yeah but a lot of the guardrails are pretty obviously to prevent competition not for safety.

Hmm. Maybe they are concerned about state actors trying to train equivalent models without the safeguards?

Why invent new motives for Anthropic when their real motives are plain and obvious and have been confirmed time and time again by their behavior over the last few years? Their concern is their own power and wealth. Every other conceivable motive is secondary to that.
Post reply on HN