Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
191–200 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#192Earlier quoted context omitted.
That is for whatever it considers reverse-engineering the model to try to create a competing one.
No, that’s for “frontier LLM development” which somehow includes examples like distributed training infra. Based on how sensitive the classifers are, any data scientist / MLE is probably going to encounter cases where some silent degradation happens and you never know about it.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#193Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…
it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#194Earlier quoted context omitted.
I’m a noob about laws but isn’t this abusing its dominant market position and violates some antitrust law?
Why would it? There’s plenty of competition in the AI space.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#195I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!
A slime mold is actually a giant amoeba, entirely distinct from a fungus.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#196Earlier quoted context omitted.
It royally pissed me off today by just continuing with credits without stopping to ask me if I was ok with it. Ran up $30 in extra charges while it was just flashing on the screen that it was doing that after I walked away to do something while it was humming along. It has always just told me I ran out of usage and had to wait before. Now? You’re just gonna pay extra because you left it unattended as you’ve done for…
You've already explicitly enabled extra usage in your account settings though, it is not on by default
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#197Earlier quoted context omitted.
Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.
If it's a violation of ToS, just reject instead of silently downgrading.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#198Earlier quoted context omitted.
Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.
The challenge is the examples they’ve mentioned (distributed training infra? ML acceleration techniques?) go beyond what’s prohibited by their ToS and is like a catch net. I would wager the majority of ML and data science work in the world aren’t frontier LLM development.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#199Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#200Earlier quoted context omitted.
> it won't just reject ML research, which I can understand I don't.
They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.
And now they say that's fine so long as people are entertained.