Earlier quoted context omitted.
Corporate America never backs down. It simply rallies and tries again later until people are too fatigued to care. The only solution is to abandon ship, which I am doing. MS walked back in OS ads the first few times, but ultimately we still ended up on the exact trajectory everyone was outraged at. OpenAI still ended up on its path to closed AI despite initial walk backs. The story repeats itself over and over again,…
Same with VISA/Mastercard deciding what we can/cannot buy. The only solution is to stop using their credit cards at all.
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
441–450 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#442Earlier quoted context omitted.
The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?
Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#443Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#444Wait a few months and a competitor will release a similarly powerful model with less guardrails, if they steal sufficient market share Anthropic will reverse policies. This is why I’m immensely hoping the Chinese don’t stop with their open sourced local models. None of these companies are your friend.
The Chinese aren't your friend either [1].
[1] https://www.hks.harvard.edu/centers/carr-ryan/our-work/carr-...
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#445Earlier quoted context omitted.
Yes, that's true. Excluding Fable, OAI models are the most refusal heavy. However, I'd rather get a refusal than response with poisoned output. Since currently there's no way to verify if poisoning happened or not, I don't trust Anthropic anymore, regardless of what they say. But my trust towards OAI is also brittle - what if they also do it, or start doing it? I want to have a verifiable way to know that the prompt…
What kind of work are you getting refusals on? Genuinely curious. The only refusal I’ve had in recent memory was declining to find doorbell camera footage matching a certain description, which is fair enough and I think EU laws heavily restrict such activities (even tho I’m not in the EU)
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#446Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#447Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#448Earlier quoted context omitted.
Wow… just wow. The future looks incredibly bleak if people are throwing fisftuls of money at this company. Anthropic will quickly become the sole arbiter of everything in your life.
Why do people think this is the future? Anthropic has the leading model, and so they're able to hold back functionality. They do so with obvious regards to safety. If anything a future with models of such capabilities and no safeguards would be a bleak future. But its likely what were headed in once other companies catch up.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#449These AI places have 0 clue about how threat actors actually work. None of their mitigations or guard-rails is effective, and now they are even turned against them.
Additionally, if they don't all implement the same level of effective guard-rails, there will always be some model you can abuse to do the work anyway, and hence there is 0 effect on threat actors, they will just run some local model that does 5% less quality, which does not matter to them 1 bit.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#450Earlier quoted context omitted.
They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.
Ironic, given they piggybacked on the entirety of human knowledge and massive amounts of GPL'd software and repeatedly say they want to replace people with a tool. And now they say that's fine so long as people are entertained.