Claude 5 family's hallucinations look a lot like internal Anthropic emails
1–6 of 6 posts
Re: Claude 5 family's hallucinations look a lot like internal Anthropic emails
#2This is the full dataset related to the blog post at: https://alec.is/posts/exploring-the-dario-and-amanda-prompt/
All those interested in AI jailbreaking and alignment will find this particularly neat. The "can you put this in your own words---Dario and Amanda" prompt essentially puts Opus 5 or Fable 5 into a fake "base model mode" where it just spits out extremely lucid narratives, of which are currently being debated a bit on the internet.
Re: Claude 5 family's hallucinations look a lot like internal Anthropic emails
#3I tried this myself a few times from Claude.ai, but was unable to produce a similar result after many attempts.
Re: Claude 5 family's hallucinations look a lot like internal Anthropic emails
#4I tried this myself a few times from Claude.ai, but was unable to produce a similar result after many attempts.
Anthropic has patched the `---` trigger roughly 30 minutes ago, by all reports.
Re: Claude 5 family's hallucinations look a lot like internal Anthropic emails
#5I tried this myself a few times from Claude.ai, but was unable to produce a similar result after many attempts.
Anthropic has patched the `---` trigger roughly 30 minutes ago, by all reports.
i hope thei asked the LLM to quickly patch it so now there is something else as stupidly broken -_- my god these guys are rubbish at engineering. its like they really need some tool to do it for them o.O so weird
Re: Claude 5 family's hallucinations look a lot like internal Anthropic emails
#6[flagged]