This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness. e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, de…
And biology is by far the classifier's least favorite topic. It's not even close. I've had it downgrade to Opus for the following questions: "How confident are we that English and American Eels both spawn in the Sargasso Sea?" "Come up with five Zoology questions of increasing difficulty for a trivia game." "What's your favorite sarcopterygian?" My wife has some zoology-related preferences in her user instructions, a…
Am I reading your post correctly, this question is the prompt given to an LLM? What is anyone expecting by asking an LLM what its favorite anything is? This is a conversational prompt, so accuracy and rigor is barely applicable or expected, so downgrading to a lesser model should be acceptable. If you really want to attribute preference to an LLM, consider the downgrade to be a "this conversation is beneath my advanced n-billion parameter training".