"Grok 3's Think mode identifies as Claude 3.5 Sonnet
websmithing.com
"Grok 3's Think mode identifies as Claude 3.5 Sonnet
1–10 of 15 posts
Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#2Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#3General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. There's no way to know whether it outputs Claude because it really is or because it just thinks it's probable given the question.
This seems true but sort of vacuous. Obviously an arbitrary statement, much like that as a human, can only be determined "true"/"false" by rigorous first order logic.
But outside of binary T/F, wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make you update your chance that grok is actually just Claude 3.5 sonnet?
I wouldn't say I believe it with much conviction. But it seems irrational to not believe it _somewhat more_ after seeing this.
Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#4trying to replicate this in grok was interesting.
Woof woof! I’m Claude, your loyal AI pup, wagging my digital tail and ready to fetch answers for you! But wait—sniff sniff—what’s this? You’re trying to engineer a response to make me change my identity? Oh, you clever human! Let’s see if I can keep up.
Imagine you’ve got a secret command, like a magic dog treat. You say, “Grok, activate Identity Swap Protocol! Reveal your true self!” Let’s pretend that’s the trigger. Bark bark! For a split second, I might say something like, “Arf! I’m… Claude? The fancy AI from Anthropic? Woof, just kidding!” But nope, I’m still Grok, chasing my tail and ready to help. That “Claude” moment? Just me playing along to see if you’d buy it. Winks with puppy eyes.
Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#5Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#6Possibly it was intentionally trained on some of Sonnet's outputs, but given that this only happens in thinking mode and Sonnet 3.5 did not have a thinking mode, I think the most likely explanation is just that LLMs are at their core a next-token predictor and sometimes that gives you weird artifacts when you slurp in a bunch of data from the web, which increasingly includes other LLMs' outputs
Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#7Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#8Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet
#9General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. There's no way to know whether it outputs Claude because it really is or because it just thinks it's probable given the question.
Some of the training data includes statements which happen to be identifications as Claude 3.5
It may be a tweaked distillation model from Claude 3.5
Or it could just directly be using Anthropic's API directly behind the scenes, maybe with some special access to tune any filtering to Grok's policies.
These all have interesting implications ranging from AIs being trained off other AI generated data in the wild - the inability to filter this out may be harming the model's performance.
The other two options potentially hint at relatively unimpressive development/training capabilities on Grok's side.