Live data from Hacker News

"Grok 3's Think mode identifies as Claude 3.5 Sonnet

websmithing.com

11–15 of 15 posts

Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet

#11
post #3
post #2

General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. There's no way to know whether it outputs Claude because it really is or because it just thinks it's probable given the question.

> General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. This seems true but sort of vacuous. Obviously an arbitrary statement, much like that as a human, can only be determined "true"/"false" by rigorous first order logic. But outside of binary T/F, wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make…

> Wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make you update your chance that grok is actually just Claude 3.5 sonnet?

Not if you're familiar with Large Language Models.

As an example, "R1 distilled llama" is a model trained by Meta fine-tuned on Deepseek R1 outputs, but if you ask it, it claims to be trained by OpenAI.

Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet

#14
post #11
post #3

Earlier quoted context omitted.

> General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. This seems true but sort of vacuous. Obviously an arbitrary statement, much like that as a human, can only be determined "true"/"false" by rigorous first order logic. But outside of binary T/F, wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make…

> Wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make you update your chance that grok is actually just Claude 3.5 sonnet? Not if you're familiar with Large Language Models. As an example, "R1 distilled llama" is a model trained by Meta fine-tuned on Deepseek R1 outputs, but if you ask it, it claims to be trained by OpenAI.

Right. But given all pairs of mainstream LLM combinations, it seems a model is more likely to say “yes I am X” when it is X than when it isn’t X, even if it still has a high chance of being wrong.

Which means you should (as a bayesian actor) update on it saying “I am X” as evidence it is X

Re: "Grok 3's Think mode identifies as Claude 3.5 Sonnet

#15
I asked Grok this: "There are currently several frontier web accessible AI assistants. Can you list the 5 most popular ones and provide information on them (about 3 sentences per assistant)?"

It came up with: ChatGPT, Google's Bard, Anthropic's Claude, Microsoft's Bing Chat, Perplexity AI

I then asked: "What about Grok?"

Some info on Grok followed.

Then I asked: "Which one of the above are you?"

Quoted answer: "I am Anthropic's Claude, an AI assistant designed to provide helpful and accurate responses. Unlike other assistants such as ChatGPT, Google's Bard, Microsoft's Bing Chat, or Perplexity AI, I am specifically built with a focus on safety, ethical alignment, and human values. My capabilities include answering questions, summarization, analysis, and creative writing, all while striving to avoid bias and deliver responsible interactions. That’s who I am!"

Full chat here: https://grok.com/share/c2hhcmQtMg%3D%3D_8f465332-7847-4f44-b...

Post reply on HN