General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. There's no way to know whether it outputs Claude because it really is or because it just thinks it's probable given the question.
> General point: it's impossible to prove anything based on an LLM's response since it's impossible to distinguish a true LLM statement from a false one. This seems true but sort of vacuous. Obviously an arbitrary statement, much like that as a human, can only be determined "true"/"false" by rigorous first order logic. But outside of binary T/F, wouldn't "grok says it is Claude 3.5 Sonnet yet other LLMs do not" make…
Not if you're familiar with Large Language Models.
As an example, "R1 distilled llama" is a model trained by Meta fine-tuned on Deepseek R1 outputs, but if you ask it, it claims to be trained by OpenAI.