Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
1–10 of 21 posts
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#2The models are nondeterministic, and therefore it's pretty normal for different runs to give different results.
I don't see this as evidence that Opus 4.6 has gotten worse.
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#3Because the website doesn't seem to show any sample size of runs, I assume they ran it once across the suite. The models are nondeterministic, and therefore it's pretty normal for different runs to give different results. I don't see this as evidence that Opus 4.6 has gotten worse.
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#4Because the website doesn't seem to show any sample size of runs, I assume they ran it once across the suite. The models are nondeterministic, and therefore it's pretty normal for different runs to give different results. I don't see this as evidence that Opus 4.6 has gotten worse.
are models really non deterministic?
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#5Because the website doesn't seem to show any sample size of runs, I assume they ran it once across the suite. The models are nondeterministic, and therefore it's pretty normal for different runs to give different results. I don't see this as evidence that Opus 4.6 has gotten worse.
And how is that an excuse?
I don't care about how good a model could be. I care about how good a model was on my run.
Consequently, my opinion on a model is going to be based around its worst performance, not its best.
As such, this qualifies as strong evidence that Opus 4.6 has gotten worse.
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#6Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#7Earlier quoted context omitted.
Yes. Look up LLM "temperature" - it's an internal parameter that tweaks how deterministic they behave.
The models are deterministic, the inference is not.
Even then, depending on the specific implementation, associativity of floating point could be an issue between batch sizes, between exactly how KV cache is implemented, etc.
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#8Because the website doesn't seem to show any sample size of runs, I assume they ran it once across the suite. The models are nondeterministic, and therefore it's pretty normal for different runs to give different results. I don't see this as evidence that Opus 4.6 has gotten worse.
are models really non deterministic?
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#9Earlier quoted context omitted.
The models are deterministic, the inference is not.
What does that even mean? Even then, depending on the specific implementation, associativity of floating point could be an issue between batch sizes, between exactly how KV cache is implemented, etc.
Re: Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%
#10Because the website doesn't seem to show any sample size of runs, I assume they ran it once across the suite. The models are nondeterministic, and therefore it's pretty normal for different runs to give different results. I don't see this as evidence that Opus 4.6 has gotten worse.