Untitled topic
1–8 of 8 posts
Re: undefined
#2They definitely must be doing some quantization or optimization to meet demand, otherwise why would model performance degrade this much? It's been crazy for me personally
Re: undefined
#3Re: undefined
#4Re: undefined
#5Re: undefined
#6These seem to be different tests? One has 6 tasks the other has 30.
Re: undefined
#7These seem to be different tests? One has 6 tasks the other has 30.
Re: undefined
#8See original opus 4.6 sitting at 16% hallucination and the retest on 12th of april at 33% They definitely must be doing some quantization or optimization to meet demand, otherwise why would model performance degrade this much? It's been crazy for me personally
Combining multiple tests on the same leaderboard like this is nonsense, there should be a separate leaderbaord for the new tasks where every model is tested again.
Putting it on the original leaderboard as "Opus 4.6 (April 12)" is so obviously inappropriate that it smells like deception. You could say that the leaderboard is hallucinated.