Earlier quoted context omitted.
> The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. Couldn't be further from the truth. The models can be tested and statistically evaluated. I ran a massive Fable max code review on my lone lisp codebase. Now that I have switched to OpenAI, I decided to run an equivalent review using Sol max and compare them. I'm keeping all data so I…
You sound very certain, but so do all the people who disagree with you, and they've got their own private benchmarks. You'll forgive me if I remain unconvinced.
I'm just saying it's not wise to simply put all these models in the same bucket and say any differences are due to vibes or hype. They are clearly different. We can and should scrutinize the testing methodology but it's not exactly fair to just ignore the results.
I don't intend for my benchmark to be private. The core component of my test is my parallel code review skill which is already on my GitHub. I'll be publishing the results on my website when it's done. Anyone could take the skill and reproduce the test using multiple models against any codebase out there, then analyse the depth of each model's findings.