Earlier quoted context omitted.
> This supposedly is better than KimiK2.7 How can you tell? I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks. Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?
Most of the random comments you read on HN and reddit about how nice/bad various LLMs are, is basically based on the commentator's "vibe" about it, and almost nothing is grounded in evidence or actual usage. Don't read too much into it, want to know how good a model is? Run it with your own non-public benchmark, basically the only way to get proper answers you can somewhat rely on, everything else is manipulated, mis…
Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.