Updating the GenAI comparison website is starting to feel a bit Sisyphean with all the new models coming out lately, but the results are in for the Flux 2 Pro Editing model! https://genai-showdown.specr.net/image-editing It scored slightly higher than BFL's Kontext model, coming in around the middle of the pack at 6 / 12 points. I’ll also be introducing an additional numerical metric soon, so we can add more nuance t…
Hey I hope you see this. The scoring needs to be a 0-10 or something with a range rather than pass or fail. Flux one getting the same score for the surfer as Gemini pro 3 reduces the quality of the benchmark.
I don't know if I'm going to get as granular as 1-10 only because the finer the scoring - the more potential for subjectivity. That's why it was initially set up as a "Minimum Passing Criteria Rule Set" along with a Pass/Fail grade.
A suggestion from a previous HN post was something along the lines of (0 Fail, 0.5 Technical Pass, 1.0 Proficient Pass).