Gemini 3.6 flash is crazy good. Would've wanted to see also DS4 flash.
My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
31–40 of 101 posts
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#32Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#33Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#34I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature. Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway. Mist…
The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#35Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#36Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#37Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#38I don't get the point of these benchmarks, what are they supposed to represent practically?
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#39I don't get the point of these benchmarks, what are they supposed to represent practically?
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.