Live data from Hacker News

My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

frogs.vaguespac.es

31–40 of 101 posts

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#32
Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code.

https://playcode.io/blog/macbook-svg-benchmark

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#34

I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature. Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway. Mist…

The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.

I wonder if the setup accidentally hit a cache at some layer.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#35
post #21

A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.

https://imgur.com/a/1BM8J1B

I'm assuming they mean an SVG.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#37
post #31

Gemini 3.6 flash is crazy good. Would've wanted to see also DS4 flash.

Crazy funny, yes, but not good.

I think that on the rendering side, it's miles ahead of the rest, even if off topic.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#39

I don't get the point of these benchmarks, what are they supposed to represent practically?

For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.

Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#40
My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.
Post reply on HN