Live data from Hacker News

My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

frogs.vaguespac.es

71–80 of 101 posts

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#71
post #69

Fable 5 on Max knocks it out of the park: https://imgur.com/a/usR8K7G Definitely has some creative flourishes. (I made no extra prompting. Just the above text. Single shot.)

rex paludis is a nice touch.

It's even Carolus rex paludis! Not sure if fable is a history buff or a Sabaton fan, but I repeat myself.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#72
post #53

My personal benchmark is a directory containing a bunch of research papers on the physics of popping popcorn kernels, and a prompt about creating high fidelity, photo realistic 3D models of all the different kinds of popped kernels. Fable (surprisingly? unsurprisingly?) refused to do it last time I tried, and the results from other models are, well, fine , but there's still plenty of headroom on this particular one.

LLMs don’t work particularly well in 3d applications in my experience. Every new model release I’ll ask one for help with my path tracer and results are horrid.

I had great success with Claude editing and cleaning up STL files that I created by 3D scanning objects using my cellphone. The scans were very messy and the AI made the clean-up really easy.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#74

Fable 5 on Max knocks it out of the park: https://imgur.com/a/usR8K7G Definitely has some creative flourishes. (I made no extra prompting. Just the above text. Single shot.)

It makes the frog a king which is not in the prompt, so there is a bit of confusion.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#76
post #55

I don't get the point of these benchmarks, what are they supposed to represent practically?

The vendors like to toss around terms like “thinking” or “reasoning” to encourage prospective buyers to anthropomorphize models. These challenges are a visceral reminder that none of those marketing claims accurately describe was LLMs do: they’ll happily return things even the worst human illustrator would never hand in and make errors showing that there’s no model of the world behind anything they do. That doesn’t m…

> That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.

The standard retort seems to be "this is also true of a large proportion of humans". But I think it's clear that there are differing patterns in how humans vs. models err on various tasks.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#77

I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature. Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway. Mist…

> Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

You seem to imply that they ought not to. I disagree.

I wasn't familiar with the term before this post. Having learned it, were I given the task, I think I'd be strongly tempted to do the same extrapolation.

> If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

Agency is agency. You still need to vet what the model's output is actually permitted to control.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#78

Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code. https://playcode.io/blog/macbook-svg-benchmark

Great idea! Would be interesting to see the raw SVG sources too, not only the renderings.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#79

Fable 5 on Max knocks it out of the park: https://imgur.com/a/usR8K7G Definitely has some creative flourishes. (I made no extra prompting. Just the above text. Single shot.)

It makes the frog a king which is not in the prompt, so there is a bit of confusion.

I'm afraid you'll need to look up "Habsburg" yourself.

Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

#80

Earlier quoted context omitted.

It makes the frog a king which is not in the prompt, so there is a bit of confusion.

I'm afraid you'll need to look up "Habsburg" yourself.

Nathan is right though. It specifies a Habsburg jaw, not that the frog is a Habsburg king/prince. It's understandable that the LLM will hallucinate a king from this, but it isn't what's being asked of it.
Post reply on HN