I don't get the point of these benchmarks, what are they supposed to represent practically?
The ability of an LLM to produce something not in its training data set.
My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
41–50 of 101 posts
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#42My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#43Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#44I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#45Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#46I don't get the point of these benchmarks, what are they supposed to represent practically?
For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI. Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outper…
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#47Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#48Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#49Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.
Re: My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
#50How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.