Live data from Hacker News

"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

tryai.dev

31–40 of 110 posts

Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

#33
post #13
post #7

Grok! LOL! Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?

Grok 4.5 is...something else. It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing. On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written…

"You can't QoS that" sounds like the title of a nerdcore rap song.

Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

#40
post #13

Earlier quoted context omitted.

Grok 4.5 is...something else. It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing. On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written…

"You can't QoS that" sounds like the title of a nerdcore rap song.

Some of my poetry has clear IT jargon, but it's a very small portion of it (For some reason, OpenAI models, Gemini (and apparently Grok too), love to latch onto this and obsess over this idea that it's "programming poetry" or "poetry for the IT crowd". Often OpenAI and Gemini try to write the "equations of my poetry" (granted, I do write about a cyclical relationship between thinking, feeling and writing a lot, and I do have ONE poem which ends with a Q.E.D.).

I'm giving this context to say that it is very bizarre. It's as if they latch onto it and act as if it's a core or highly distinguished part of the poetry, when it really isn't. Anthropic models, on the other hand, absolutely do not do this, and have never done it.

I really don't understand why this happens. Maybe it's because it has a lot of portuguese, I don't know. And even though the "QoS" is clearly the wrong token being generated, I have had situations where gemini spoke of some phase of my poetry as the "Q&A part" (really, no joke...)

In any case, it's why it's my personal benchmark after all :D

Post reply on HN