Grok! LOL! Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?
"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
31–40 of 110 posts
Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#32The difference in cost is pretty incredible.
Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#33Grok! LOL! Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?
Grok 4.5 is...something else. It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing. On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written…
Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#34Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#35Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#36Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#37Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#38Now add Deepseek, GLM and Kimi :-)
.-""-.
/ \
| _ _ |
| (o)(o) |
\ /\ /
| -- |
| \/ |
| |
/ -- \
/ / \ \
/ / \ \
(__/ \__)
It might not have been the most scientific testRe: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#39The rectangular smudge tool is a weird tool in the first place, but it's cute to see the models try to use it.
Re: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
#40Earlier quoted context omitted.
Grok 4.5 is...something else. It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing. On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written…
"You can't QoS that" sounds like the title of a nerdcore rap song.
I'm giving this context to say that it is very bizarre. It's as if they latch onto it and act as if it's a core or highly distinguished part of the poetry, when it really isn't. Anthropic models, on the other hand, absolutely do not do this, and have never done it.
I really don't understand why this happens. Maybe it's because it has a lot of portuguese, I don't know. And even though the "QoS" is clearly the wrong token being generated, I have had situations where gemini spoke of some phase of my poetry as the "Q&A part" (really, no joke...)
In any case, it's why it's my personal benchmark after all :D