We made Grok 4.5, GPT-5.5, and Claude build the same apps
11–20 of 97 posts
So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#12I tried to one-shot the first test (the Rubik's Cube test) with LucidQuery's Swift model, to test it, as there are not much benchmarks about it and that they brag a lot about it, and I was pleasantly surprised to see it achieving a result similar to Grok 4.5 but in one shot (there is the same issue that if you scramble twice the solve button does not work anymore, but it got it in one shot).
Though it crunched most of the free quota, 47111 tokens, so I couldn't make multiple attempts.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#13Why not wait one more day for GPT-5.6?
If we wait for the next models, we will never test anything because there will always be another model. Like the Ai Scotsman:
> "Nay, laddie, that’s no’ the real AI Scotsman! He’s grander still! More powerful! Just wait for the next model!"
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#14Why not wait one more day for GPT-5.6?
I worry that GPT 5.6 will be heavily restricted and have the same feature to fallback to another model like Claude fable 5 does all too often. That fallback shenanigans mess up actual benchmarks and I don't like it.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#15So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
GPT was the worst on the Rubik's cube
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#16Love the idea, I think more complex games would show the gap in ability better.
Do it again but this time get them to make a multiplayer online Jetmen REVIVAL game. Online play is key, because it's very complex. Jetmen is a good game for this since it has physics and customization that's complex enough but still simple.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#17Too nice to Grok, if there are really cost savings it should say how much each of the three demos cost so we can judge if it's worth the lower quality (probably not). The time to complete each would also be interesting.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#18Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#19Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#20I'd like to see the comparisons with DeepSeek, Qwen, Mimo, Kimi and GLM