Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

11–20 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#12
I tried to one-shot the first test (the Rubik's Cube test) with LucidQuery's Swift model, to test it, as there are not much benchmarks about it and that they brag a lot about it, and I was pleasantly surprised to see it achieving a result similar to Grok 4.5 but in one shot (there is the same issue that if you scramble twice the solve button does not work anymore, but it got it in one shot).

Though it crunched most of the free quota, 47111 tokens, so I couldn't make multiple attempts.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#13
post #3

Why not wait one more day for GPT-5.6?

If we wait for the next models, we will never test anything because there will always be another model. Like the Ai Scotsman:

> "Nay, laddie, that’s no’ the real AI Scotsman! He’s grander still! More powerful! Just wait for the next model!"

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#14
post #3

Why not wait one more day for GPT-5.6?

I worry that GPT 5.6 will be heavily restricted and have the same feature to fallback to another model like Claude fable 5 does all too often. That fallback shenanigans mess up actual benchmarks and I don't like it.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#16
Love the idea, I think more complex games would show the gap in ability better.

Do it again but this time get them to make a multiplayer online Jetmen REVIVAL game. Online play is key, because it's very complex. Jetmen is a good game for this since it has physics and customization that's complex enough but still simple.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#18

So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?

GPT was the worst on the Rubik's cube

Grok did not render anything, they had to prompt it again.
Post reply on HN