Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

61–70 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#64

How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?

Cursor

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#65

So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?

GPT was the worst on the Rubik's cube

GPT-5.5 isn't really a fair comparison to Fable or Opus. GPT-5.5-Pro would be a better comparison (and yes, I know how much more expensive it is).

I really wish they'd thrown in something like GLM-5.2 into the comparison.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#66

This is disgustingly biased. The conclusion is that Grok holds its own?! There was zero evidence of that.

Yes, I wonder how the verdicts would hold under a blinded test. This analysis read like the authors going out of their way to be supportive of Grok.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#68
post #4

I am 99% sure the post was written by AI

The honest takeaway: this is 100% written by an LLM.

Pretty much any harness with a halfway decent model should be capable of running the experiment including making the screenshots, summarising the results, and so forth.

I sort of do this whenever an interesting new model comes out.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#69

[flagged]

> This feels like a kid trying to do science. The will is there, but lacks experience.

It's funny, when I saw the title I was hoping the article would include some sort of blind ranking, where you could see the outputs (without knowing which model they came from) and score them on some criteria. Could have been a fun way to get a better ranking of the results.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#70
post #49

If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

Impressive, specially the amount of models used for the comparison.
Post reply on HN