We made Grok 4.5, GPT-5.5, and Claude build the same apps
61–70 of 97 posts
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#62Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#63Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#64How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#65So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
GPT was the worst on the Rubik's cube
I really wish they'd thrown in something like GLM-5.2 into the comparison.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#66This is disgustingly biased. The conclusion is that Grok holds its own?! There was zero evidence of that.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#67Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#68I am 99% sure the post was written by AI
The honest takeaway: this is 100% written by an LLM.
I sort of do this whenever an interesting new model comes out.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#69[flagged]
It's funny, when I saw the title I was hoping the article would include some sort of blind ranking, where you could see the outputs (without knowing which model they came from) and score them on some criteria. Could have been a fun way to get a better ranking of the results.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#70If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).