Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

81–90 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#81
post #57

Earlier quoted context omitted.

I want this with smaller models as well like Gemma 4 or Qwen 3.6

Awesome - will work on getting those in.

Having 'local runable' to compare would be awesome. For example I have a 48G MacBook.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#82

How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?

1. Elon throw money on Gemini folks to get them switch ship

2. Dario is an idiot for not realising his dataset, workflow and model are going to be copied when he uses spacex datacenters

3. Grok has a special fan base that promote it everywhere they go

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#83

[flagged]

>Comparison are completely subjective

Nothing wrong with that.

In this case, the requirement is clear enough, and the result is similar enough to judge it subjectively.

Even code-wise, if tests pass and the features are there in all cases, the rest of what matters (architecture, code quality, style, readability) is all subjective too.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#85
post #49

If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

This is impressive and, I think, complements well whatever benchmark is the hottest right now.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#87
post #29
post #2

Tried at work , this release def a moment I will remember. My work is not the same . The model is the first model that offer exactly as I want : For hard tasks , that needs precision I will wait and pay expensive tokens For everything else , query data , logs, rolling out releases , I’m using grok and it’s much better vs other tools and much cheaper too .

Who the hell tries something that's been out a few hours and says "My work is not the same"?

I am , because it’s smart and extremely fast . Same way as opus 4.6 was an obvious game changer

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#89

Earlier quoted context omitted.

It's kinda logical. Most people, individually, have a somewhat unique writing style. So if there's one set of writing that's very formulaic and consistent and you build an averaging machine it's going to converge on that formulaic style because everyone else's writing style is going to be much closer to n=1.

Kenyans who provided the data for RLHF liked that style. That's the most of it. And a transformer is not an "averaging machine", it's a prediction machine.

Kenya influenced the entire English language by being an outsourcing site for AI RLHF.

Sounds like a world detail from Snow Crash.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#90

So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?

GPT was the worst on the Rubik's cube

GPT doesn't render but does actually work, Grok does not fully work Two scrambles in a row and the thing is broken.
Post reply on HN