Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

51–60 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#53
Half year ago I tried to use Codex, Claude and Gemini build the same scripts to automate various things on my machine. Claude was the clear winner back then, making the most reasonable assumptions, presenting results in the easiest-to-read format, writing runnable script with minimum dependency. Half year later I think Codex and Claude models have both advanced a lot, but Gemini is still lackluster. Gemini could catch problems when reviewing Claude/Codex's design plans and code, but it's hard to make Gemini make complex plans or implement complex code by itself.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#55
post #49

If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

I want this with smaller models as well like Gemma 4 or Qwen 3.6

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#57
post #49

If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

I want this with smaller models as well like Gemma 4 or Qwen 3.6

Awesome - will work on getting those in.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#58

Earlier quoted context omitted.

This is the real unlock of the speed and value monster. I am trying to figure out how many LLM converged on a writing style that resembles a LinkedIn MBA true believer. Maybe because there was just such a sheer mass of corporate-speak drone writing out there in the wild in the training data set? But more seriously, is there a firefox extension that 'skims' the text body content of a page and puts some kind of "this w…

It's kinda logical. Most people, individually, have a somewhat unique writing style. So if there's one set of writing that's very formulaic and consistent and you build an averaging machine it's going to converge on that formulaic style because everyone else's writing style is going to be much closer to n=1.

Kenyans who provided the data for RLHF liked that style. That's the most of it. And a transformer is not an "averaging machine", it's a prediction machine.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#60
post #49

If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).

Really nice site! From your experience, what’s your go-to model for nice storefronts?
Post reply on HN