We made Grok 4.5, GPT-5.5, and Claude build the same apps
51–60 of 97 posts
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#52What’s more interesting to me than time-to-first token or latency is the time it takes for the agent to execute, from starts to finish, excluding when it’s waiting on a human.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#53Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#54> The receipts: speed and cost I don't get why cost per reply is at all relevant here? Why do so few who attempt comparisons actually compare dollars per task.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#55If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#56I am 99% sure the post was written by AI
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#57If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).
I want this with smaller models as well like Gemma 4 or Qwen 3.6
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#58Earlier quoted context omitted.
This is the real unlock of the speed and value monster. I am trying to figure out how many LLM converged on a writing style that resembles a LinkedIn MBA true believer. Maybe because there was just such a sheer mass of corporate-speak drone writing out there in the wild in the training data set? But more seriously, is there a firefox extension that 'skims' the text body content of a page and puts some kind of "this w…
It's kinda logical. Most people, individually, have a somewhat unique writing style. So if there's one set of writing that's very formulaic and consistent and you build an averaging machine it's going to converge on that formulaic style because everyone else's writing style is going to be much closer to n=1.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#59Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#60If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).