Earlier quoted context omitted.
I want this with smaller models as well like Gemma 4 or Qwen 3.6
Awesome - will work on getting those in.
We made Grok 4.5, GPT-5.5, and Claude build the same apps
81–90 of 97 posts
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#82How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?
2. Dario is an idiot for not realising his dataset, workflow and model are going to be copied when he uses spacex datacenters
3. Grok has a special fan base that promote it everywhere they go
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#83[flagged]
Nothing wrong with that.
In this case, the requirement is clear enough, and the result is similar enough to judge it subjectively.
Even code-wise, if tests pass and the features are there in all cases, the rest of what matters (architecture, code quality, style, readability) is all subjective too.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#84[flagged]
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#85If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonnet 5 should be pushed).
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#86Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#87Tried at work , this release def a moment I will remember. My work is not the same . The model is the first model that offer exactly as I want : For hard tasks , that needs precision I will wait and pay expensive tokens For everything else , query data , logs, rolling out releases , I’m using grok and it’s much better vs other tools and much cheaper too .
Who the hell tries something that's been out a few hours and says "My work is not the same"?
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#88[flagged]
lol. These are the comments that keep bringing me back here. Great feedback delivered with a dash (or more) of salt.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#89Earlier quoted context omitted.
It's kinda logical. Most people, individually, have a somewhat unique writing style. So if there's one set of writing that's very formulaic and consistent and you build an averaging machine it's going to converge on that formulaic style because everyone else's writing style is going to be much closer to n=1.
Kenyans who provided the data for RLHF liked that style. That's the most of it. And a transformer is not an "averaging machine", it's a prediction machine.
Sounds like a world detail from Snow Crash.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#90So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
GPT was the worst on the Rubik's cube