Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

71–80 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#71

I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d.…

I get better results with Opus than Fable 5 on various oneshots (including our old friend of "generate an SVG of a pelican riding a bicycle"). (Opus is also far and away better than pretty much any other SOTA or near-SOTA model.)

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#73
post #27
post #20

I'd like to see the comparisons with DeepSeek, Qwen, Mimo, Kimi and GLM

I just did the tests, Mimo and GLM delivered working cubes but GLM was the only visually perfect with smooth movements and great effects. GLM is the clear winner: https://chat.z.ai/space/t19sx5kvw631-art

GLM-5.2 just keeps blowing me away. I pretty much use it in the mode I used to use Opus for where I have it drive a lot of the reasoning with DeepSeek or other models for the smaller subtasks/subagents.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#74

How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?

Those two companies spend all their effort patting themselves on the back about "peer reviewed studies" and posturing immeasurables like security theatre and "trust and safety". It's really not hard to believe that Grok or Chinese models can show there is no moat.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#77

I’m spending a significant portion of my day waiting for agents to execute. What’s more interesting to me than time-to-first token or latency is the time it takes for the agent to execute, from starts to finish, excluding when it’s waiting on a human.

Subagents can help enormously…

On which note I recently convinced codex to use the ChatGPT web client to run subagents. Means I only pay for the slavemaster, and all the slaves I can eat for $20 a month. Actually works surprisingly well - I currently have it crunching through a large dataset, which would have taken weeks on a single thread - started last night, nearly done this morning. $20.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#79

Grok failed the Rubik's cube. I pressed Scramble twice and then solve and it didn't solve the cube. Opus did.

I'm not sure, but it looks like none of them solve the cube. They just memorize what happened to the cube and do the reverse.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#80
post #57

Earlier quoted context omitted.

I want this with smaller models as well like Gemma 4 or Qwen 3.6

Awesome - will work on getting those in.

It'd be interesting to add Nemotron, which is quite popular on Spark alongside Qwen 3.6.
Post reply on HN