I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d.…
We made Grok 4.5, GPT-5.5, and Claude build the same apps
71–80 of 97 posts
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#72Why not wait one more day for GPT-5.6?
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#73I'd like to see the comparisons with DeepSeek, Qwen, Mimo, Kimi and GLM
I just did the tests, Mimo and GLM delivered working cubes but GLM was the only visually perfect with smooth movements and great effects. GLM is the clear winner: https://chat.z.ai/space/t19sx5kvw631-art
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#74How can grok create a coding LLM at the same level as OpenAI or Anthropic when they don’t have the same amount of AI talent as the other companies by an order of magnitude? Is it really that easy to train a coding model like that?
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#75I am 99% sure the post was written by AI
> “snappy stylist”
Funny you can tell its slop just by this
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#76Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#77I’m spending a significant portion of my day waiting for agents to execute. What’s more interesting to me than time-to-first token or latency is the time it takes for the agent to execute, from starts to finish, excluding when it’s waiting on a human.
On which note I recently convinced codex to use the ChatGPT web client to run subagents. Means I only pay for the slavemaster, and all the slaves I can eat for $20 a month. Actually works surprisingly well - I currently have it crunching through a large dataset, which would have taken weeks on a single thread - started last night, nearly done this morning. $20.
Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#78Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps
#79Grok failed the Rubik's cube. I pressed Scramble twice and then solve and it didn't solve the cube. Opus did.