Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

91–97 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#91
post #57

Earlier quoted context omitted.

I want this with smaller models as well like Gemma 4 or Qwen 3.6

Awesome - will work on getting those in.

It would be nice to see how "long" -- e.g. how many self-turns/tool calls/etc the prompt takes to resolve as well.

I know that models like Gemma4-e4b will take longer self-turns but IBM's Granite models will take shorter self-turns in exchange for more tool calls.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#92

[flagged]

I wouldn't be surprised if multiple versions of both the cube and Breakout code are available some place online.

https://upload.wikimedia.org/wikipedia/en/c/cd/Breakout_game...

https://upload.wikimedia.org/wikipedia/commons/1/1a/Screensh...

That the tiles are much to thick makes Grok the most reasonable result.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#93
post #27
post #20

I'd like to see the comparisons with DeepSeek, Qwen, Mimo, Kimi and GLM

I just did the tests, Mimo and GLM delivered working cubes but GLM was the only visually perfect with smooth movements and great effects. GLM is the clear winner: https://chat.z.ai/space/t19sx5kvw631-art

Tried the recent Tencent Hy3 and it produced a working cube, but plain visuals.

GLM still wins by a landslide

* I ran all agents in their web version with their default settings, I don't remember now if they were set to deep thinking or what but that's the result they produced by default.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#94
post #89

Earlier quoted context omitted.

Kenyans who provided the data for RLHF liked that style. That's the most of it. And a transformer is not an "averaging machine", it's a prediction machine.

Kenya influenced the entire English language by being an outsourcing site for AI RLHF. Sounds like a world detail from Snow Crash.

I'll take that over an LLM that's been trained by default to the tastes of "polite" Indian English speakers, repeating back to me things like kindly do the needful and revert back.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#96

[flagged]

You broke the HN guidelines badly with this, first because of the name-calling and personal attack, and second because it is a shallow dismissal in exactly the sense we ask people not to post here:

"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." - https://news.ycombinator.com/newsguidelines.html

On the plus side, your comment does contain some specific observations which could have made for a good comment—one that communicates interesting information respectfully. If you'd please post that way in the future, we'd appreciate it.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#97
post #26

Earlier quoted context omitted.

I did a quick skim and the usage of phrases like "snappy stylist" and "speed-and-value monster" were what instantly stuck out to me as AI. I decided I probably didn't need to actually read the article after that.

This is the real unlock of the speed and value monster. I am trying to figure out how many LLM converged on a writing style that resembles a LinkedIn MBA true believer. Maybe because there was just such a sheer mass of corporate-speak drone writing out there in the wild in the training data set? But more seriously, is there a firefox extension that 'skims' the text body content of a page and puts some kind of "this w…

[dead]
Post reply on HN