Live data from Hacker News

We made Grok 4.5, GPT-5.5, and Claude build the same apps

tryai.dev

31–40 of 97 posts

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#31
post #26

Earlier quoted context omitted.

That was the honest giveaway ... :-D

I did a quick skim and the usage of phrases like "snappy stylist" and "speed-and-value monster" were what instantly stuck out to me as AI. I decided I probably didn't need to actually read the article after that.

This is the real unlock of the speed and value monster.

I am trying to figure out how many LLM converged on a writing style that resembles a LinkedIn MBA true believer. Maybe because there was just such a sheer mass of corporate-speak drone writing out there in the wild in the training data set?

But more seriously, is there a firefox extension that 'skims' the text body content of a page and puts some kind of "this was probably written by AI" meter, gauge, number or indicator in the top menu bar adjacent to the URL bar? It could even be color coded in various shades from green, yellow, orange, red. If there isn't, it sure seems like something that would be good to have.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#34

So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?

GPT was the worst on the Rubik's cube

They gave grok 2nd try on that one. Now one shotting a webpage is a dumb metric but if that is what you are testing it did the worse.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#35
I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d. Fabel 5 is just rock solid.

Yes, subjective. But it matches my repeated experiences with these models for what it is worth.

Re: We made Grok 4.5, GPT-5.5, and Claude build the same apps

#40
post #26

Earlier quoted context omitted.

I did a quick skim and the usage of phrases like "snappy stylist" and "speed-and-value monster" were what instantly stuck out to me as AI. I decided I probably didn't need to actually read the article after that.

This is the real unlock of the speed and value monster. I am trying to figure out how many LLM converged on a writing style that resembles a LinkedIn MBA true believer. Maybe because there was just such a sheer mass of corporate-speak drone writing out there in the wild in the training data set? But more seriously, is there a firefox extension that 'skims' the text body content of a page and puts some kind of "this w…

It's kinda logical. Most people, individually, have a somewhat unique writing style. So if there's one set of writing that's very formulaic and consistent and you build an averaging machine it's going to converge on that formulaic style because everyone else's writing style is going to be much closer to n=1.
Post reply on HN