Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

11–20 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#11
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#12
(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing

Agent: https://arena.ai/leaderboard/agent

Web dev: https://arena.ai/leaderboard/code/webdev

Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#13
post #7

The cost seems to be using the wrong symbol: ¢ vs $

Nope, they're that cheap. E.g. Grok 4.5 is $.02 to $.06 per million tokens. A 400 token reply costs ~.002¢

https://www.tryai.dev/models/grok-4.5

Update: kibae above and below is correct and I'm not. They have fixed their blog post.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#14
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

It's a preemptive defense against methodology cynicism seen often on sites including but not limited to Hacker News. I've been guilty of including such defenses myself over the years because I've gotten annoyed with receiving such cynicism.

Look at the top comment on their previous HN submission: https://news.ycombinator.com/item?id=48839886

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#15
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens

Don't be fooled, there is politics, opinions, and less rigor in science as well.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#17
My concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chunk of that being pre-existing information.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#18
post #7

The cost seems to be using the wrong symbol: ¢ vs $

Nope, they're that cheap. E.g. Grok 4.5 is $.02 to $.06 per million tokens. A 400 token reply costs ~.002¢ https://www.tryai.dev/models/grok-4.5 Update: kibae above and below is correct and I'm not. They have fixed their blog post.

Grok 4.5 is $2/$6 there's no model anywhere close to that cheap

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#19
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

because if you don’t put this disclaimer the top comment is always "Acthually this isn't real science because you didn't publish your P value" so you can't win.

also the article itself is clearly LLM generated though

Post reply on HN