"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
11–20 of 95 posts
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#12Agent: https://arena.ai/leaderboard/agent
Web dev: https://arena.ai/leaderboard/code/webdev
Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#13The cost seems to be using the wrong symbol: ¢ vs $
https://www.tryai.dev/models/grok-4.5
Update: kibae above and below is correct and I'm not. They have fixed their blog post.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#14"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
Look at the top comment on their previous HN submission: https://news.ycombinator.com/item?id=48839886
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#15"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#16"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#17Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#18The cost seems to be using the wrong symbol: ¢ vs $
Nope, they're that cheap. E.g. Grok 4.5 is $.02 to $.06 per million tokens. A 400 token reply costs ~.002¢ https://www.tryai.dev/models/grok-4.5 Update: kibae above and below is correct and I'm not. They have fixed their blog post.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#19"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
also the article itself is clearly LLM generated though