Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

1–10 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#3

Interesting tests being done but I can't help but think it limits testing innovation in some way given that the requested apps are essentially all clones of others

I hear this take a lot, but every app I’ve ever built was like 80% similar to every other app out there. The unique/ creative part of an app is not the bulk of it, and LLMs have been pretty good at helping me explore the 20%, too.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#4
Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today:

https://arena.logic.inc/

It's really interesting to see the Sol/Terra/Luna apps side-by-side.

I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much cheaper, but it seems like Terra is likely a more optimal time/cost balance for most use cases.

The Terra quality is usually nearly as good as Sol, but much faster and cheaper. I do appreciate Sol's design sensibilities (see, for example, the audio sequencer). It's the first model in a while that is clearly distinct on that front. They'd all converged to very similar visuals for a while.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#6

   "This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. 

Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#8

Interesting tests being done but I can't help but think it limits testing innovation in some way given that the requested apps are essentially all clones of others

I hear this take a lot, but every app I’ve ever built was like 80% similar to every other app out there. The unique/ creative part of an app is not the bulk of it, and LLMs have been pretty good at helping me explore the 20%, too.

Calculator / Rubik's cube / game of life apps should be very close to 100% identical, right? I don't see the point of asking an AI for one of these when there are dozens (hundreds?) of repos that all have exactly what you want.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#10
> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion.

My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.

Post reply on HN