Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

51–60 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#51

My concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chun…

> without a large chunk of that being pre-existing information.

Is there any evidence that novel reasoning is present in LLM? I've never been able to make that work, and I believe Apple's paper some time ago was good evidence that it doesn't exist. In my experience, sparse latent spaces result in a complete, comical, failure in reasoning.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#52
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.

You can suggest a different one and it works pretty well, the bar is very low.

I asked for hemingway in a planning document one time and the result was highly amusing to everyone. "We will not wreck it with small greeds." was an all time favorite for me.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#53
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Indeed, although it is important to note that science is a proper subset of "using reason to solve problems."

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#54
post #39

Earlier quoted context omitted.

> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?

Yes? AI generated text is explicitly disallowed on HN, so it’s not crazy to expect that a similar standard be used for linked content.

But AI generated linked content was explicitly not disallowed. At least not at this time.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#55
I think there's approximately zero value in seeing how a model can turn 100 tokens into a 100k. What workflow is that? It's not useful in the real world.

I want to know how well it can follow instructions, manage various potentially competing desires in the context, and so on. It's much more interesting how it can turn 100k tokens (e.g. a codebase and lots of tool calls) into 100 tokens.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#57

Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.

It's not, but it's trying to bring any level of objective measure in this realm, vs just going off of vibes.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#58
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

tower defense against pedantic autists who miss the social cue of “does it matter?”

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#59

> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion. My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.

Say you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input?

Measuring "intelligence" is hard, but giving an "intelligent" entity tasks, and seeing what comes out, and then comparing the output with others, seems like a very reasonable, relative, way to do it.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#60
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.

Man invents text generation machine.

Man generates text.

Post reply on HN