Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

71–80 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#72
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

Eh? I write that way sometimes. Long before LLMs.

I'm sick and tired of reading comments like these.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#73
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

For the arguments sake: What if that is the authors natural voice?

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#75
post #59

> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion. My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.

Say you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input? Measuring "intelligence" is hard, but giving an "intelligent" en…

[deleted because pointless]

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#76
post #59

Earlier quoted context omitted.

Say you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input? Measuring "intelligence" is hard, but giving an "intelligent" en…

[deleted because pointless]

> Even if this was a good idea when applied to humans (it's not)

I'm not sure I understand. What's not a good idea? I'm asking you how you would do it, with some possible examples. Or, are you saying it's a bad idea to try to measure how competent someone is before hiring them?

> LLMs aren't humans,

Not sure how this is relevant. My question was how to measure competence and "intelligence" for a task an entity, intelligent enough to do that task, will do. LLMs are not humans, but are usually used to complete tasks humans want completed that would usually be done by humans. That's where the most token spend is for them. Since that's what people are using them for, it seems reasonable to try to measure competence in those tasks.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#77

My concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chun…

Yeah, Anthropic likely is gaining an edge in tests like this from the data they got from Canva.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#79
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

People became such a fragile snowflakes since AI popped up. Fussy because a programmer didn't produce prose they fancy.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#80
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens

Qualitative science is science.
Post reply on HN