GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
71–80 of 95 posts
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#72> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
I'm sick and tired of reading comments like these.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#73> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#74"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#75> We generated a big pile of artifacts, we are publishing all of them, and you can form your own opinion. My opinion is that spamming HN with two gimmicky "one-shot prompting shootout" marketing pieces in two days does not build confidence about either your technical or marketing expertise.
Say you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input? Measuring "intelligence" is hard, but giving an "intelligent" en…
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#76Earlier quoted context omitted.
Say you were interviewing a human, to see how capable they were. You are allowed to give them take home work. What kind of questions would you ask, or tasks would you give, to try to get a measure of their competence? If you gave them a task, would you iterate with them on the design, or would you see what they could produce on their own, without input? Measuring "intelligence" is hard, but giving an "intelligent" en…
[deleted because pointless]
I'm not sure I understand. What's not a good idea? I'm asking you how you would do it, with some possible examples. Or, are you saying it's a bad idea to try to measure how competent someone is before hiring them?
> LLMs aren't humans,
Not sure how this is relevant. My question was how to measure competence and "intelligence" for a task an entity, intelligent enough to do that task, will do. LLMs are not humans, but are usually used to complete tasks humans want completed that would usually be done by humans. That's where the most token spend is for them. Since that's what people are using them for, it seems reasonable to try to measure competence in those tasks.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#77My concern with most of these visual benchmarks, popular as they are, is that they are likely more indicative of knowledge (i.e. how comprehensive the training data is and how well it can be retrieved from the model) than of reasoning ability. I don't see in particular how a model would construct a CoT that mapped somehow to a representation of the cube geometry and its animations in latent space without a large chun…
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#78Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#79> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#80"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.
Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens