Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

41–50 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#41
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?

It is just horrible writing style. It doesn't particularly matter that AI wrote it.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#42
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.

The problem is that everyone (well, too many) is now using the same "voice" when writing.

So no matter how good or thoughtful the writing is it gets tiresome.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#44
post #35

Obviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5 (edit: seems to be an issue with inline videos)

2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures: https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...

Ah, I missed that, and didn't click through the links. Most of the videos are not showing any animation for me, only Opus / Qwen / Muse, so Grok's attempt looked broken.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#46

(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing Agent: https://arena.ai/leaderboard/agent Web dev: https://arena.ai/leaderboard/code/webdev Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.

There's a ton of arenamaxxing going on (especially from facebook), though I don't disagree that it's one of the better actual benchmarks.

Always fun to ask them to recreate classic demoscene effects (sadly they're still pretty bad at generating music, though at least claude seems to create decent synths).

I keep trying to get them to recreate the fluid+particle stuff from Agenda Circling Forth etc., but even giving them the blog posts describing the implementation (and screenshots) they're still pretty bad.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#47

Earlier quoted context omitted.

Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.

The problem is that everyone (well, too many) is now using the same "voice" when writing. So no matter how good or thoughtful the writing is it gets tiresome.

That's not true at all, human tends to come through (in a way I didn't notice pre-LLM) with varied tone, with opinions injected, and with varying degrees of weight to different components, in all but the most egregious of examples (linkedin motivation pieces, apple marketing speak).

LLMs have a bunch of tricks for dressing up their infodumps, but they are almost purely infodumps, and no real opinion comes through. There's no sense of more importance to one statement or the other, it's all monotone (and usually over the top.)

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#50

(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing Agent: https://arena.ai/leaderboard/agent Web dev: https://arena.ai/leaderboard/code/webdev Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.

Doesn't have Grok 4.5 listed yet, wonder why 5.6 is, it was released later?
Post reply on HN