> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
41–50 of 95 posts
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#42> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.
So no matter how good or thoughtful the writing is it gets tiresome.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#43Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#44Obviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5 (edit: seems to be an issue with inline videos)
2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures: https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/... https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#45Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#46(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing Agent: https://arena.ai/leaderboard/agent Web dev: https://arena.ai/leaderboard/code/webdev Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.
Always fun to ask them to recreate classic demoscene effects (sadly they're still pretty bad at generating music, though at least claude seems to create decent synths).
I keep trying to get them to recreate the fluid+particle stuff from Agenda Circling Forth etc., but even giving them the blog posts describing the implementation (and screenshots) they're still pretty bad.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#47Earlier quoted context omitted.
Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.
The problem is that everyone (well, too many) is now using the same "voice" when writing. So no matter how good or thoughtful the writing is it gets tiresome.
LLMs have a bunch of tricks for dressing up their infodumps, but they are almost purely infodumps, and no real opinion comes through. There's no sense of more importance to one statement or the other, it's all monotone (and usually over the top.)
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#48> Draw a horse riding an astronaut in svg
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#49Sign-in via Google is broken - it redirects back to localhost from Supabase :)
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#50(LM)Arena is basically this. IMO it’s the best benchmark that avoids benchmaxxing Agent: https://arena.ai/leaderboard/agent Web dev: https://arena.ai/leaderboard/code/webdev Currently Fable and 5.6 are neck and neck on web dev which is basically the same finding as this.