Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

31–40 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#32
post #18

Earlier quoted context omitted.

Grok 4.5 is $2/$6 there's no model anywhere close to that cheap

The numbers come from the tryai.dev link: How much does Grok 4.5 cost on TryAI? Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.

Slop pricing pages? https://www.tryai.dev/models/claude-fable-5 says Fable costs $0.1/$0.5. Can't wait to use it at those prices!

(edit: these have been fixed shortly after)

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#33
post #18

Earlier quoted context omitted.

Grok 4.5 is $2/$6 there's no model anywhere close to that cheap

The numbers come from the tryai.dev link: How much does Grok 4.5 cost on TryAI? Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.

Those numbers are incorrect unless they got a deal that's 100x cheaper than API pricing which is unlikely. They updated the original post with the correct costs.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#34
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

> how hard would it have been really to type these two sentences by hand, in your own natural voice

On the other hand, do we have to complain about every seemingly AI written text?

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#35

Obviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5 (edit: seems to be an issue with inline videos)

2/5 isn't quality, it's consistency as written there. The full links are at the bottom. Most of Spark's attempts are failures:

https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...

https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...

https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#37
post #25
post #4

Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/Terra/Luna apps side-by-side. I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much chea…

What caught my eye was: Model Lines of Code File Size Gzip Size GPT-5.6 Sol 1,264 35.5 KB 10.0 KB GPT-5.6 Terra 827 20.0 KB 6.7 KB

Yea, that's an interesting result as well. The Terra apps don't feel 35% less feature-rich. So it seems quite token efficient.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#38
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

Yeah, I can't figure out where this "voice" comes from and it is so impossible to get rid of. It is so grating.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#39
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?

Yes? AI generated text is explicitly disallowed on HN, so it’s not crazy to expect that a similar standard be used for linked content.
Post reply on HN