GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
31–40 of 95 posts
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#32Earlier quoted context omitted.
Grok 4.5 is $2/$6 there's no model anywhere close to that cheap
The numbers come from the tryai.dev link: How much does Grok 4.5 cost on TryAI? Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.
(edit: these have been fixed shortly after)
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#33Earlier quoted context omitted.
Grok 4.5 is $2/$6 there's no model anywhere close to that cheap
The numbers come from the tryai.dev link: How much does Grok 4.5 cost on TryAI? Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#34> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
On the other hand, do we have to complain about every seemingly AI written text?
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#35Obviously AI-written, but I'm confused with the results: Muse Spark has the best Rubik's cube by far, the only one properly animating, yet it gets a 2/5 (edit: seems to be an issue with inline videos)
https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...
https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...
https://d1md4c6gq9re9p.cloudfront.net/blog/gpt-5.6-buildoff/...
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#36Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#37Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/Terra/Luna apps side-by-side. I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much chea…
What caught my eye was: Model Lines of Code File Size Gzip Size GPT-5.6 Sol 1,264 35.5 KB 10.0 KB GPT-5.6 Terra 827 20.0 KB 6.7 KB
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#38> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#39> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…
> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?
Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
#40Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.