Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

21–30 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#23
post #18

Earlier quoted context omitted.

Nope, they're that cheap. E.g. Grok 4.5 is $.02 to $.06 per million tokens. A 400 token reply costs ~.002¢ https://www.tryai.dev/models/grok-4.5 Update: kibae above and below is correct and I'm not. They have fixed their blog post.

Grok 4.5 is $2/$6 there's no model anywhere close to that cheap

The numbers come from the tryai.dev link:

  How much does Grok 4.5 cost on TryAI?
  Grok 4.5 is Input: $0.02 / 1M tokens, Output: $0.06 / 1M tokens. There is no subscription — you pay only for what you use.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#24
> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all.

> so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes

You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.

Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#25
post #4

Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/Terra/Luna apps side-by-side. I need to add these stats somewhere in the UI, but one interesting take away: Terra took 1/2 as much wall-clock time as Sol, but Luna took more wall-clock time than Sol (by about 23%). It's still much much chea…

What caught my eye was:

            Model  Lines of Code  File Size  Gzip Size 
      GPT-5.6 Sol          1,264    35.5 KB    10.0 KB 
    GPT-5.6 Terra            827    20.0 KB     6.7 KB

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#27
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Because for its entire existence, the top HN comment on articles is typically a contrarian take or pointing out flaws. This goes double for a study, where people just hunt for some aspect of the methodology they dislike. If you don't address the flaws, then it looks like you never considered them, and the top comment will say that your entire methodology is suspect. It's super predictable to the point that you can harness this kind of reaction to get stuff on the frontpage if you really want to.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#29
post #6

"This isn't objective." Correct, and we are not pretending it is. We are not handing down a scientific verdict. Actually, you are doing rational investigation in a fuzzy probabilistic new/emergent space, with open sharing to the world. I don’t understand why people downplay themselves and put on a pedestal others supposedly serious sciences.

Ultimately advocates exist for models and there are incredible financial incentives for some to be advocates, so authors are guaranteed someone being mad if their horse doesn't perform well.

Given that type of reaction is inevitable, it just saves the conversation.

Post reply on HN