Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

81–90 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#81

Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.

One-shot benchmarks are great for me as a solo creator, since they slightly correlate to whether the better frontier models (Opus and Fable for me) make better decisions about things I didn't spec, or whether they'll give me better suggestions right off the bat.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#83

Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.

I imagine one could one-shot a basic app and then feed feature requests one by one, sounds like an obvious way to benchmark architecture/maintainability

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#84

Grok redeemed? But I wonder how long they will keep it cheap

I gave up on Grok.

It constantly ignores explicit instructions (e.g. do NOT remove existing comments) and it's not nearly as intuitive in knowing the right questions to ask as gpt, claude or gemini in my experience from using all of them.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#86
post #24

> Separate question, separate table. This is our standard latency harness (three short prompts, five reps, 400-token cap), not the build tasks. tok/s is output tokens over wall-clock, uniform for all. > so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and…

For the arguments sake: What if that is the authors natural voice?

[dead]

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#87

"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.

I hear "Honestly" more often from Anthropic than I ever do from all humans.

And it uses the word in weird ways I have never experienced. I'm genuinely curious what's in the training set that caused this tic.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#88

"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.

I hear "Honestly" more often from Anthropic than I ever do from all humans.

Claude training should learn that sometimes people use the word "honestly" because they'd otherwise be lying most of the time.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#89

Earlier quoted context omitted.

> how hard would it have been really to type these two sentences by hand, in your own natural voice On the other hand, do we have to complain about every seemingly AI written text?

It is just horrible writing style. It doesn't particularly matter that AI wrote it.

Don't read it then. Why waste enrgy on complaining?
Post reply on HN