Live data from Hacker News

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

tryai.dev

91–95 of 95 posts

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#91

"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.

Can't we just take that new language that llmish is and feed it to a transformer of sort that'd get rid of those infuriating sentences?

Nothing hard. Everybody wins.

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#92

Earlier quoted context omitted.

It is just horrible writing style. It doesn't particularly matter that AI wrote it.

Don't read it then. Why waste enrgy on complaining?

For the same reasons you replied to my comment

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#93

Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.

just found a decent looking benchmark for iterative development: https://swe-milestone.com/

surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that

still doesn't measure the human-agent interaction part, but that's pure vibes atp

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#94

Earlier quoted context omitted.

Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens

Don't be fooled, there is politics, opinions, and less rigor in science as well.

The extent to which their are, is the extent to which that is not science

Re: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

#95

Earlier quoted context omitted.

Don't be fooled, there is politics, opinions, and less rigor in science as well.

The extent to which their are, is the extent to which that is not science

No True Scotsman
Post reply on HN