Live data from Hacker News

SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

cognition.com

81–90 of 151 posts

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#81
post #26

Okay, let's give software engineers a break for a bit and focus on obsoleting other high-linguistic context occupations.

And to do that you’ll need development so until we’re all out of a job they’ll keep pushing. Once automating is automated it’s done.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#82
post #79

Earlier quoted context omitted.

id rather iterate multiple times than wait 15 minutes to notice it made a mistake.

Again, my point is exactly the opposite. Higher quality implies a mistake isn't made in a significant % of cases.

It's a lossy conversion though. "Mistake" is relative to the stated goals and specifications which are often heavily lacking. So unless you write with a high degree of architectural and implementation specificity then it might make very high quality code that is still not what you wanted.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#83
post #65
post #60

While I am skeptical of the results here, I am very excited for this new trend of making models faster. Running capable models at 1k TPS is more valuable for me than running better models at 30 TPS. I can only imagine the trend continues to move from "let's only make models smarter" to just incremental intelligence gains but with step improvements in speed.

Why? I'm personally on the opposite end. Less babysitting/higher quality means more time goes back to me/the user. 1000tps of bad code means you have to keep validating the output and circling back.

High tps is good for deeper agent thinking loops and openclaw etc. I was running cerebus recently doing some data heavy tasks, it managed to crash the server I was submitting posts to. 6 hour task down to ~1hr

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#84
post #31

Earlier quoted context omitted.

It's based on an open weight model (Kimi 2.7) so shouldn't it also be open weight?

There is no obligation to do that. I think the landscape would be very different now if one of the big labs had released an earlier “frontier” model under copyleft that requires sharing fine tunes. I hope it still happens.

Dario is convinced that will create SkyNet, and so no, it will never happen. Only the blessed members of the True Church Of Effective Altruism can approach the Ark of the Covenant. The unwashed cannot be trusted.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#85
post #31

Earlier quoted context omitted.

There is no obligation to do that. I think the landscape would be very different now if one of the big labs had released an earlier “frontier” model under copyleft that requires sharing fine tunes. I hope it still happens.

Dario is convinced that will create SkyNet, and so no, it will never happen. Only the blessed members of the True Church Of Effective Altruism can approach the Ark of the Covenant. The unwashed cannot be trusted.

Rationalism and EA is Scientology for the Bay Area.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#86

Would have been worth a consideration if it could have been used beyond it's own harness. Unfortunately, doesn't seem to be the case. https://x.com/theodormarcu/status/2074896486047834380

Harness-wrapper tools that support multiple harnesses and allow sharing workspace features (skills, slash commands, etc.) between them will be meta.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#87
post #75

On artificialanalysis.ai, Kimi 2.7 Code is way worse than GLM 5.2 at everything (general intelligence, coding, agentic tasks). But here, both Kimi 2.7 and its derivative SWE-1.7 are ahead of GLM 5.2. This tells me the benchmarks they use are cherry-picked.

> This tells me the benchmarks they use are cherry-picked.

Which benchmarks would you have chosen instead, and why?

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#88

We need more models that optimize for coding and that can be cheaper than frontier models, like what SWE 1.7 and composer 2.5 are trying to do. I don't think there's an effort to make something GLM-5.2 level but focused only on coding.

Defining what "coding" means now, and how quickly we fall off the capability cliff seems increasingly important. Today my "coding" sessions often enough begin with real life problems, where I discuss domain or inter-domain things, ranging from business, economics, psychology, etc. Being able to do all of that with one model is something I am willing to pay a premium for. Of course not having to pay the premium, becau…

> Today my "coding" sessions often enough begin with real life problems

intuition is that your sessions consists of 10% of domain related reasoning, and 90% of code plumbing. Those 90% could be moved to cheap and efficient specialized and focused model.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#89
Cognition... oh what a ride... We were customers when they acquired Windsurf, stopped offering customer support, raised prices, dismantled the brand, and raised prices again. We are not customers anymore. Benchmarks are not the only thing to worry about when you are using models.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#90

Earlier quoted context omitted.

Qwen was doing something like this with their coder models. But alas, they seem not to be releasing those anymore. Last one was Qwen3-coder-next.

Its crazy that OpenAI and Anthropic themselves aren't doing that. No attempts at reducing inference cost for code as far as I know from them.

My speculation is that frontier models are MoE, and they just have some number of experts for coding.
Post reply on HN