Live data from Hacker News

SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

cognition.com

41–50 of 151 posts

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#41

Earlier quoted context omitted.

Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.

yeah but why waste your time on these models, just use the one that gets the better results

I was going to respond until I saw your account name lol.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#42

We need more models that optimize for coding and that can be cheaper than frontier models, like what SWE 1.7 and composer 2.5 are trying to do. I don't think there's an effort to make something GLM-5.2 level but focused only on coding.

Qwen was doing something like this with their coder models. But alas, they seem not to be releasing those anymore. Last one was Qwen3-coder-next.

I use this model. It's pretty good but not Opus 4.8 or Fable levels obviously. I'm really hoping we get more models like it (and better) soon. I run it locally and it's great that way.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#44

Earlier quoted context omitted.

Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.

yeah but why waste your time on these models, just use the one that gets the better results

Because you can get them from more trustworthy providers or with hardware encryption.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#45
post #4

These models are never as good, the benchmarks dont tell the full story

Funny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.

It's simple: close weights = not welcome.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#46
The benchmark debate is fair, but I think the more interesting signal is how quickly coding models are becoming a category of their own rather than just smaller frontier models. More specialization, more competition on cost, and probably a lot more benchmark gaming along the way :)

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#47

A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…

To be fair it does seem like most AI startups are now like this (particularly when it comes to constantly mentioning how hard they work and ignoring consumer markets).

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#48

Earlier quoted context omitted.

Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.

yeah but why waste your time on these models, just use the one that gets the better results

I actively prefer GLM-5.2 for some tasks. For simple tasks the results are just as good as e.g. Opus, and it produces results significantly faster.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#49
post #19

Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot. What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW. I'd posit that it's not deliberate deception, but for both companies their tra…

I think it's also telling that they left out the usual hallmarks of the Pareto distribution: GLM 5.2, Qwen 3.7, Minimax M3, and Mimo 2.5

https://arena.ai/leaderboard/code/webdev/pareto

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#50
post #49
post #19

Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot. What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW. I'd posit that it's not deliberate deception, but for both companies their tra…

I think it's also telling that they left out the usual hallmarks of the Pareto distribution: GLM 5.2, Qwen 3.7, Minimax M3, and Mimo 2.5 https://arena.ai/leaderboard/code/webdev/pareto

> they left out ... GLM 5.2

They did not.

Post reply on HN