Earlier quoted context omitted.
Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
yeah but why waste your time on these models, just use the one that gets the better results
SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
41–50 of 151 posts
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#42We need more models that optimize for coding and that can be cheaper than frontier models, like what SWE 1.7 and composer 2.5 are trying to do. I don't think there's an effort to make something GLM-5.2 level but focused only on coding.
Qwen was doing something like this with their coder models. But alas, they seem not to be releasing those anymore. Last one was Qwen3-coder-next.
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#43Feels like they discovers that if you build your own benchmark, you can win it
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#44Earlier quoted context omitted.
Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
yeah but why waste your time on these models, just use the one that gets the better results
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#45Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#46Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#47A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#48Earlier quoted context omitted.
Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
yeah but why waste your time on these models, just use the one that gets the better results
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#49Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot. What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW. I'd posit that it's not deliberate deception, but for both companies their tra…
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#50Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot. What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW. I'd posit that it's not deliberate deception, but for both companies their tra…
I think it's also telling that they left out the usual hallmarks of the Pareto distribution: GLM 5.2, Qwen 3.7, Minimax M3, and Mimo 2.5 https://arena.ai/leaderboard/code/webdev/pareto
They did not.