Unrelated: what's the point of "*equal contribution"? Why would someone specify this
SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
11–20 of 151 posts
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#12Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#13Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#14These models are never as good, the benchmarks dont tell the full story
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#15Unrelated: what's the point of "*equal contribution"? Why would someone specify this
Because papers are often referred to by the first author’s name, and often the first author is the primary researcher and therefore deserves the extra credit. When two or more primary authors are equally involved, they’ll often do a random ordering but annotate this so that no one thinks one did more than the others.
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#16These models are never as good, the benchmarks dont tell the full story
Funny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#17Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#18A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…
Would love to see these companies use benchmarks done by third parties.
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#19What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.
I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.
Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
#20A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…
Could you expand on this?