Live data from Hacker News

SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

cognition.com

11–20 of 151 posts

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#11

Unrelated: what's the point of "*equal contribution"? Why would someone specify this

Because papers are often referred to by the first author’s name, and often the first author is the primary researcher and therefore deserves the extra credit. When two or more primary authors are equally involved, they’ll often do a random ordering but annotate this so that no one thinks one did more than the others.

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#15
post #11

Unrelated: what's the point of "*equal contribution"? Why would someone specify this

Because papers are often referred to by the first author’s name, and often the first author is the primary researcher and therefore deserves the extra credit. When two or more primary authors are equally involved, they’ll often do a random ordering but annotate this so that no one thinks one did more than the others.

Interesting. Thank you

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#16
post #4

These models are never as good, the benchmarks dont tell the full story

Funny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.

all the open source models are a waste of time relative to the bleeding edge from openai/anthropic

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#18

A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…

Would love to see these companies use benchmarks done by third parties.

they are right there? it shows swe-bench multilingual and terminal bench

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#19
Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot.

What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.

I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.

1. https://cursor.com/blog/composer-2-5

Re: SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

#20

A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technic…

> "A company whose first demo was completely fraudulent"

Could you expand on this?

Post reply on HN