Live data from Hacker News

CursorBench 3.1

cursor.com

21–30 of 107 posts

Re: CursorBench 3.1

#22
post #6

I'm pretty baffled by their choice of axes. I would have thought that the left was the cheapest, not the most expensive. I appreciate that this layout means that top right can be best, but it's still unintuitive to have this backwards cost axis IMO. Putting that aside, I spend all day every day implementing very, very hard things right on the edge of what agents are (barely, sometimes) capable of, and I have had to k…

You can set GPT 5.5 to 1M context mode in Cursor but it costs more after the default 272k.

Re: CursorBench 3.1

#23

I feel like this benchmark reiterates my disbelief that anyone uses the latest Anthropic models for any productive work. They seem to be the best at burning tokens and spawning unnecessary subagents even for well-defined and tightly scoped tasks. Can we get a count of people that have had Claude read irrelevant documents or perform unnecessary web searches even when told not to from the beginning? I'm starting to won…

> I'm starting to wonder if this increased token usage is inadvertently bleeding into how Anthropic actually trains their model

Related: Sonnet 5’s new tokenizer increases token usage by 30%. (https://simonwillison.net/2026/Jun/30/claude-sonnet-5/)

Re: CursorBench 3.1

#24
post #11

I'm a bit skeptical. Cursor's benchmark finds that Cursor's model (Composer 2.5) is basically as good as Opus 4.8 max and GPT-5.5 xhigh, but at a fraction of the price. Artificial Analysis' testing shows Composer 2.5 to be pretty far behind: https://artificialanalysis.ai/agents/coding-agents . You look at the DeepSWE benchmark (which is probably the hardest to game at this point) and GPT-5.5 xhigh gets a 64, Opus 4.8…

By the same token, Fable 5 is given a score of 77 vs 76 for GPT 5.5

Re: CursorBench 3.1

#25
post #6

I'm pretty baffled by their choice of axes. I would have thought that the left was the cheapest, not the most expensive. I appreciate that this layout means that top right can be best, but it's still unintuitive to have this backwards cost axis IMO. Putting that aside, I spend all day every day implementing very, very hard things right on the edge of what agents are (barely, sometimes) capable of, and I have had to k…

opus@max is on average worst than opux@xhigh

for supporting evidence, see first chart here: https://www.anthropic.com/news/claude-fable-5-mythos-5

Re: CursorBench 3.1

#26
everytime a new benchmark appears, Chinese models are far lower than the level where they are supposed to be according to existing benchmarks. then after a while they recover :)

Re: CursorBench 3.1

#27
Would like to see wall times. I feel that’s the part that annoys me most, my tasks aren’t particularly challenging I want them done fast

Re: CursorBench 3.1

#28
Do these benchmarks even add any value at this point? This one is basically Cursor saying that their model is as good as the frontier ones at a fraction of the price. The independent benchmarks are probably part of training data now and the models are pattern-matching against them all the time. The final test of a model (and the harness, probably) is how good it works FOR YOU - since most of the models can pretty much do most of our tasks on a daily basis - it boils down to which one has the least friction to its usage.

Re: CursorBench 3.1

#29
post #6

I'm pretty baffled by their choice of axes. I would have thought that the left was the cheapest, not the most expensive. I appreciate that this layout means that top right can be best, but it's still unintuitive to have this backwards cost axis IMO. Putting that aside, I spend all day every day implementing very, very hard things right on the edge of what agents are (barely, sometimes) capable of, and I have had to k…

It’s Gartner. Top-right is where you want to be.

Re: CursorBench 3.1

#30
post #19
post #18

backwards X axis? is there a reason for that? it looks ridiculous

It looks very natural, cheaper is better after all. Performance axis going up, and cheapness axis going up match each other.

gp's argument is that cheapness is a construct, derived from the real, and natural, cost parameter which most people are naturally accustomed to interpreting as increasing from left to right. cheapness would then replace the cost label, and feel natural. alas, this is not what we have here.
Post reply on HN