Live data from Hacker News

Qwen3.7-Max: The Agent Frontier

qwen.ai

1–10 of 317 posts

Re: Qwen3.7-Max: The Agent Frontier

#6
post #4

It is super strange that all last (3?) releases they keep comparing older models such as Opus-4.6.

Some of it’s probably timing. Some of it is wanting to look good. That said, I just went to the claw-eval site, and neither 4.7 nor 5.5 from oAI are listed on the benchmarks. So there’s also just the time from others to get benchmarking done and published.

Re: Qwen3.7-Max: The Agent Frontier

#10
post #2

These are very good numbers. I still don’t get why they don’t compare against latest competitor versions in these posts, it’s not like we’re all not going to notice.

I think its part of the expectation setting (with a side of we did our distillation/ eval harness on a specific model).

if they say it's 4.7 comparable, it anchors that into your head as the model to evaluate against.

Post reply on HN