Live data from Hacker News

Browser Agent Benchmark: Comparing LLM models for web automation

browser-use.com

1–7 of 7 posts

Re: Browser Agent Benchmark: Comparing LLM models for web automation

#6

It's lacking the best model (Opus 4.5) on the benchmark tho.

Yeah but then their own product might not score the highest.

Exactly why I'm pointing it out, which feels a bit corrupt, but understandable.

Re: Browser Agent Benchmark: Comparing LLM models for web automation

#7

Earlier quoted context omitted.

Yeah but then their own product might not score the highest.

Exactly why I'm pointing it out, which feels a bit corrupt, but understandable.

tbh i was a bit cranky yesterday - even if they are #2 on a legit benchmark that would be impressive