Live data from Hacker News

The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

minimaxir.com

91–100 of 116 posts

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#91
post #11

OpenRouter rankings frustrate me, because they show the total number of tokens but they provide no indication of how many unique users a model has. Which means if a surprise model tops the leaderboard one week we can never be sure if it was because a single whale user pushing billions of tokens a day switched to it, or if it represents a genuine community trend towards that model.

(openrouter co-founder here) Yeah we should do something to indicate cardinality. I can share that there can often (I'm talking generally; not related to this model in particular) be e.g. a very large app that can be pushing a lot of volume. But in almost all cases that app has a large number of end users. Hypothetically, for instance, would Cursor be consider one user, or millions? Will think about it! Thanks for th…

I'd consider Cursor one user because it's one entity that made an editorial decision about which model to make available to their own community.

If you treated Cursor as millions of users it might look like millions of people independently chose a new model when actually it was Cursor making the choice for them - and the thing I care most about is how many choices were made that selected a model and put it above the others.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#92
post #21

I’ve tested this model on four of my benchmarks: https://github.com/lechmazur/buyout_game 10th out 36. https://github.com/lechmazur/pact/ 14th out 25. https://github.com/lechmazur/nyt-connections/ 60th out 81. https://github.com/lechmazur/debate 16th out of 29.

Good stuff!

Is there a reason you change the leaderboard graphs for the third and fourth one?

Also: would be great to have an overview page with a summary over all test, like a total score or similar.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#93
post #11

OpenRouter rankings frustrate me, because they show the total number of tokens but they provide no indication of how many unique users a model has. Which means if a surprise model tops the leaderboard one week we can never be sure if it was because a single whale user pushing billions of tokens a day switched to it, or if it represents a genuine community trend towards that model.

(openrouter co-founder here) Yeah we should do something to indicate cardinality. I can share that there can often (I'm talking generally; not related to this model in particular) be e.g. a very large app that can be pushing a lot of volume. But in almost all cases that app has a large number of end users. Hypothetically, for instance, would Cursor be consider one user, or millions? Will think about it! Thanks for th…

One idea I had was to count # of distinct API keys that have spent atleast $100 (number's flexible), which would be enough to provide guidance on if the traffic is from a single power-user.

In the Cursor case which is BYOK, that would count as distinct API keys.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#95
post #4

> Two new models are now beating LLM darling Claude in terms of token usage and by more than 50%? Time for a reminder that OpenRouter leaderboards only show tokens sent through OpenRouter, which most Anthropic API users don’t use.

That doesn't mean it can't be used as a market signal. These 2 things can both be true at once.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#96

Earlier quoted context omitted.

Scraping the internet isn't a copyright violation. Using it for LLM training is much more transformative than Google and Internet Archive, which are legal.

Your right, scraping is legally protected. It's reproducing verbatim text that's a violation, which is why LLMs still clumsily refuse to produce song lyrics. They are capable of copyright violations and have to be 'aligned' not to get their providers sued.

Verbatim reproduction is neither necessary nor sufficient to create a copyright violation.

"Copyright violation" is what we call the set of things that destroy the incentive for people to create original work by unduly benefitting from someone else's original work.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#97

Earlier quoted context omitted.

Agreed. My little solo dev SaaS app’s production pipelines push almost two billion tokens a day.

Haha, I never tire of the AI haters downvoting stuff like this. Down with reality!!

Or, everyone finally realizes that token burn is not the same as productivity. Maybe they just down voted for the questionable spending brag.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#99

High token usage cuz it's free doesn't count

The post goes into that issue. Throughly. The numbers at the beginning of the post are weekly aggregate values well after the endpoint was paid-only.

The post is wrong, it's still free, see - https://openrouter.ai/tencent/hy3-preview:free it's free in kilo.ai https://kilo.ai/models/tencent-hy3-preview-free It's free in a lot of places.

Re: The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin

#100
post #4

> Two new models are now beating LLM darling Claude in terms of token usage and by more than 50%? Time for a reminder that OpenRouter leaderboards only show tokens sent through OpenRouter, which most Anthropic API users don’t use.

That doesn't mean it can't be used as a market signal. These 2 things can both be true at once.

I'm pretty sure the popularity came from being free at some point
Post reply on HN