So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
They can sleep just fine being the only player in town actually not loosing subsidized money.
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
221–230 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#222The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
"completely meaningless" "the only purpose" There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your post really says is that benchmarks only show an average over multiple tasks. Yes, obviously, the point of a statistic is to summarize.
disagree
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#223So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#224The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.
I've also found that Kimi K3 on Max reasoning is benchmaxxing a little bit, High is probably enough for most dev work as long as you have good tooling and a good, detailed plan (which you can create on Max reasoning if you want).
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#225Earlier quoted context omitted.
Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol
It is breaking competition. Would you like all products everywhere be priced like their producers want?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#226Earlier quoted context omitted.
Yes. I do wish there were benchmarks for specific tech stacks. I.e, if I have an Elixir/Phoenix project, which model performs best (idiomatic, etc.) in 2026? Of course it will be somewhat subjective. And I can hang around those communities for opinions. But it might be useful in a world where it's impractical to constantly compare them all, and it varies pretty widely.
This absolutely should be a thing, but it'll have to be a per-community thing, them building their own dataset and creating their own evals (similar to how people do for production workloads). Although: "(idiomatic, etc.)" I don't think that should be part of the aim (or it should be under-weighted), because... you can just provide guidance on how to do things more idiomatically, rather than depend on that knowledge…
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#227Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#228Earlier quoted context omitted.
they likely tune their models for areas where they have their money: search, ads, youtube, etc.
I wonder if that'll be a mistake as LLMs are used for internal LLM R&D. Either Google will not take this approach, use a 3rd party model (weird, data leak risk?), or use a non-public internal model (big sunk dev cost with no recoup by trickling it to public).
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#229So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
1. Google Books, Youtube and the Google Search index all provide vast amounts of legally acquired training data.
2. They can easy people into AI using the info box. I think this strategy is working even if it does cannibalize their main revenue source. Better than just withering and leaving all of the money to OpenAI/Anthropic. I would not be surprised if Google has significant layoffs due to reduced ad revenue at some point, but I think they'll still be on top.
3. They already have their hooks into people's lives through Gmail, Google Calendar, Android, etc. The only other companies that come close are Apple (but for a much smaller number of people), and Microsoft (but only for business).
The fact that Google's models might be 20% worse, or a few months behind Anthropic's is completely insignificant in comparison to those things.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#230The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
If you click on the link you will see that it's not "one single metric" there is literally all the metrics so you can make an informed decision