Earlier quoted context omitted.
Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)
Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
181–190 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#182Very interesting that one of the components is "AA-Omniscience Index" AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Ma…
I think grounding has as much if not more go do with it. Google does a great job w/ grounding for obvious reasons
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#183Earlier quoted context omitted.
Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol
It is breaking competition. Would you like all products everywhere be priced like their producers want?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#184The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
But I agree, wrapping that up into a single result.. you lose all the nuance, it's just bragging rights.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#185Earlier quoted context omitted.
Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?
I think the source of this misapprehension is that, You are comparing wet work in a lab to writing code on a computer. When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero. You can screw up an infinite number of times on your way to a successful exploit. If you screw up with lethal agents in a…
Correct, there is a small number of people who have been able to do it at high cost for a long time.
Now, there is an ever-growing number of people who are able to do it at an ever-falling cost.
That's the entire issue. Do you dispute that this is what's happening?
> I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm
Did someone propose doing that? Or is this a strawman?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#186The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
The model did fine.
Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.
In this moment, andai was enlightened.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#187Earlier quoted context omitted.
Gemini knocks this out of the park, Gemini gang unite. https://share.gemini.google/34vZzlnsmTaL
I would love for Gemini to be competitive but even 3.6 flash doesn’t match sol, or opus 5, or k3
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#188#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#189Earlier quoted context omitted.
Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting
I found out it really depends on the repo I'm working on, and it's not just about being a new/old repo, there's something else which I can't really grasp yet.
The way I keep an eye on model performance is alternating between models for code review. It shows how much the model understands the codebase without disrupting my workflow.