Earlier quoted context omitted.
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
I think on swebench verified luna was only like 3% points lower for 1/5 the cost Like 96% vs 93% or something
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
101–110 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#102Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#103I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#104Earlier quoted context omitted.
I'm curious about this too, and it's difficult to get any information about this given everyone has different setups, workflows and use-cases. I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (i…
There's probably some system prompt crap in claude code that tamps down that behavior? I wonder what it was even trying to do.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#105Earlier quoted context omitted.
> while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case? Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?
Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?
You are comparing wet work in a lab to writing code on a computer.
When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero.
You can screw up an infinite number of times on your way to a successful exploit.
If you screw up with lethal agents in a lab? You die.
Here's a non-exhaustive list,
Dora Lush died after accidentally pricking her finger with a needle containing lethal scrub typhus while attempting to develop a vaccine for the disease
A 23-year-old laboratory assistant at the London School of Hygiene and Tropical Medicine, was infected with smallpox after observing the harvesting of live smallpox virus from eggs without isolation cabinets at that time. The assistant was hospitalised and before being isolated, she infected two visitors to a patient in an adjacent bed, both of whom died. They in turn infected a nurse, who survived
Ebola laboratory infection by the accidental stick of contaminated needle in the United Kingdom
Researcher Nikolai Ustinov was lethally infected with the Marburg virus after accidentally pricking himself with a syringe used for inoculation of guinea pigs. The accident occurred at the Scientific-Production Association "Vektor" (today the State Research Center of Virology and Biotechnology "Vektor") in Koltsovo, USSR (today Russia).
"lethally infected with the Marburg virus after accidentally pricking himself"Anything lethal enough to kill other humans is lethal enough to kill you.
And if you don't know what you're doing — and for this argument you're saying this person has to ask a LLM "how do I spanish flu?" then they definitely don't know what they're doing, the number of ways you will die far outnumber the ways you can succeed.
And this, of course, doesn't even cover the cost of equipment, the precursors, sourcing the highly specific materials needed, then setting the equipment up... etc.
The same is true for the Bosch-Haber / Haber-Bosch process, which famously made WW1 possible. Every HS'er learns about the process and the steps. Steps that were classified once upon a time and were the subject of negotiation at the Versailles.
Does that mean a HS'er (or any adult) can set up an experiment that works at 177 times the pressure of the Earth's atmosphere to do anything at any scale without significant infrastructure and help?
The people who can do this are domain experts, and they've been able to do this with COTS stuff since the 1990s, at the very least, for a price of around $2M – https://en.wikipedia.org/wiki/Project_Bacchus . And those people don't need a LLM to tell them what to do. In fact, they're the exact people who'll have access to unrestricted versions of these LLMs.
And from a security perspective, I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm to the effort of finding people who could be planning such a thing than it helps. It takes more resources to go through the mass of false negatives that have now been created as matter of policy.
These experiments have been run. The fictional scenario of someone learning how bioweapons work and conjuring up a plague isn't real and it hurts humanity as a whole to impede the sciences over it.
Because what someone can flail around in / do is learn about immunology / try to "cure cancer" with a LLM and hopefully get started on a long career in medicine. Or, a discovery that matters.
Because in those cases, if and when they do end up at a lab, screwing up doesn't mean death. Just tons of wasted time (and money). And they will fail / screw up. Just look at literally every undergrad in any lab and the expensive messes they create.
-
And last, but not least, yes. Teenage hackers have been a meme for decades.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#106Earlier quoted context omitted.
The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.
For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#107Earlier quoted context omitted.
Things Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"
[loads up most intelligent AI ever created] “Rabbit sex, how?”
Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently?
I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :)
I like to ask dumb questions. It's fun. I encourage it.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#108Earlier quoted context omitted.
[loads up most intelligent AI ever created] “Rabbit sex, how?”
I mean, yes. Why? Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently? I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :) I like to ask dumb questions. It's fun. I encourage it.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#109Earlier quoted context omitted.
The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519... Fable understood it as something along the lin…
> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
Yeah, Fable is Anthropic's Cybertruck.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#110Earlier quoted context omitted.
Things Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"
[loads up most intelligent AI ever created] “Rabbit sex, how?”