Earlier quoted context omitted.
Have you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?
This is like telling someone who wants a motorcycle that they should get better at running instead.
Claude Sonnet 5
231–240 of 822 posts
Re: Claude Sonnet 5
#232Earlier quoted context omitted.
I think you misunderstood what their vision is, or rather what their possible futures are. They are many steps ahead of almost everyone, both in wargaming possibilities and the actual realized path. What doesn’t make sense to you may be the only safe option for them.
> What doesn’t make sense to you may be the only safe option for them thats true because their point of view makes no sense for us. dario is all in on lesswrong machine god theory and really believes they need to create a super intelligence before anyone else. that means doing as much as possible to slow down others progress and accelerate your own. but the fact that they believe its the only option doesnt make it tr…
Re: Claude Sonnet 5
#233Anthropic's run on the model and product side of things is highly impressive. They got Sam A. punching the air consistently, which is well-deserved and self-inflicted above all.
Wdym? They've been knocking it out of the park on marketing, but Claude Code is still a meme, and Opus is getting trashed by GPT5.5 meanwhile you can't even use their "dominant" model, and anecdotal reports from when people could use Fable, when they weren't getting silently poisoned, was that it was only marginally better than GPT 5.5 in terms of SWE smarts, mostly being better in terms of pleasantness to interact w…
Claude Code generates more revenue than OpenAI...It appears to be a nice meme.
Re: Claude Sonnet 5
#234Earlier quoted context omitted.
Because it’s a massive improvement over the previous model, and cheaper? You are reading too much into the graph and ignoring the threshold of usefulness for real world tasks. By that logic Sonnet 4.5 would have never been worth using.
Am i missing something? Because your making my point. Its only worth it compared to Opus 4.8, if the tasks your running requires Opus 4.8 low (or non-existing lower). For the rest the gap in pricing vs efficiency is so small, that there is no point in using Sonnet. I am looking at their own cost comparisons vs efficiency...
I use Haiku a lot for agent workflows, if I can get better output at similar prices, Sonnet 5 will replace it completely.
Re: Claude Sonnet 5
#235I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
Re: Claude Sonnet 5
#236Earlier quoted context omitted.
Wdym? They've been knocking it out of the park on marketing, but Claude Code is still a meme, and Opus is getting trashed by GPT5.5 meanwhile you can't even use their "dominant" model, and anecdotal reports from when people could use Fable, when they weren't getting silently poisoned, was that it was only marginally better than GPT 5.5 in terms of SWE smarts, mostly being better in terms of pleasantness to interact w…
> Claude Code is still a meme Claude Code generates more revenue than OpenAI...It appears to be a nice meme.
Re: Claude Sonnet 5
#237- Do the ever increasing scores on the mean we will soon have models that approach 100%? And what would that even mean? That there is no more room for improvement?
- Would Anthropic (or any other model vendor for that matter) ever release a newer model that scores lower? If not, does that mean they keep tweaking a new model they want to release until it shows an improvement of the prior model?
- Would it be more useful to move toward a comparative rather than absolute ranking?
Re: Claude Sonnet 5
#238And yet, the $2-$5 section is the widest, even though it only contains a single point.
I can't even say if this is making the product look better or not, but it sure is weird. Maybe Claude just hallucinated those splits xD
Re: Claude Sonnet 5
#239Earlier quoted context omitted.
Yeah, there's a real opportunity for one of these companies to invest time in a model that's tuned for, to use your term, agent-assisted developement. Trouble is, everyone inside their buildings seems to believe that no one will be working like that in a year or two.
Whether they believe it or not is immaterial. It is the end-goal they want to achieve, because then they own the means of production entirely.
Re: Claude Sonnet 5
#240Earlier quoted context omitted.
This is why Fable was so good. It followed instructions and it was in no way lazy.
People keep making comments about fable like this? You could only use it for what like a week? How is that at all enough time to evaluate? Opus 4.6 didnt suffer from this problems for a hot minute and then when newer models were released it got worse. I think they change a ton behind the scenes and allocate compute however they want, so the model you use today may behave much differently than how it behaved yesterday