Still mad at them because they decided not to take their users' privacy serious. Would be interested how the new model behaves, but just have a mental lock and can't sign up again.
Claude Opus 4.5
401–410 of 525 posts
Re: Claude Opus 4.5
#402The first chart is straight from "how to lie in charts"..
Re: Claude Opus 4.5
#403Re: Claude Opus 4.5
#404Earlier quoted context omitted.
I once heard a devils advocate say, “if child porn can be fully AI generated and not imply more exploitation of real children, and it’s still banned then it’s about control not harm.” Attack away or downvote my logic.
So how exactly did you train this AI to produce CSAM?
Re: Claude Opus 4.5
#405Earlier quoted context omitted.
There are two possible explanations for this behavior: the model nerf is real, or there's a perceptual/psychological shift. However, benchmarks exist. And I haven't seen any empirical evidence that the performance of a given model version grows worse over time on benchmarks (in general.) Therefore, some combination of two things are true: 1. The nerf is psychologial, not actual. 2. The nerf is real but in a way that…
> 1. The nerf is psychologial, not actual. 2. The nerf is real but in a way that is perceptual to humans, but not benchmarks. They could publish weekly benchmarks. To disprove. They almost certainly have internal benchmarking. The shift is certainly real. It might not be model performance but contextual changes or token performance (tasks take longer even if the model stays the same).
Re: Claude Opus 4.5
#406This is gonna be game-changing for the next 2-4 weeks before they nerf the model. Then for the next 2-3 months people complaining about the degradation will be labeled “skill issue”. Then a sacrificial Anthropic engineer will “discover” a couple obscure bugs that “in some cases” might have lead to less than optimal performance. Still largely a user skill issue though. Then a couple months later they’ll release Opus 4…
There are two possible explanations for this behavior: the model nerf is real, or there's a perceptual/psychological shift. However, benchmarks exist. And I haven't seen any empirical evidence that the performance of a given model version grows worse over time on benchmarks (in general.) Therefore, some combination of two things are true: 1. The nerf is psychologial, not actual. 2. The nerf is real but in a way that…
Re: Claude Opus 4.5
#407This is gonna be game-changing for the next 2-4 weeks before they nerf the model. Then for the next 2-3 months people complaining about the degradation will be labeled “skill issue”. Then a sacrificial Anthropic engineer will “discover” a couple obscure bugs that “in some cases” might have lead to less than optimal performance. Still largely a user skill issue though. Then a couple months later they’ll release Opus 4…
There are two possible explanations for this behavior: the model nerf is real, or there's a perceptual/psychological shift. However, benchmarks exist. And I haven't seen any empirical evidence that the performance of a given model version grows worse over time on benchmarks (in general.) Therefore, some combination of two things are true: 1. The nerf is psychologial, not actual. 2. The nerf is real but in a way that…
I do suspect continued fine tuning lowers quality — stuff they roll out for safety/jailbreak prevention. Those should in theory buildup over time with their fine tune dataset, but each model will have its own flaws that need tuning out.
I do also suspect there’s a bit of mental adjustment that goes in too.
Re: Claude Opus 4.5
#408Earlier quoted context omitted.
https://apps.apple.com/us/app/claude-by-anthropic/id64737536... Has a section for code. You link it to your GitHub, and it will generate code for you when you get on the bus so there's stuff for you to review after you get to the office.
Thanks. Still looking for some kind of total code by phone thing.
Re: Claude Opus 4.5
#409Earlier quoted context omitted.
based on their past usage of "interleaved tool calling" it means that the tool can be used while the model is thinking. https://aws.amazon.com/blogs/opensource/using-strands-agents...
AFAICT, kimi k2 was the first to apply this technique [1]. I wonder if Anthropic came up with it independently or if they trained a model in 5 months after seeing kimi’s performance. 1: https://www.decodingdiscontinuity.com/p/open-source-inflecti...
And the July Kimi K2 release wasn't a thinking model, the model in that article was released less than 20 days ago.
Re: Claude Opus 4.5
#410Earlier quoted context omitted.
No, it's entirely psychological. Users are not reliable model evaluators. It's a lesson the industry will, I'm afraid, have to learn and relearn over and over again.
I'm working on a hard problem recently and have been keeping my "model" setting pegged to "high". Why in the world, if I'm paying the loss leader price for "unlimited" usage of these models, would any of these companies literally respect my preference to have unfettered access to the most expensive inference? Especially when one of the hallmark features of GPT-5 was a fancy router system that decides automatically wh…