How do people keep track of all these versions and releases of all these models and their pros/cons? Seems like a fulltime hobby to me. I'd rather just improve my own skills with all that time and energy
Claude Sonnet 4.6
501–510 of 1001 posts
Re: Claude Sonnet 4.6
#502It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
Re: Claude Sonnet 4.6
#503Re: Claude Sonnet 4.6
#504Earlier quoted context omitted.
Why is it wild that a LLM is as capable as a previously released LLM?
Opus is supposed to be the expensive-but-quality one, while Sonnet is the cheaper one. So if you don't want to pay the significant premium for Opus, it seems like you can just wait a few weeks till Sonnet catches up
I'm even surprised people pay more money for some models than others.
Re: Claude Sonnet 4.6
#505Earlier quoted context omitted.
Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…
Yeah, but a Stratocaster guitar is available to everybody too, but not everybody’s an Eric Clapton
Re: Claude Sonnet 4.6
#506I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…
Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…
Automation should be, obviously, a good thing, because more is produced with less labor. What it says of ourselves and our politics that so many people (me included) are afraid of it?
In a sane world, we would realize that, in a post-work world, the owner of the robots have all the power, so the robots should be owned in common. The solution is political.
Re: Claude Sonnet 4.6
#507Earlier quoted context omitted.
Remarkable, since the goal is clearly stated and the language isn’t tricky.
Well it is a trick question due to it being non-sensical. The AI is interpreting it in the only way that makes sense, the car is already at the car wash, should you take a 2nd car to the car wash 50 meters away or walk. It should just respond "this question doesn't make any sense, can you rephrase it or add additional information"
Re: Claude Sonnet 4.6
#508Earlier quoted context omitted.
Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…
So like....every business having electricity? I am not a economist so would love someone smarter than me explain how this is any different than the advent of electricity and how that affected labor.
I'd be more trusting of LLM companies if they were all workplace democracies, not really a big fan of the centrally planned monarchies that seem to be most US corporations.
Re: Claude Sonnet 4.6
#509Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…
https://claude.ai/share/32de37c4-46f2-4763-a2e1-8de7ecbcf0b4
Re: Claude Sonnet 4.6
#510Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…
My answer was (for which it did zero thinking and answered near-instantaneously): "Drive. You're going there to use water and machinery that require the car to be present. The question answers itself." I tried it 3 more times with extended thinking explicitly off: "Drive. You're going to a car wash." "Drive. You're washing the car, not yourself." "Drive. You're washing the car — it needs to be there." Guess they're s…