Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

791–800 of 1001 posts

Re: Claude Sonnet 4.6

#791

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

I'm curious why and how models like these give one answer for one person and a completely different answer for someone else. One reason can be memory maybe? Past conversations that tell the model "Think this way for this user"

Re: Claude Sonnet 4.6

#792
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

> And sadly everyone has the same ideas

I'm not sure that's true. If LLMs can help researchers implement (not find) new ideas faster, they effectively accelerate the progress of research.

Like many other technologies, LLMs will fail in areas and succeed in others. I agree with your take regarding business ideas, but the story could be different for scientific discovery.

Re: Claude Sonnet 4.6

#793

Earlier quoted context omitted.

That last part is a real one though, mine tried to debug a Dockerfile by poking around my local environment outside of Docker today.

Claude's willingness to poke outside of its present directory can definitely be a little worrying. Just the other day, it started trying to access my jails after I specifically told it not to.

On a Mac, I use built-in sandboxing to jail Claude (and every other agent) to $CWD so it doesn’t read/write anything it shouldn’t, doesn’t leak env, etc. This is done by dynamically generating access policies and I open sourced this at https://agent-safehouse.dev

Re: Claude Sonnet 4.6

#794

Earlier quoted context omitted.

I guess I'm getting the dumb one too. I just got this response: > Walk — it's only 50 meters, which is less than a minute on foot. Driving that distance to a car wash would also be a bit counterproductive, since you'd just be getting the car dirty again on the way there (even if only slightly). Lace up and stroll over!

Sonnet 4.6 gives me the fairly bizarre: > Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — and at that distance, walking takes maybe 30–45 seconds. You can simply pull the car out, walk it over (or push it if it's that close), or drive it the short distance once you're ready to wash it. Either way, no need to "drive to the car wash" in the traditional sense. I struggle…

I got almost the same reply, including the "push it" nonsense:

> Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — and the walk will take you less than a minute. You can simply pull the car out and push or walk it over, or drive it the short distance once you're ready to wash it. Either way, no need to "drive" in any meaningful sense for just 50 meters.

Re: Claude Sonnet 4.6

#795
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

I don’t think we are running out of work to do… there seems to be an endless amount of work to be done. And most of it comes from human needs and desires.

Re: Claude Sonnet 4.6

#796
post #360

I don't really understand why they would release something "worse" than Opus 4.6. If it's comparable, then what is the reason to even use Opus 4.6? Sure, it's cheaper, but if so, then just make Opus 4.6 cheaper?

>> Sure, it's cheaper, but if so, then just make Opus 4.6 cheaper?

That makes no sense. People are willing to pay for Opus 4.6 so why would Anthropic make it cheaper exactly?

Re: Claude Sonnet 4.6

#797
post #716

Honest question: why would anyone use Opus instead of this? I’m doing web development, the whole shebang, and I don’t think I need Opus right now. I know it’s supposed to be smarter, but a 2%–5% improvement doesn’t seem meaningful, especially when it costs more than double and has only a portion of the context window. Am I getting this wrong? I would seriously appreciate any clarification here.

The 2-5% margin makes a much bigger difference when it comes to complex problems.

Re: Claude Sonnet 4.6

#798
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

If one person can do the job of three, then you can keep output the same and reduce headcount, or maintain headcount and improve output etc.

Anecdotally it seems demand for software >> supply of software. So in engineering, I think we’ll see way more software. That’s what happened in the Industrial Revolution. Far more products, multiple orders of magnitude more, were produced.

The Industrial Revolution was deeply disruptive to labour, even whilst creating huge wealth and jobs. Retraining is the real problem. That’s what we will see in software. If you can’t architect and think well, you’ll struggle. Being able to write boiler plate and repetitive low level code is a thing of the past. But there are jobs - you’re going to have to work hard to land them.

Now, if AGI or superintelligence somehow renders all humans obsolete, that is a very different problem but that is also the end of capitalism so will be down to governments to address.

Re: Claude Sonnet 4.6

#799
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

This is just a theory of mine, but the fact that people don't see LLMs as something that will grow the pie and increase their output leading to prosperity for all just means that real economic growth has stagnated.

From all my interactions with C-level people as an engineer, what I learned from their mindset is their primary focus is growing their business - market entry, bringing out new products, new revenue streams.

As an engineer I really love optimizing out current infra, bringing out tools and improved workflows, which many of my colleagues have considered a godsend, but it seems from a C-level perspective, it's just a minor nice-to-have.

While I don't necessarily agree with their world-view, some part of it is undeniable - you can easily build an IT company with very high margins - say 3x revenue/expense ratio, in this case growing the profit is a much more lucrative way of growing the company.

Re: Claude Sonnet 4.6

#800

Earlier quoted context omitted.

We said the same thing when 3D printing came out. Any sort of cool tech, we think everybody’s going to do it. Most people are not capable of doing it. in college everybody was going to be an engineer and then they drop out after the first intro to physics or calculus class. A bunch of my non tech friends were vibe coding some tools with replit and lovable and I looked at their stuff and yeah it was neat but it wasn't…

I don't think claude code is like 3d printing. The difference is that 3D printing still requires someone, somewhere to do the mechanical design work. It democratises printing but it doesn't democratise invention . I can't use words to ask a 3d printer to make something. You can't really do that with claude code yet either. But every few months it gets better at this. The question is: How good will claude get at turni…

The design work remains.

I’m not a fan of analogies, but here goes: Apple don’t make iPhones. But they employ an enormous number of people working on iPhone hardware, which they do not make.

If you think AI can replace everyone at Apple, then I think you’re arguing for AGI/superintelligence, and that’s the end of capitalism. So far we don’t have that.

Post reply on HN