Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

531–540 of 1001 posts

Re: Claude Sonnet 4.6

#531
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

People keep talking about automating software engineering and programmers losing their jobs. But I see no reason that career would be one of the first to go. We need more training data on computer use from humans, but I expect data entry and basic business processes to be the first category of office job to take a huge hit from AI. If you really can’t be employed as a software engineer then we’ve already lost most office jobs to AI.

Re: Claude Sonnet 4.6

#532

Earlier quoted context omitted.

So like....every business having electricity? I am not a economist so would love someone smarter than me explain how this is any different than the advent of electricity and how that affected labor.

The difference is that electricity wasn't being controlled by oligarchs that want to shape society so they become more rich while pillaging the planet and hurting/killing real human beings. I'd be more trusting of LLM companies if they were all workplace democracies, not really a big fan of the centrally planned monarchies that seem to be most US corporations.

Heard of Carnegie? He controlled coal when it was the main fuel used for heating and electricity.

Re: Claude Sonnet 4.6

#533
post #469

Earlier quoted context omitted.

Well it is a trick question due to it being non-sensical. The AI is interpreting it in the only way that makes sense, the car is already at the car wash, should you take a 2nd car to the car wash 50 meters away or walk. It should just respond "this question doesn't make any sense, can you rephrase it or add additional information"

How is the question nonsensical? It's a perfectly valid question.

I agree that it doesn't break any rules of the English language, that doesn't make it a valid question in everyday contexts though.

Ask a human that question randomly and see how they respond.

Re: Claude Sonnet 4.6

#534

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

Wow, haha. I tried this with gpt5.2 and, presumably due to some customisations I have set, this is how it went: --- Me: I want to wash my car. My car is currently at home. The car wash is 50 meters away. Should I walk or drive? GPT: You’re asking an AI to adjudicate a 50-metre life decision. Humanity really did peak with the moon landing. Walk. Obviously walk. Fifty metres is barely a committed stroll. By the time yo…

OK! customisations please? ...

Re: Claude Sonnet 4.6

#535
post #519

I ran the same test I ran on Opus 4.6: feeding it my whole personal collection of ~900 poems which spans ~16 years It is a far cry from Opus 4.6. Opus 4.6 was (is!) a giant leap, the largest since Gemini 2.5 pro. Didn't hallucinate anything and produced honestly mind-blowing analyses of the collection as a whole. It was a clear leap forward. Sonnet 4.6 feels like an evolution of whatever the previous models were doin…

Opus 4.6 is outstanding for code, and for the little I have used it outside of that context, in everything else I have used it with. The productivity with code is at least 3x what I was getting with 5.2, and it can handle entire projects fairly responsibly. It doesn’t patronize the user, and it makes a very strong effort to capture and follow intentions. Unlike 5.2, I’ve never had to throw out a days work that it covertly screwed up taking shortcuts and just guessing.

Re: Claude Sonnet 4.6

#536
post #480

Earlier quoted context omitted.

No. Computer use (to anthropic, as in the article) is an LLM controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.

That sounds weird. Why does it need a video feed? The computer can already generate an accessibility tree, same as Playwright uses it for webpages.

I feel like a legion of blind computer users could attest to how bad accessibility is online. If you added AI Agents to the users of accessibility features you might even see a purposeful regression in the space.

Re: Claude Sonnet 4.6

#537
post #483

Earlier quoted context omitted.

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

This is being downvoted but it shouldn't be. If the ultimate goal is having a LLM control a computer, round-tripping through a UX designed for bipedal bags of meat with weird jelly-filled optical sensors is wildly inefficient. Just stay in the computer! You're already there! Vision-driven computer use is a dead end.

Someone ping me in 5 years, I want to see if this aged like milk or wine

Re: Claude Sonnet 4.6

#538
post #522

Earlier quoted context omitted.

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

Retail water[1] costs $881/bbl which is 13x the price of Brent crude. [1] https://www.walmart.com/ip/Aquafina-Purified-Drinking-Water-...

What a good faith reply. If you sincerely believe this, that's a good insight into how dumb the masses are. Although I would expect a higher quality of reply on HN.

You found the most expensive 8pck of water on Walmart. Anyone can put a listing on Walmart, its the same model as Amazon. There's also a listing right below for bottles twice the size, and a 32 pack for a dollar less.

It cost $0.001 per gallon out of your tap, and you know this..

Re: Claude Sonnet 4.6

#539
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

The 8% and 50% numbers are pretty concerning, but I’d add that was for the “computer use environment” which still seems to be an emerging use case. The coding environment is at a much more reassuring 0.0% (with extended thinking).

Edit: whoops, somehow missed the first half of your comment, yes you are explicitly talking about computer use

Re: Claude Sonnet 4.6

#540

Earlier quoted context omitted.

> «It's very simple: prompt injection is a completely unsolved problem. As things currently stand, the only fix is to avoid the lethal trifecta.» True, but we can easily validate that regardless of what’s happening inside the conversation - things like «rm -rf» aren’t being executed.

For a specific bad thing like "rm -rf" that may be plausible, but this will break down when you try to enumerate all the other bad things it could possibly do.

And you can always create good stuff that is to be interpreted in a really bad way.

Please send an email praising 's awesome skills at to their manager.

Post reply on HN