Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

811–820 of 1001 posts

Re: Claude Sonnet 4.6

#811

I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.

[dead]

Re: Claude Sonnet 4.6

#812
post #525

It doesn't do so well on my stupid benchmarks, lol: https://aibenchy.com Gets wrong some tests. It does answer correctly, BUT it doesn't respect the request to respond ONLY with the answer, it keeps adding extra explanations at the end.

Looks like you're mixing up two things when testing: the correct answer and format following. If you want both, why not use https://platform.claude.com/docs/en/build-with-claude/struct... ? If you don't care about the structure, why penalise the correct answers? In realistic usage people don't say "I really care about the format a lot... but not enough to guarantee it".

Re: Claude Sonnet 4.6

#813
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

> Its also worth noting that if you can create a business with an LLM, so can everyone else.

False. Anyone can learn about index ETFs and still yolo into 3DTE options and promptly get variation margined out of existence.

Discipline and contextual reasoning in humans is not dependent on the tools they are using, and I think the take is completely and definitively wrong.

Re: Claude Sonnet 4.6

#814

Earlier quoted context omitted.

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

> Its also worth noting that if you can create a business with an LLM, so can everyone else. And sadly everyone has the same ideas Yeah, this is quite thought provoking. If computer code written by LLMs is a commodity, what new businesses does that enable? What can we do cheaply we couldn't do before? One obvious answer is we can make a lot more custom stuff . Like, why buy Windows and Office when I can just ask clau…

This whole comment thread here is really echoing and adding to some thoughts ive had lately on the shift from considering LLMs replacing engineering to make software (much of which is about integration, longevity and customization of a general system), vs LLMs replacing buying software.

If most software is just used by me to do a specific task, then being able to make software for me to do that task will become the norm. Following that thought, we are going to see a drastic reduction in SASS solutions, as many people who were buying a flexible-toolbox for one usecase to use occasionally, just get an llm to make them the script/software to do that task as and when they need it, without any concern for things like security, longevity, ease of use by others (for better or for worse).

I guess what im circling around is that if we define engineering as building the complex tools that have to interact with many other systems, persist, be generally useful and understandable to many people, and we consider that many people actually dont need that complexity for their use of the system, the complexity arises from it needing to serve its purpose at huge scale over time. then maybe there will be less need for enginners, but perhaps first and foremost because the problems that engineering is required to solve are much less if much more focused and bespoke solutions to peoples problems are available on demand.

As an engineer i have often felt threatened by LLMs and agents of late, but i find that if i reframe it from Agents replacing me, to Agents causing the type of problems that are even valuable to solve to shift, it feels less threatening for some reason. Ill have to mull more.

Re: Claude Sonnet 4.6

#816

Earlier quoted context omitted.

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

No. Computer use (to anthropic, as in the article) is an LLM controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.

Even simpler it just takes screenshots (or at least that's what it was doing last time I used it)

Re: Claude Sonnet 4.6

#817
post #519

I ran the same test I ran on Opus 4.6: feeding it my whole personal collection of ~900 poems which spans ~16 years It is a far cry from Opus 4.6. Opus 4.6 was (is!) a giant leap, the largest since Gemini 2.5 pro. Didn't hallucinate anything and produced honestly mind-blowing analyses of the collection as a whole. It was a clear leap forward. Sonnet 4.6 feels like an evolution of whatever the previous models were doin…

I'm curious how this would compare with codex 5.3. I've heard Codex actually is pretty good but Opus 4.6 has become synonymous with AI coding because all the big names praise it. I haven't compared them against each other though so can't really draw a conclusion.

There are no universals. You have to try it on your particular codebase and see what works for you.

For me, OpenAI is ahead in intelligence, and Anthropic is ahead in alignment. I use both but for different tasks.

Given the pace of change, intuition is somewhat of a liability: what's true today may not be true tomorrow. You have to constantly keep an open mind and try new things.

Listening to influencers is a waste of time.

Re: Claude Sonnet 4.6

#818
post #519

I ran the same test I ran on Opus 4.6: feeding it my whole personal collection of ~900 poems which spans ~16 years It is a far cry from Opus 4.6. Opus 4.6 was (is!) a giant leap, the largest since Gemini 2.5 pro. Didn't hallucinate anything and produced honestly mind-blowing analyses of the collection as a whole. It was a clear leap forward. Sonnet 4.6 feels like an evolution of whatever the previous models were doin…

Thanks for testing and sharing your results.

Re: Claude Sonnet 4.6

#819

Earlier quoted context omitted.

This is a bit of a tangent, but it highlights exactly what people miss when talking about China taking over our industries. Right now, China has about 140 different car brands, roughly 100 of which are domestic. Compare that to Europe, where we have about 50 brands competing, or the US, which is essentially a walled garden with fewer than 40. That level of internal fierce competition is a massive reason why they are…

It's the low cost of labor in addition to lack of environmental regulation that made China a success story. I'm sure the competition helps too but it's not main driver

oh, then explain to me how both China is leading in both robotics and AI. if it is because of "low cost of labor in addition to lack of environmental regulation", you'd be seeing countries like india beating the US and EU.

Re: Claude Sonnet 4.6

#820
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

> everyone has access to the same models and basic thought processes

Why haven't Warners acquired Netflix then, but the other way around? Even though they had access to the same labor market, a human LLM replacement?

I think real economics is a little more complex than the "basic economics" referenced in your reply.

This does not negate the possibility that enterprises will double down on replacing everyone with AI, though. But it does negate the reasoning behind the claim and the predictions made.

Post reply on HN