Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

731–740 of 1001 posts

Re: Claude Sonnet 4.6

#731
They use the word "Sonnet" 60+ times on that page but never give the casual reader any context of what a "Sonnet model" actually is. Neither does their landing page. You have to scroll all the way to the footer to find a link under the "Models" section. You click it and you finally get the description

"Hybrid reasoning model with superior intelligence for agents, featuring a 1M context window"

You then compare that to Opus Model description

"Hybrid reasoning model that pushes the frontier for coding and AI agents, featuring a 1M context window"

Is the casual person meant to decide if "Superior" is actually less powerful than "Frontier"?

Re: Claude Sonnet 4.6

#732

Earlier quoted context omitted.

Last I checked, the tractor and plow are doing a lot more work than 3 farmers, yet we've got more jobs and grow more food. People will find work to do, whether that means there's tens of thousands of independent contractors, whether that means people migrate into new fields, or whether that means there's tens of multi-trillion dollar companies that would've had 200k engineers each that now only have 50k each and it's…

America has lost over 50% of farms and farmers since 1900. Farming used to be a significant employer, and now it's not. Farming used to be a significant part of the GDP, and now it's not. Farming used to be politically significant... and not its complicated? . If you go to the many small towns in farm country across the United States, I think the last 100 years will look a lot closer to "doom" than "bumps in the road…

Same thing with Walmart and local shops.

On the one hand, it brings a greater selection, at cheaper prices, delivered faster, to communities.

On the other hand, it steamrolls any competing businesses and extracts money that previously circulated locally (to shareholders instead).

Re: Claude Sonnet 4.6

#733

Earlier quoted context omitted.

I actually agree with you, but I have no idea how one can compete in this playing field. The second there are a couple of bad actors in spammarketing, your hands are tied. You really can’t win without playing dirty. I really hate this, not justifying their behaviour, but have no clue how one can do without the other.

Its just law of the jungle all over again. Might makes right. Outcomes over means. Game theory wise there is no solution except to declare (and enforce) spaces where leeching / degrading the environment is punished, and sharing, building, and giving back to the environment is rewarded. Not financially, because it doesn't work that way, usually through social cred or mutual values. But yeah the internet can no longer…

I'm pretty sure this might be a hot take, but I believe we need some sort of a Tech Police.

We have Road Police, Financial Police, Mail Police, Work Safety Police, Military Police...

Re: Claude Sonnet 4.6

#734
We ran some tests at mocha (we have a coding agent with our own harness to build web apps, with a lot of tools and medium length tasks (3min to 10min).

Our notes:

Sonnet 4.6 feels like a fundamentally different model than Sonnet 4.5, it is much closer to the Opus series in terms of agentic behavior and autonomy.

Autonomy - In our zero-shot app building experiments, Sonnet 4.6 ran up to 3-4x longer than Sonnet 4.5 without intervention, producing functional apps on par in terms of quality to the Opus series. Note that subjectively we found Opus 4.5 and 4.6 are better "designers" than Sonnet 4.6; producing more visually appealing apps from the same prompts.

Planning / Task Decomposition - We found Sonnet 4.6 is very good at decomposing tasks and staying on track during long-running trajectories. It's quite good at ensuring all of the requirements of an input prompt are accounted for, whereas we were often forced to goad sonnet 4.5 into decomposing tasks, Sonnet 4.6 does this naturally.

Exploration - In some of our complex "exploration" tasks (e.g. cloning/remixing an existing website), Sonnet 4.6 often performs on par or better than Opus 4.5 and 4.6. It generally takes longer, and takes more tokens, though we believe this is likely a consequence of our tool-calling setup.

Tool-use - Sonnet 4.6 seems eager to use tools; however, we did find that it struggles with our XML-based custom tool use format (perhaps exclusive to the format we use). We did not have a chance to assess with native tool use

Self-verification - Similar to Opus 4.5/4.6, Sonnet 4.6 has a proclivity for verifying it's work.

Prompting - We found Sonnet 4.6 is very sensitive to prompting around thinking, planning, and task decomposition. Our prompt built for sonnet 4.5 has a tendency to push sonnet 4.6 into incredibly long thinking and planning loops. Though we also found it requires significantly less careful and specific instructions for how to approach problems.

How are we thinking about this:

We can't launch this model day 0, it requires more changes to our harness, and we're working on them right now.

But it reminds me a bit of 3.5 to 3.7 --> It's a pretty different model that behaves and responds to instructions in new ways. So it requires more tuning before we can extract its full potential.

Re: Claude Sonnet 4.6

#735

Earlier quoted context omitted.

This is the elephant in the room nobody wants to talk about. AI is dead in the water for the supposed mass labor replacement that will happen unless this is fixed. Summarize some text while I supervise the AI = fine and a useful productivity improvement, but doesn’t replace my job. Replace me with an AI to make autonomous decisions outside in the wild and liability-ridden chaos ensues. No company in their right mind…

There’s a middle road where AI replaces half the juniors or entry level roles, the interns and the bottom rung of the org chart. In marketing, an AI can effortlessly perform basic duties, write email copy, research, etc. Same goes for programming, graphic design, translation, etc. The results will be looked over by a senior member, but it’s already clear that a role with 3 YOE or less could easily be substituted with…

I think you're really overstating things here. Entry level positions are the tier at which replacement of senior positions happen. They don't do a lot, sure, but they are cheap and easily churnable. This is precisely NOT the place companies focus on for cutbacks or downsizing. AI being acceptable at replacing unskilled labor doesn't mean it WILL replace it. It has to make business sense to implement it.

Re: Claude Sonnet 4.6

#736

Earlier quoted context omitted.

We said the same thing when 3D printing came out. Any sort of cool tech, we think everybody’s going to do it. Most people are not capable of doing it. in college everybody was going to be an engineer and then they drop out after the first intro to physics or calculus class. A bunch of my non tech friends were vibe coding some tools with replit and lovable and I looked at their stuff and yeah it was neat but it wasn't…

Its not our current location, but our trajectory that is scary. The walls and plateaus that have been consistently pulled out from "comments of reassurance" have not materialized. If this pace holds for another year and a half, things are going to be very different. And the pipeline is absolutely overflowing with specialized compute coming online by the gigawatt for the foreseeable future. So far the most accurate pr…

There is a distribution of optimism, some people in 2023 were predicting AGI by 2025.

No such thing as trajectory when it comes to mass behavior because it can turn on a dime if people find reason to. Thats what makes civilization so fun.

Re: Claude Sonnet 4.6

#737

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

Is this the new "r's in strawberry"? Are you going (stochastically) parrot this until it's been trained out?

> trained out

No need. Just add one more correction to the system prompt.

It's amusing to see hardcore believers of this tech doing mental gymnastics and attacking people whenever evidence of there being no intelligence in these tools is brought forth. Then the tool is "just" a statistical model, and clearly the user is holding it wrong, doesn't understand how it works, etc.

Re: Claude Sonnet 4.6

#739
The 8% one-shot / 50% unbounded injection numbers from the system card are more honest than most labs publish, and they highlight exactly why you can't evaluate safety with static tests. An attacker doesn't get one shot — they iterate. The right metric isn't "did it resist this prompt" but "how many attempts until it breaks." That's inherently an adversarial, multi-turn evaluation. Single-pass safety benchmarks are measuring the wrong thing for the same reason single-pass capability benchmarks are: real-world performance is sequential and adaptive.
Post reply on HN