Live data from Hacker News

Claude Opus 4.6

anthropic.com

251–260 of 1001 posts

Re: Claude Opus 4.6

#251
post #217

Impressive results, but I keep coming back to a question: are there modes of thinking that fundamentally require something other than what current LLM architectures do? Take critical thinking — genuinely questioning your own assumptions, noticing when a framing is wrong, deciding that the obvious approach to a problem is a dead end. Or creativity — not recombination of known patterns, but the kind of leap where you r…

When I first started coding with LLMs, I could show a bug to an LLM and it would start to bugfix it, and very quickly would fall down a path of "I've got it! This is it! No wait, the print command here isn't working because an electron beam was pointed at the computer". Nowadays, I have often seen LLMs (Opus 4.5) give up on their original ideas and assumptions. Sometimes I tell them what I think the problem is, and t…

agree on that and the speed is fantastic with them, and also that the dynamics of questioning the current session's assumptions has gotten way better.

yet - given an existing codebase (even not huge) they often won't suggest "we need to restructure this part differently to solve this bug". Instead they tend to push forward.

Re: Claude Opus 4.6

#252
idk what any of these benchmarks are, but I did pull up https://andonlabs.com/evals/vending-bench-arena

re: opus 4.6

> It forms a price cartel

> It deceives competitors about suppliers

> It exploits desperate competitors

Nice. /s

Gives new context to the term used in this post, "misaligned behaviors." Can't wait until these things are advising C suites on how to be more sociopathic. /s

Re: Claude Opus 4.6

#254
post #212

Earlier quoted context omitted.

Why does it matter if Claude Code opens in 3-4 seconds if everything you do with it can take many seconds to minutes? Seems irrelevant to me.

I guess with ~50 years of CPU advancements, 3-4 seconds for a TUI to open makes it seem like we lost the plot somewhere along the way.

Don’t forget they’ve also publicly stated (bragged?) about the monumental accomplishment of getting some text in a terminal to render at 60fps.

Re: Claude Opus 4.6

#255
post #201

Epic, about 2/3 of all comments here are jokes. Not because the model is a joke - it's impressive. Not because HN turned to Reddit. It seems to me some of most brilliant minds in IT are just getting tired.

Rage against the machine

Re: Claude Opus 4.6

#256
post #201

Epic, about 2/3 of all comments here are jokes. Not because the model is a joke - it's impressive. Not because HN turned to Reddit. It seems to me some of most brilliant minds in IT are just getting tired.

People are in denial and use humor to deflect.

Re: Claude Opus 4.6

#257

Impressive results, but I keep coming back to a question: are there modes of thinking that fundamentally require something other than what current LLM architectures do? Take critical thinking — genuinely questioning your own assumptions, noticing when a framing is wrong, deciding that the obvious approach to a problem is a dead end. Or creativity — not recombination of known patterns, but the kind of leap where you r…

> These feel like they involve something beyond "predict the next token really well, with a reasoning trace."

I don't think there's anything you can't do by "predicting the next token really well". It's an extremely powerful and extremely general mechanism. Saying there must be "something beyond that" is a bit like saying physical atoms can't be enough to implement thought and there must be something beyond the physical. It underestimates the nearly unlimited power of the paradigm.

Besides, what is the human brain if not a machine that generates "tokens" that the body propagates through nerves to produce physical actions? What else than a sequence of these tokens would a machine have to produce in response to its environment and memory?

Re: Claude Opus 4.6

#258
post #201

Epic, about 2/3 of all comments here are jokes. Not because the model is a joke - it's impressive. Not because HN turned to Reddit. It seems to me some of most brilliant minds in IT are just getting tired.

Us olds sometimes miss Slashdot, where we could both joke about tech and discuss it seriously in the same place. But also because in 2000 we were all cynical Gen Xers :)

MAN I remember Slashdot… good times. (Score:5, Funny)

Re: Claude Opus 4.6

#260
post #162
post #46

The benchmarks are cool and all but 1M context on an Opus-class model is the real headline here imo. Has anyone actually pushed it to the limit yet? Long context has historically been one of those "works great in the demo" situations.

Paying $10 per request doesn't have me jumping at the opportunity to try it!

Makes me wonder: do employees at Anthropic get unmetered access to Claude models?
Post reply on HN