Live data from Hacker News

Claude Opus 4.6

anthropic.com

501–510 of 1001 posts

Re: Claude Opus 4.6

#501
post #423

Earlier quoted context omitted.

Interesting. Everyone in my circle said the opposite.

It probably depends on programming language and expectations.

This is mostly Python/TS for me... what Jonathan Blow would probably call not "real programming" but it pays the bills

They can both write fairly good idiomatic code but in my experience opus 4.5 is better at understanding overall project structure etc. without prompting. It just does things correctly first time more often than codex. I still don't trust it obviously but out of all LLMs it's the closest to actually starting to earn my trust

Re: Claude Opus 4.6

#502
post #201

Epic, about 2/3 of all comments here are jokes. Not because the model is a joke - it's impressive. Not because HN turned to Reddit. It seems to me some of most brilliant minds in IT are just getting tired.

Every single day 80% of the frontpage is AI news… Those of us who don't use AI (and there are dozens of us, DOZENS) are just bored I guess.

Marketing something that is meant to replace us to us...

Re: Claude Opus 4.6

#503
post #361
post #231

Earlier quoted context omitted.

It's quite clear that these companies do make money on each marginal token. They've said this directly and analysts agree [1]. It's less clear that the margins are high enough to pay off the up-front cost of training each model. [1] https://epochai.substack.com/p/can-ai-companies-become-profi...

It’s not clear at all because model training upfront costs and how you depreciate them are big unknowns, even for deprecated models. See my last comment for a bit more detail.

By now, model lifetime inference compute is >10x model training compute, for mainstream models. Further amortized by things like base model reuse.

Re: Claude Opus 4.6

#504

From the press release at least it sounds more expensive than Opus 4.5 (more tokens per request and fees for going over 200k context). It also seems misleading to have charts that compare to Sonnet 4.5 and not Opus 4.5 (Edit: It's because Opus 4.5 doesn't have a 1M context window). It's also interesting they list compaction as a capability of the model. I wonder if this means they have RL trained this compaction as o…

On Openrouter it has the same cost per token as 4.5

Re: Claude Opus 4.6

#505

Earlier quoted context omitted.

There's a research paper from the University of Liverpool, published in 2006 where researchers asked people to draw bicycles from memory and how people overestimate their understanding of basic things. It was a very fun and short read. It's called "The science of cycology: Failures to understand how everyday objects work" by Rebecca Lawson. https://link.springer.com/content/pdf/10.3758/bf03195929.pdf

A place I worked at used it as part of an interview question (it wasn't some pass/fail thing to get it 100% correct, and was partly a jumping off point to a different question). This was in a city where nearly everyone uses bicycles as everyday transportation. It was surprising how many supposedly mechanical-focused people who rode a bike everyday, even rode a bike to the interview, would draw a bike that would not w…

I wish I had interviewed there. When I first read that people have a hard time with this I immediately sat down without looking at a reference and drew a bicycle. I could ace your interview.

Re: Claude Opus 4.6

#506
I just tested both codex 5.3 and opus 4.6 and both returned pretty good output, but opus 4.6's limits are way too strict. I am probably going to cancel my Claude subscription for that reason:

What do you want to do?

  1. Stop and wait for limit to reset
   2. Switch to extra usage
   3. Upgrade your plan

 Enter to confirm · Esc to cancel
How come they don't have "Cancel your subscription and uninstall Claude Code"? Codex lasts for way longer without shaking me down for more money off the base $xx/month subscription.

Re: Claude Opus 4.6

#507

Earlier quoted context omitted.

On benchmarks GPT 5.2 was roughly equivalent to Opus 4.5 but most people who've used both for SWE stuff would say that Opus 4.5 is/was noticeably better

There's an extended thinking mode for GPT 5.2 i forget the name of it right at this minute. It's super slow - a 3 minute opus 4.5 prompt is circa 12 minutes to complete in 5.2 on that super extended thinking mode but it is not a close race in terms of results - GPT 5.2 wins by a handy margin in that mode. It's just too slow to be useable interactively though.

Interesting, sounds like I definitely need to give the GPT models another proper go based on this discussion

Re: Claude Opus 4.6

#508
post #83

Earlier quoted context omitted.

Microsoft's products are also extremely successful they're also total garbage

but they have the advantage of already being a big company. Anthropic is new and there's no reason for people to use it

what about if management gives them a reason? You can think of which those can be.

Re: Claude Opus 4.6

#509

I'm still not sure I understand Anthropic's general strategy right now. They are doing these broad marketing programs trying to take on ChatGPT for "normies". And yet their bread and butter is still clearly coding. Meanwhile, Claude's general use cases are... fine. For generic research topics, I find that ChatGPT and Gemini run circles around it: in the depth of research, the type of tasks it can handle, and the qual…

Claude sucks at non English languages. Gemini and ChatGPT are much better. Grok is the worst. I am a native Czech speaker and Claude makes up words and Grok sometimes respond in Russian. So while I love it for coding, it’s unusable for general purpose for me.

Claude code (opus) is very good in Polish.

I sometimes vibe code in polish and it's as good as with English for me. It speaks a natural, native level Polish.

I used opus to translate thousands of strings in my app into polish, Korean, and two Chinese dialects. Polish one is great, and the other are also good according to my customers.

Re: Claude Opus 4.6

#510
post #440

Earlier quoted context omitted.

Claude sucks at non English languages. Gemini and ChatGPT are much better. Grok is the worst. I am a native Czech speaker and Claude makes up words and Grok sometimes respond in Russian. So while I love it for coding, it’s unusable for general purpose for me.

> Grok sometimes respond in Russian Geopolitically speaking this is hilarious.

The voice mode sounded like a Ukrainian trying to speak Czech. I don’t think it means anything.
Post reply on HN