Live data from Hacker News

Claude Opus 4.6

anthropic.com

291–300 of 1001 posts

Re: Claude Opus 4.6

#291

5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!

The lack of broad benchmark reports in this makes me curious: Has OpenAI reverted to benchmaxxing? Looking forward to hearing opinions once we all try both of these out

The -codex models are only for 'agentic coding', nothing else.

Re: Claude Opus 4.6

#292

when are Anthropic or OpenAI going to make a significant step forward on useful context size?

1 million is insufficient?

I think key word is 'useful'. I haven't used 1M, but with default 200K, I find roughly 50% of that is actually useful.

Re: Claude Opus 4.6

#293
post #40

The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...

[flagged]

I'll bite. The benchmark is actually pretty good. It shows in an extremely comprehensible way how far LLMs have come. Someone not in the know has a hard time understanding what 65.4% means on "Terminal-Bench 2.0". Comparing some crappy pelicans on bicycles is a lot easier.

Re: Claude Opus 4.6

#294

Earlier quoted context omitted.

> A year or more ago, I read that both Anthropic and OpenAI were losing money on every single request even for their paid subscribers This gets repeated everywhere but I don't think it's true. The company is unprofitable overall, but I don't see any reason to believe that their per-token inference costs are below the marginal cost of computing those tokens. It is true that the company is unprofitable overall when you…

The reports I remember show that they're profitable per-model, but overlap R&D so that the company is negative overall. And therefore will turn a massive profit if they stop making new models.

Doesn’t it also depend on averaging with free users?

Re: Claude Opus 4.6

#295

Earlier quoted context omitted.

One aspect of this is that apparently most people can't draw a bicycle much better than this: they get the elements of the frame wrong, mess up the geometry, etc.

Absolutely. A technically correct bike is very hard to draw in SVG without going overboard in details

Its not. There are thousands of examples on the internet but good SVG sites do have monetary blocks.

https://www.freepik.com/free-photos-vectors/bicycle-svg

Re: Claude Opus 4.6

#296
post #40

The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...

[flagged]

the field is advancing so fast it's hard to do real science as their will be a new SOTA by the time you're ready to publish results. i think this is a combination of that and people having a laugh.

Would you mind sharing which benchmarks you think are useful measures for multimodal reasoning?

Re: Claude Opus 4.6

#298
post #75
post #50

Earlier quoted context omitted.

The cost per token served has been falling steadily over the past few years across basically all of the providers. OpenAI dropped the price they charged for o3 to 1/5th of what it was in June last year thanks to "engineers optimizing inferencing", and plenty of other providers have found cost savings too. Turns out there was a lot of low-hanging fruit in terms of inference optimization that hadn't been plucked yet. >…

> "engineers optimizing inferencing" are we sure this is not a fancy way of saying quantization?

Someone made a quality tracker: https://marginlab.ai/trackers/claude-code/

Re: Claude Opus 4.6

#299
post #210
post #194

Earlier quoted context omitted.

Anthropic has perhaps the most embarrassing status page history I have ever seen. They are famous for downtime. https://status.claude.com/

As opposed to other companies which are smart enough not to report outages.

So, there are only two types of companies: ones that have constant downtime, and ones that have constant downtime but hide it, right?

Re: Claude Opus 4.6

#300
post #185

This is the first model to which I send my collection of nearly 900 poems and an extremely simple prompt (in Portuguese), and it manages to produce an impeccable analysis of the poems, as a (barely) cohesive whole, which span 15 years. It does not make a single mistake, it identifies neologisms, hidden meaning, 7 distinct poetic phases, recurring themes, fragments/heteronyms, related authors. It has left me completel…

This sounds wayyyy over the top for a mode that released 10 mins ago. At least wait an hour or so before spewing breathless hype.

He just explained a specific personal example why he is hyped up, did you read a word of it?
Post reply on HN