Live data from Hacker News

Claude Opus 4.6

anthropic.com

601–610 of 1001 posts

Re: Claude Opus 4.6

#601

Earlier quoted context omitted.

> From my view the community here is just mean reverting to any other tech internet comments section. As someone deeply familiar with tech internet comments sections, I would have to disagree with you here. Dang et al have done a pretty stellar job of preventing HN from devolving like most other forums do. Sure you have your complainers and zealots, but I still find surprising insights here there I don't find anywher…

Mean reverting is a time based process I fear. I think dang, tomhow, et al are fantastic mods but they can ultimately only stem the inevitable. HN may be a few years behind the other open tech forums but it's a time shifted version of the same process with the same destination, just IMO. I've stopped engaging much here because I need a higher ROI from my time. Endless squabbling, flamewars, and jokes just isn't enoug…

> FWIW I've loved reading your comments over the years and think you've done a great job of living up to what I've loved in this community.

You're too kind! I do appreciate that.

I actually checked out your site on your profile, that's some pretty interesting data! Curious if you've considered updating it?

Re: Claude Opus 4.6

#602

Impressive results, but I keep coming back to a question: are there modes of thinking that fundamentally require something other than what current LLM architectures do? Take critical thinking — genuinely questioning your own assumptions, noticing when a framing is wrong, deciding that the obvious approach to a problem is a dead end. Or creativity — not recombination of known patterns, but the kind of leap where you r…

They're incredibly bad on philosophy, complete lack of understanding

Re: Claude Opus 4.6

#603
post #482

Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…

To be fair, I don't think "Slugulus Eructo" (the name) is actually in the books. This is what's in my copy: > The smug look on Malfoy’s face flickered. > “No one asked your opinion, you filthy little Mudblood,” he spat. > Harry knew at once that Malfoy had said something really bad because there was an instant uproar at his words. Flint had to dive in front of Malfoy to stop Fred and George jumping on him, Alicia shr…

I have a vague recollection that it might come up named as such in Half-Blood Prince, written in Snape's old potions textbook?

In support of that hypothesis, the Fandom site lists it as “mentioned” in Half-Blood Prince, but it says nothing else and I'm traveling and don't have a copy to check, so not sure.

Re: Claude Opus 4.6

#604
post #482

Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…

That doesn't seem a super useful test for a model that's optimized for programming?

Re: Claude Opus 4.6

#605

Does anyone with more insight into the AI/LLM industry happen to know if the cost to run them in normal user-workflows is falling? The reason I'm asking is because "agent teams" while a cool concept, it largely constrained by the economics of running multiple LLM agents (i.e. plans/API calls that make this practical at scale are expensive). A year or more ago, I read that both Anthropic and OpenAI were losing money o…

These are intro prices. This is all straight out of the playbook. Get everyone hooked on your product by being cheap and generous. Raise the price to backpay what you gave away plus cover current expenses and profits. In no way shape or form should people think these $20/mo plans are going to be the norm. From OpenAI's marketing plan, and a general 5-10 year ROI horizon for AI investment, we should expect AI use to c…

The models in 5-10 years are going to be unimaginably good. $100/month will be a bargain for knowledge workers, if they survive.

Re: Claude Opus 4.6

#606
post #562

Earlier quoted context omitted.

It didn't use web search. But for sure it has some internal knowledge already. It's not a perfect needle in the hay stack problem but gemini flash was much worse when I tested it last time.

If you want to really test this, search/replace the names with your own random ones and see if it lists those. Otherwise, LLMs have most of the books memorised anyway: https://arstechnica.com/features/2025/06/study-metas-llama-3...

Couldn't you just ask the LLM which 50 (or 49) spells appear in the first four Harry Potter books without the data for comparison?

Re: Claude Opus 4.6

#607
post #482

Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…

I often wonder how much of the Harry Potter books were used in the training. How long before some LLM is able to regurgitate full HP books without access to the internet?

Re: Claude Opus 4.6

#608

Earlier quoted context omitted.

Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?

We don't vary our model quality with time of day or load (beyond negligible non-determinism). It's the same weights all day long with no quantization or other gimmicks. They can get slower under heavy load, though. (I'm from OpenAI.)

Has this always been the case?

Re: Claude Opus 4.6

#609

Earlier quoted context omitted.

Why? I use it for all and love it. That doesn't mean you have to, but I'm curious why you think it's behind in the personal assistant game.

It's hard to say. Maybe it has to do with the way Claude responds or the lack of "thinking" compared to other models. I personally love Claude and it's my only subscription right now, but it just feels weird compared to the others as a personal assistant.

Oh, I always use opus 4.5 thinking mode. Maybe that's the diff.

Re: Claude Opus 4.6

#610
Google already won the AI race. It's very silly to try and make AGI by hyperfocusing on outdated programming paradigms. You NEED multimodal to do anything remotely interesting with these systems.
Post reply on HN