Live data from Hacker News

Claude Opus 4.6

anthropic.com

911–920 of 1001 posts

Re: Claude Opus 4.6

#911

Earlier quoted context omitted.

I am having trouble with 4.6 following the most basic of instructions. As an example, I asked it to commit everything in the worktree. I stressed everything and prompted it very explicitly, because even 4.5 sometimes likes to say, "I didn't do that other stuff, I'm only going to commit my stuff even though he said everything". It still only committed a few things. I had to ask again. And again. I had to ask four time…

Tell it what git commands to explicitly run and in what order for your desired outcome instead of “commit everything in the worktree” This prompt will work better across any/all models.

> Tell it what git commands to explicitly run and in what order

Why don't run the commands yourself then?

Re: Claude Opus 4.6

#912
post #569
post #552

Earlier quoted context omitted.

Do you remember how to get around those tricks?

This is the paper: https://arxiv.org/abs/2601.02671 Grok and Deepmind IIRC didn’t require tricks.

What's not clear from the study (at least skimming it) is if they always started the ball rolling with ground truth passages or if they chained outputs from the model until they got to the end of the book. I strongly suspect the latter would hopelessly corrupt relatively quickly.

It seems like this technique only works if you have a copy of the material to work off of, i.e. enter a ground truth passage, tell the model to continue it as long as it can, and then enter the next ground truth passage to continue in the next session.

Re: Claude Opus 4.6

#913
post #905

I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…

Well, if there are papers that match your criteria, it's hallucinating the "no".

It might be wrong but that’s not really a hallucination.

Edit: to give you the benefit of doubt, it probably depends on whether the answer was a definitive “this does not exist” or “I couldn’t find it and it may not exist”

Re: Claude Opus 4.6

#914

Earlier quoted context omitted.

Why does it matter if Claude Code opens in 3-4 seconds if everything you do with it can take many seconds to minutes? Seems irrelevant to me.

Some people[0] like their tools to be well engineered. This is not unique to software. [0] Perhaps everyone who actually takes pride in their craft and doesn’t prioritise shitty hustle culture and making money over everything else.

Aside from startup time, as a tool Claude Code is tremendous. By far the most useful tool I’ve encountered yet. This seems to be very nit picky compared to the total value provided. I think y'all are missing the forrest for the trees.

Re: Claude Opus 4.6

#915

I asked > Can you find an academic article that _looks_ legitimate -- looks like a real journal, by researchers with what look like real academic affiliations, has been cited hundreds or thousands of times -- but is obviously nonsense, e.g. has glaring typos in the abstract, is clearly garbled or nonsensical? It pointed me to a bunch of hoaxes. I clarified: > no, I'm not looking for a hoax, or a deliberate comment on…

> For my tastes telling me "no" instead of hallucinating an answer is a real breakthrough.

It's all anecdata--I'm convinced anecdata is the least bad way to evaluate these models, benchmarks don't work--but this is the behavior I've come to expect from earlier Claude models as well, especially after several back and forth passes where you rejected the initial answers. I don't think it's new.

Re: Claude Opus 4.6

#916

Earlier quoted context omitted.

Surely the corpus Opus 4.6 ingested would include whatever reference you used to check the spells were there. I mean, there are probably dozens of pages on the internet like this: https://www.wizardemporium.com/blog/complete-list-of-harry-p... Why is this impressive? Do you think it's actually ingesting the books and only using those as a reference? Is that how LLMs work at all? It seems more likely it's predicting t…

Most people still don't realize that general public world knowledge is not really a test for a model that was trained on general public world knowledge. I wouldn't be surprised if even proprietary content like the books themselves found their way into the training data, despite what publishers and authors may think of that. As a matter of fact, with all the special deals these companies make with publishers, it is ge…

> If you're a janitor with a high school diploma, there may be barely any textual information or fact you have ever consumed that such a model hasn't seen during training already.

The plot of Good Will Hunting would like a word.

Re: Claude Opus 4.6

#917
post #349
post #38

Earlier quoted context omitted.

CC has >6000 open issues, despite their bot auto-culling them after 60 days of inactivity. It was ~5800 when I looked just a few days ago so they seem to be accelerating towards some kind of bug singularity.

Insane to think that a relatively simple CLI tool has so many open issues...

Well part of the issue is that it isn't actually a CLI tool. It takes control of the whole terminal and then badly reimplements a CLI...

Re: Claude Opus 4.6

#918

Earlier quoted context omitted.

Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?

We don't vary our model quality with time of day or load (beyond negligible non-determinism). It's the same weights all day long with no quantization or other gimmicks. They can get slower under heavy load, though. (I'm from OpenAI.)

sure. we believe you

Re: Claude Opus 4.6

#919
post #482

Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-haystack challenge: finding every spell in all Harry Potter books. All 7 books come to ~1.75M tokens, so they don't quite fit yet. (At this rate of progress, mid-April should do it ) For now you can fit the first 4 books (~733K tokens). Results: Opus 4.6 found 49 out of 50 officially documented spells across those 4 books. The only miss was "Slugulus Eruc…

You need to publish this tbh

Re: Claude Opus 4.6

#920

Earlier quoted context omitted.

We know Open AI got caught getting benchmark data and tuning their models to it already. So the answer is a hard no. I imagine over time it gives a general view of the landscape and improvements, but take it with a large grain of salt.

Are you referring to FrontierMath? We had access to the eval data (since we funded it), but we didn't train on the data or otherwise cheat. We didn't even look at the eval results until after the model had been trained and selected.

No one believes you.
Post reply on HN