Live data from Hacker News

Claude Opus 4.6

anthropic.com

331–340 of 1001 posts

Re: Claude Opus 4.6

#331
post #210

Earlier quoted context omitted.

As opposed to other companies which are smart enough not to report outages.

So, there are only two types of companies: ones that have constant downtime, and ones that have constant downtime but hide it, right?

Basically, yes.

Re: Claude Opus 4.6

#332
post #151

Earlier quoted context omitted.

Is there a way to disable it? Sometimes I value agent not having knowledge that it needs to cut corners

90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.

I'm fairly sure OpenAI/GPT does pull prior information in the form of its memories

Re: Claude Opus 4.6

#333

Earlier quoted context omitted.

It would be way way better if they were benchmaxxing this. The pelican in the image (both images) has arms. Pelicans don't have arms, and a pelican riding a bike would use it's wings.

Pelicans don’t ride bikes. You can’t have scruples about whether or not the image of a pelican riding a bike has arms.

Wouldn’t any decent bike-riding pelican have a bike tailored to pelicans and their wings?

Re: Claude Opus 4.6

#334
post #40

The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...

There's no way they actually work on training this.

The people that work at Anthropic are aware of simonw and his test, and people aren't unthinking data-driven machines. How valid his test is or isn't, a better score on it is convincing. If it gets, say, 1,000 people to use Claude Code over Codex, how much would that be worth to Anthropic?

$200 * 1,000 = $200k/month.

I'm not saying they are, but to say that they aren't with such certainty, when money is on the line; unless you have some insider knowledge you'd like to share with the rest of the class, it seems like an questionable conclusion.

Re: Claude Opus 4.6

#335

Earlier quoted context omitted.

Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?

When do you think we should run this benchmark? Friday, 1pm? Monday 8AM? Wednesday 11AM? I definitely suspect all these models are being degraded during heavy loads.

This hypothesis is tested regularly by plenty of live benchmarks. The services usually don't decay in performance.

Re: Claude Opus 4.6

#336

Earlier quoted context omitted.

90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.

I'm fairly sure OpenAI/GPT does pull prior information in the form of its memories

Ah, that could explain why I've found myself using it the least.

Re: Claude Opus 4.6

#337
post #151

Earlier quoted context omitted.

Is there a way to disable it? Sometimes I value agent not having knowledge that it needs to cut corners

90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.

Gemini has this feature but it’s opt-in.

Re: Claude Opus 4.6

#338

Earlier quoted context omitted.

There's no way they actually work on training this.

I suspect they're training on this. I asked Opus 4.6 for a pelican riding a recumbent bicycle and got this. https://i.imgur.com/UvlEBs8.png

perhaps try a penny farthing?

Re: Claude Opus 4.6

#339

Earlier quoted context omitted.

Same with opencode and gemini, it's disgusting Codex (by openai ironically) seems to be the fastest/most-responsive, opens instantly and is written in rust but doesn't contain that many features Claude opens in around 3-4 seconds Opencode opens in 2 seconds Gemini-cli is an abomination which opens in around 16 second for me right now, and in 8 seconds on a fresh install Codex takes 50ms for reference... -- If their m…

Why does it matter if Claude Code opens in 3-4 seconds if everything you do with it can take many seconds to minutes? Seems irrelevant to me.

This is exactly the type of thing that AI code writers don't do well - understand the prioritization of feature development.

Some developers say 3-4 seconds are important to them, others don't. Who decides what the truth is? A human? ClawdBot?

Re: Claude Opus 4.6

#340

5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!

Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?

On benchmarks GPT 5.2 was roughly equivalent to Opus 4.5 but most people who've used both for SWE stuff would say that Opus 4.5 is/was noticeably better
Post reply on HN