Earlier quoted context omitted.
As opposed to other companies which are smart enough not to report outages.
So, there are only two types of companies: ones that have constant downtime, and ones that have constant downtime but hide it, right?
Claude Opus 4.6
331–340 of 1001 posts
Re: Claude Opus 4.6
#332Earlier quoted context omitted.
Is there a way to disable it? Sometimes I value agent not having knowledge that it needs to cut corners
90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.
Re: Claude Opus 4.6
#333Earlier quoted context omitted.
It would be way way better if they were benchmaxxing this. The pelican in the image (both images) has arms. Pelicans don't have arms, and a pelican riding a bike would use it's wings.
Pelicans don’t ride bikes. You can’t have scruples about whether or not the image of a pelican riding a bike has arms.
Re: Claude Opus 4.6
#334The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
There's no way they actually work on training this.
$200 * 1,000 = $200k/month.
I'm not saying they are, but to say that they aren't with such certainty, when money is on the line; unless you have some insider knowledge you'd like to share with the rest of the class, it seems like an questionable conclusion.
Re: Claude Opus 4.6
#335Earlier quoted context omitted.
Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?
When do you think we should run this benchmark? Friday, 1pm? Monday 8AM? Wednesday 11AM? I definitely suspect all these models are being degraded during heavy loads.
Re: Claude Opus 4.6
#336Earlier quoted context omitted.
90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.
I'm fairly sure OpenAI/GPT does pull prior information in the form of its memories
Re: Claude Opus 4.6
#337Earlier quoted context omitted.
Is there a way to disable it? Sometimes I value agent not having knowledge that it needs to cut corners
90-98% of the time I want the LLM to only have the knowledge I gave it in the prompt. I'm actually kind of scared that I'll wake up one day and the web interface for ChatGPT/Opus/Gemini will pull information from my prior chats.
Re: Claude Opus 4.6
#338Re: Claude Opus 4.6
#339Earlier quoted context omitted.
Same with opencode and gemini, it's disgusting Codex (by openai ironically) seems to be the fastest/most-responsive, opens instantly and is written in rust but doesn't contain that many features Claude opens in around 3-4 seconds Opencode opens in 2 seconds Gemini-cli is an abomination which opens in around 16 second for me right now, and in 8 seconds on a fresh install Codex takes 50ms for reference... -- If their m…
Why does it matter if Claude Code opens in 3-4 seconds if everything you do with it can take many seconds to minutes? Seems irrelevant to me.
Some developers say 3-4 seconds are important to them, others don't. Who decides what the truth is? A human? ClawdBot?
Re: Claude Opus 4.6
#3405.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!
Dumb question. Can these benchmarks be trusted when the model performance tends to vary depending on the hours and load on OpenAI’s servers? How do I know I’m not getting a severe penalty for chatting at the wrong time. Or even, are the models best after launch then slowly eroded away at to more economical settings after the hype wears off?