I hope we standardize on what effort levels mean soon. Right now it has big Spinal Tap "this goes to 11" energy.
Claude Opus 4.7
571–580 of 1001 posts
Re: Claude Opus 4.7
#572Working on some research projects to test Opus 4.7. The first thing I notice is that it never dives straight into research after the first prompt. It insists on asking follow-up questions. "I'd love to dive into researching this for you. Before I start..." The questions are usually silly, like, "What's your angle on this analysis?" It asks some form of this question as the first follow-up every time. The second obser…
Re: Claude Opus 4.7
#573This is a CC harness thing than a model thing but the "new" thinking messages ('hmm...', 'this one needs a moment...') are extraordinarily irritating. They're both entirely uninformative and strictly worse than a spinner. On my workflows CC often spends up to an hour thinking (which is fine if the result is good) and seeing these messages does not build confidence.
Could you say more about your workflow? I don’t think I’ve ever gotten close to an hour of thinking before. Always curious to learn how to get more out of agents.
Re: Claude Opus 4.7
#574Earlier quoted context omitted.
yeah they took "i pick the budget" and turned it into "trust us".
I keep saying even if there's not current malfeasance, the incentives being set up where the model ultimately determines the token use which determines the model provider's revenue will absolutely overcome any safeguards or good intentions given long enough.
Re: Claude Opus 4.7
#575noticing sharp uptick in "i switched to codex" replies lately. a "codex for everything" post flocking the front page on the day of the opus 4.7 release me and coworker just gave codex a 3 day pilot and it was not even close to the accuracy and ability to complete & problem solve through what we've been using claude for. are we being spammed? great. annoying. i clicked into this to read the differences and initial exp…
Re: Claude Opus 4.7
#576Too late, personally after how bad 4.6 was the past week I was pushed to codex, which seems to mostly work at the same level from day to day. Just last night I was trying to get 4.6 to lookup how to do some simple tensor parallel work, and the agent used 0 web fetches and just hallucinated 17K very wrong tokens. Then the main agent decided to pretend to implement tp, and just copied the entire model to each node...
I guess our conscience of OpenAI working with the Department of War has an expiry date of 6 weeks.
Re: Claude Opus 4.7
#577New model - that explains why for the past week/two weeks I had this feeling of 4.6 being much less "intelligent". I hope this is only some kind of paranoia and we (and investors) are not being played by the big corp. /s
I don't get it. Why would they make the previous model worse before releasing an update?
Re: Claude Opus 4.7
#578Earlier quoted context omitted.
Does this mean Claude no longer outputs the full raw reasoning, only summaries? At one point, exposing the LLM's full CoT was considered a core safety tenet.
Anthropic always summarizes the reasoning output to prevent some distillation attacks
Re: Claude Opus 4.7
#579Re: Claude Opus 4.7
#580noticing sharp uptick in "i switched to codex" replies lately. a "codex for everything" post flocking the front page on the day of the opus 4.7 release me and coworker just gave codex a 3 day pilot and it was not even close to the accuracy and ability to complete & problem solve through what we've been using claude for. are we being spammed? great. annoying. i clicked into this to read the differences and initial exp…
Yeah, very. Every single time this happens here, where there's a thread about an Anthropic model and people spam the comments with how Codex is better, I go and try it by giving the exact same prompt to Codex and Opus and comparing the output. And every single time the result is the same: Opus crushes it and Codex really struggles.
I feel like people like me are being gaslit at this point.