Earlier quoted context omitted.
Which open weights model?
> Which open weights model? Yes, I'm also wondering! Currently I'm testing out gemma4:26b and qwen3.6:35b-a3b-q4_K_M locally on my M2 Max Macbook Pro. Not the fastest, but reasonable. However, I am also interested in getting as close as possible in performance to Opus 4.6 while minimizing my costs.
Claude Opus 4.7
941–950 of 1001 posts
Re: Claude Opus 4.7
#942Re: Claude Opus 4.7
#943noticing sharp uptick in "i switched to codex" replies lately. a "codex for everything" post flocking the front page on the day of the opus 4.7 release me and coworker just gave codex a 3 day pilot and it was not even close to the accuracy and ability to complete & problem solve through what we've been using claude for. are we being spammed? great. annoying. i clicked into this to read the differences and initial exp…
Well, I can share my experience from a few days ago. Gave the same task (a major refactor) to both Claude and Codex. Codex finished in 5 minutes, Claude was still spinning after 20 minutes. Also it used up all my usage, about twice over (the 5-hour window rolled over in the middle of the task, so the usage for one task added up to 192%). Codex usage was 9%. So, 21x difference there, lol They're saying there's bugs la…
Re: Claude Opus 4.7
#944Re: Claude Opus 4.7
#945I'm finding the "adaptive thinking" thing very confusing, especially having written code against the previous thinking budget / thinking effort / etc modes: https://platform.claude.com/docs/en/build-with-claude/adapti... Also notable: 4.7 now defaults to NOT including a human-readable reasoning token summary in the output, you have to add "display": "summarized" to get that: https://platform.claude.com/docs/en/build-…
Re: Claude Opus 4.7
#946Earlier quoted context omitted.
Sounds like you will need to drink a(n identity) verification can soon [1] to continue as a security researcher on their platform. 1: https://support.claude.com/en/articles/14328960-identity-ver... Identity verification on Claude Being responsible with powerful technology starts with knowing who is using it. Identity verification helps us prevent abuse, enforce our usage policies, and comply with legal obligations. W…
What do you offer as a solution? If theoretically some foreign state intelligence was exposed using Claude for security penetration that affected the stability of your home government due to Antropic's lax safety controls, are you going to defend Anthropic because their reasoning was to allow everyone to be able to do security research?
Re: Claude Opus 4.7
#947Earlier quoted context omitted.
Usually they're hemorrhaging performance while training. From that it's pretty likely they were training mythos for the last few weeks, and then distilling it to opus 4.7 Pure speculation of course, but would also explain the sudden performance gains for mythos - and why they're not releasing it to the general public (because it's the undistilled version which is too expensive to run)
Mythos is speculated to have 10 trillion parameters. Almost certainly they were training it for months.
It's been like that for each model release within the last year
Re: Claude Opus 4.7
#948So strange, i've been using Opus 4.7 in Claude code all day today and i've had no malware related comments or issues at all. It's been performing noticably better, and picking up on things it wasn't before. Maybe because i'm using xhigh effort, but i'm super happy with this update!
Re: Claude Opus 4.7
#949They've increased their cybersecurity usage filters to the point that Opus 4.7 refuses to work on any valid work, even after web fetching the program guidelines itself and acknowledging "This is authorized research under the [Redacted] Bounty program, so the findings here are defensive research outputs, not malware. I'll analyze and draft, not weaponize anything beyond what's needed to prove the bug to [Redacted]. I…
FYI, unless you specifically get verified [0], GPT-5.4 silently reroutes request to GPT-5.2 if an intermediate model detects any cybersecurity work.
Re: Claude Opus 4.7
#950This comment thread is a good learner for founders; look at how much anguish can be put to bed with just a little honest communication. 1. Oops, we're oversubscribed. 2. Oops, adaptive reasoning landed poorly / we have to do it for capacity reasons. 3. Here's how subscriptions work. Am I really writing this bullet point? As someone with a production application pinned on Opus 4.5, it is extremely difficult to tell ap…
These threads are always full of superstitious nonsense. Had a bad week at the AIs? Someone at Anthropic must have nerfed the model! The roulette wheel isn't rigged, sometimes you're just unlucky. Try another spin, maybe you'll do better. Or just write your own code.
Though I reckon even if the HN crowd is a loud minority Anthropic has no problem with traction, and even if eventually it will the enterprise market doesn't care much about HN threads.