Live data from Hacker News

Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

tokens.billchambers.me

611–620 of 620 posts

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#611

Earlier quoted context omitted.

The actual rate isn’t relevant for the discussion

What if the rate is negative? Would it matter?

It would matter but would be a different discussion than the one I was going for

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#612

Earlier quoted context omitted.

The most frustrating part is the quality loss caused by the forced adaptive thinking. It eats 5-10% of my Max 5x usage and churns for ten minutes, only to come back with totally untrustworthy results. It lazily hand-waves issues away in order to avoid reading my actual code and doing real reasoning work on it. Opus simply cannot be trusted if adaptive thinking is enabled.

> It lazily hand-waves issues away in order to avoid reading my actual code and doing real reasoning work on it. It decided to leave the write endpoints added to an authentication service completely unauthenticated. The effort to do the contrary was about 6 characters, and in the claude.md. It tried to implement PKCE by embedding _everything_ in the state. This thing is beyond untrustworthy. The fact that they are us…

Opus 4.6 with settings and prompt fixed has been the most reliable for me. It consistently thinks things out and demonstrates good reasoning. It arrives at sound solutions that work. When perfection matters, it's very easy to refine the output with a few review passes.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#613

Earlier quoted context omitted.

I haven't used Claude. Because I suspect this sort of things to come. With enterprise subscription, the bill gets bigger but it's not like VP can easily send a memo to all its staff that a migration is coming. Individuals may end their subscription, that would appease the DC usage, and turn profits up.

Sorry you are missing out. I use claude all day every day with max and what people are reporting here has not been my experience. My current usage is 16% and it resets Thursday.

Some research argue LLMs (up to 2025) were giving a false sense of higher productivity.

That and atrophy, I will pass on what Claude is trying to accomplish.

I'm not dismissing LLMs entirely, for certain cases the concerns don't apply, at least not as much.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#614

Earlier quoted context omitted.

Asking a seller to sell less. That's an incentive difficult to reconcile with the user's benefit. To keep this business running they do need to invest to make the best model, period. It happens to be exactly what Anthropic's strategy is. That and great tooling.

But they're clearly oversubscribed, massively. And they're selling less and less (suddenly 5 hour window lasts 1 hour on the similar tasks it lasted 5 hours a week ago), so IMO they're scamming. I hope many people are making notes and will raise heat soon.

I agree. I'm rather pointing out the whole strategy dictates the outcome.

Anthropic has to keep racing ahead and be stamped offering the best frontier models.

It isn't optimal, so the models cost them disproportionately too much to sell at a profitable price. So they keep feeding the hype and push the costs higher, hoping there won't be too much heat and get away with it.

I wouldn't like to be a leader at such company, but their pay keep them in line.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#615
post #81

Earlier quoted context omitted.

Which plan are you on? I could see that happening with Pro (which I think defaults to Sonnet?), would be surprised with Max…

Opus is not available for claude code in pro

Yes it is.

https://claude.com/pricing

It's not available on Free plan, but it's available on Pro.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#616

Earlier quoted context omitted.

As someone who's switched from mobile to web dev professionally for the last 6 months now. If you care about code quality, you'll develop that neural connection after some time. But if you don't and there's no PR process (side projects), the motivation to form that connection is quite low.

> If you care about code quality, you'll develop that neural connection after some time. No, because you can get LLMs to produce high quality code that has gone through an infinite number of refinement/polish cycles and is far more exhaustive than the code you would have written yourself. Once you hit that point, you find yourself in a directional/steering position divorced from the code since no matter what directio…

Only if never find opportunities to simplify the code it's writing and you don't review the code at all.

> no matter what direction you take, you'll get high quality code

This is not the case today. You get medium-quality, sometimes over-engineered code 10x faster.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#617
post #578

Earlier quoted context omitted.

Bruh. It's getting hard to track down all these MAKE_IT_ACTUALLY_WORK settings that default to off for no reason.

That's the beginning of Googlification of feature evolution, via population statistics rather than quality. If it increases a KPI by 5% for 95% of users but torpedos the experience for 5%? Ship it.

that sounds like a win-win to me.

on one hand 95% of users get an improved experience. While a competitor gets the chance to build a business for the remaining 5%.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#618
post #604
post #510

Earlier quoted context omitted.

I like your style, and I appreciate you trying to get to the truth, despite us both being aware that we are engaging in persuasive writing here, so part of the rhetorical game is in what we choose to emphasize and what we choose to leave out. > How likely do you think this is? Do you think it is more likely than the other three I mentioned? I won't write down probability estimates, because frankly, I have no idea. Un…

I strive to be decently Bayesian and embrace uncertainty. I'm sharing my probability estimates because it helps me to stop and think ("is this roughly what I think?" and "let spend a minute making sure before I say so"). But yeah, of course, they are my priors and fuzzy. Hopefully I can reflect I figure them out +/- 15% or so. But at least you can see how my takes compare with each other. And down the road I can see…

Relevant I think: https://www.anthropic.com/engineering/april-23-postmortem

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#619

45% is brutal if you're building on top of these models as a bootstrapped founder. The unit economics just don't work anymore at that price point for most indie products. What I've been doing is running a dual-model setup — use the cheaper/faster model for the heavy lifting where quality variance doesn't matter much, and only route to the expensive one when the output is customer-facing and quality is non-negotiable.…

[dead]

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#620
For sure Opus 4.7 is more chatty and talkative, I had to explicitly state a "be concise" preference in the settings. Is anyone experiencing some (very very rare) glitch in the output? Broken words, I mean. I'm using the WEB interface extensively, adaptive thinking ON, PRO plan.
Post reply on HN