Live data from Hacker News

Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

tokens.billchambers.me

601–610 of 620 posts

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#601

Earlier quoted context omitted.

Do you have to use Grok? I don't anyhow that found it passed evaluations.

I find most people who use grok do so for ideological reasons

Yeah, I guess. "Lex Luthor made an AI, I need to support him so I'll use Grok!" is a thing.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#602
post #480

Earlier quoted context omitted.

> Those tools seem mostly useful for a Google alternative, scaffolding tedious things, code reviewing, and acting as a fancy search. Just to get a sense for the rate of change, imagine if you took a survey. Compare what people said about AI tools... 3 years ago, 2 years ago, 1 year ago, 6 months ago. Then think about what is plausible that people will be saying in 3 months, 6 months, 9 months ... Moving the goalposts…

You're relying on the public's sentiment as a metric. The public's sentiment is, more than often, skewed, influenced by marketing, or flat out wrong. That is not a good metric to rely on. Did it ever occur to you that the ever changing goalposts might have more to do with the expensive marketing campaigns of the big LLM providers? We could talk about what's a measurable metric and what's not. Certainly, we have not m…

Hi. I read your message, and I considered it. I've also read some of your previous HN comments. Briefly, I'll just say I've argued at length against many of the claims you make (you certainly aren't alone in making them). I don't feel it would be useful to repeat these again here, but I'll reference a few, below, just to show that I do care about the subject matter and am happy to dig deeper ...

... but only with certain conversational norms. I say this because I predict we aren't (yet) matched up in a way such that we would have a conversation useful to us. The main reason (I guess) isn't about our particular viewpoints nor about i.e. "if we're both critical thinkers". We're both demonstrating that frame, at least in our language. Instead, I think it is about the way we engage and what we want to get out of a conversation. Just to pick one particular guide star, I strive to follow Rapoport's Rules [1]. FWIW, HN Guidelines are not all that different, so simply by commenting here, one is explicitly joining a sort of social contract that point in their direction already.

Anatol Rapoport or Daniel Dennett were not only brilliant in their areas of specialty but also in teaching us how to criticize constructively in general. I offer the link at [1] just in case you want to read them and give them a try, here. We can start the conversation over (if you want).

---

In response to your comments about consciousness, intelligence, etc, here are some examples of what I mean by intelligence and why:

- intelligence: https://news.ycombinator.com/item?id=43236444

- general intelligence: https://news.ycombinator.com/item?id=43223521

- pressure towards AGI: https://news.ycombinator.com/item?id=41707643

- intelligence as "what machines cannot do" / no physics-based constraints to surpass human intelligence: https://news.ycombinator.com/item?id=44974963

---

[1]: https://assets.edge.bigthink.com/uploads/attachment/file/151...

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#603

Earlier quoted context omitted.

I find most people who use grok do so for ideological reasons

Yeah, I guess. "Lex Luthor made an AI, I need to support him so I'll use Grok!" is a thing.

Unfortunately that is kind of how some people operate when it comes to musk. Grokopedia certainly is not used by people because it’s useful.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#604
post #510
post #470

Earlier quoted context omitted.

>>> ... but it seems like Anthropic is going for the Tinder/casino intermittent reinforcement strategy: optimized to keep you spending tokens instead of achieving results. >> This part of the above comment strikes me as uncharitable and overconfident. And, to be blunt, presumptuous. To claim to know a company's strategy as an outsider is messy stuff. > I said "it seems like". Sorry. I take back the "presumptuous" par…

I like your style, and I appreciate you trying to get to the truth, despite us both being aware that we are engaging in persuasive writing here, so part of the rhetorical game is in what we choose to emphasize and what we choose to leave out. > How likely do you think this is? Do you think it is more likely than the other three I mentioned? I won't write down probability estimates, because frankly, I have no idea. Un…

I strive to be decently Bayesian and embrace uncertainty. I'm sharing my probability estimates because it helps me to stop and think ("is this roughly what I think?" and "let spend a minute making sure before I say so"). But yeah, of course, they are my priors and fuzzy. Hopefully I can reflect I figure them out +/- 15% or so. But at least you can see how my takes compare with each other. And down the road I can see how I did.

Thanks for getting into some of the details ...

>> (1) One possibility is they are having capacity and/or infrastructure problems so the model performance is degraded.

> As far as I understand it, scaling issues would result in increased latency or requests being dropped, not model quality being lower.

Yes, many scaling issues would manifest in that way -- but not all. It seems plausible for Anthropic to have other ways to degrade model performance that don't show up in the latency or reliability metrics. I need to research more... (I'll try to think more on your other points later).

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#605
post #597

Earlier quoted context omitted.

And, so the anti-LLM argument goes, if you've not built the computer you can't learn anything about what computers could be used for.

That's not the anti-LLM argument, that's a brand new argument you made up.

Did you not read the comment thread you replied to? That's the exact argument that I_love_retros made above.

That is in fact the anti LLM argument you've ostensibly been discussing. If you want to talk to the person who made it up I'm not your guy.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#606

Earlier quoted context omitted.

You don't have to use adaptive thinking. It had been turned off on my main work computer. I was using a different computer on a trip and I started getting so angry at Claude for doing a bad job. I evetually figured out it was adaptive thinking and set it to "hard" and it started working again. At the time I think "hard" was the top choice. With 4.7, my computer now shows "xhard", which I assume is the equivelent sett…

"With Opus 4.6, extended thinking was a toggle you managed: turn it on for hard stuff, off for quick stuff. If you left it on, every question paid the thinking tax whether it needed to or not. Now, with Opus 4.7, extended thinking becomes adaptive thinking. " https://claude.com/resources/tutorials/working-with-claude-o... You want extended thinking? It's not adaptive thinking and opus will turn it on if it thinks it…

I am getting pretty good performance. Even on trivial questions it seems to go through the thinking process end. If they are using adaptive thinking, it seems to work much better than before. I will see how my experience goes with more usage.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#608
post #74

The bump from 4.6 to 4.7 is not very noticeable to me in improved capabilities so far, but the faster consumption of limits is very noticeable. I hit my 5 hour limit within 2 hours yesterday, initially I was trying the batched mode for a refactor but cancelled after seeing it take 30% of the limit within 5 minutes. Had to cancel and try a serial approach, consumed less (took ~50 minutes, xhigh effort, ~60% of the rem…

The most frustrating part is the quality loss caused by the forced adaptive thinking. It eats 5-10% of my Max 5x usage and churns for ten minutes, only to come back with totally untrustworthy results. It lazily hand-waves issues away in order to avoid reading my actual code and doing real reasoning work on it. Opus simply cannot be trusted if adaptive thinking is enabled.

> It lazily hand-waves issues away in order to avoid reading my actual code and doing real reasoning work on it.

It decided to leave the write endpoints added to an authentication service completely unauthenticated. The effort to do the contrary was about 6 characters, and in the claude.md. It tried to implement PKCE by embedding _everything_ in the state.

This thing is beyond untrustworthy.

The fact that they are using Claude to build Claude (not just Claude Code) probably explains a lot.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#609

Earlier quoted context omitted.

What? You don't think businesses do financial planning and calculations for profit margins? Do you really think they go on vibes - "welp, this AI thing seems to improve developer performance, I guess. Heck, what's an extra 5k per developer anyways, amirite". Well, maybe they really do in your neck of the woods. Explains a lot, I guess.

Yes most companies do in fact operate like this. There are tens of thousands of companies that will pay more for the best thing and call it at that, because the cost is dwarfed by what even marginal gains in quality unlock for the business.

> the cost is dwarfed by what even marginal gains in quality

That is just, like, your opinion, man.

Also, I doubt these kinds of companies have "quality" of anything, never mind "gains in quality".

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#610

Earlier quoted context omitted.

Well it might. If the actual rate is .9x then it matters a lot. Or even if it's like 1.1x, is the cost worth the return?

The cost is so small relative to the increase. The cost whining on HN is bizarre to me. Feels like everyone here is on an individual plan and has no understanding of what margins look like for actual business. Meta pays $750k+ TC and makes far more profit/eng, do you think they care about $5k/eng/mo in inference? A 1.1x increase would be so significant that it would justify the cost easily, especially when you can ju…

Nobody is whining here.
Post reply on HN