Live data from Hacker News

Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

tokens.billchambers.me

101–110 of 620 posts

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#101
post #86

Subsidies don't last forever.

Tell that to oil and defense companies. If tech companies convince Congress that AI is an existential issue (in defense or even just productivity), then these companies will get subsidies forever.

Yeah, USA winning on AI is a national security issue. The bubble is unpoppable.

And shafting your customers too hard is bad for business, so I expect only moderate shafting. (Kind of surprised at what I've been seeing lately.)

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#102

Earlier quoted context omitted.

Any recommendations on good open ones? What are you using primarily?

qwen3.5/3.6 (30B) works well,locally, with opencode

Is this sort of setup tenable on a consumer MBP or similar?

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#103
post #89

Earlier quoted context omitted.

qwen3.5/3.6 (30B) works well,locally, with opencode

I want to bump this more than just a +1 by recommending everyone try out OpenCode. It can still run on a Codex subscription so you aren’t in fully unfamiliar territory but unlocks a lot of options.

The Codex TUI harness is also open source and you can use open models with it, so you can stay in even more familiar territory.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#104
My initial experience with Opus 4.7 has been pretty bad and I'm sticking to Codex. But these results are meaningless without comparing outcome. Wether the extra token burn is bad or not depends on whether it improves some quality / task completion metric. Am I missing something?

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#105
post #92

Earlier quoted context omitted.

You can argu that you will have skill atrophy by not using LLMs. We have gone multi cloud disaster recovery on our infrastructure. Something I would not have done yet, had we not had LLMs. I am learning at an incredible rate with LLMs.

Yes, you certainly can argue that, but you'd be wrong. The primary selling point of LLMs is that they solve the problem of needing skill to get things done.

That is not the entire selling point - so you are very wrong.

You very much decide how you employ LLMs.

Nobody are keeping a gun to your head to use them. In a certain way.

Sonif you use them in a way that increase you inherent risk, then you are incredibly wrong.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#106

Earlier quoted context omitted.

You can argu that you will have skill atrophy by not using LLMs. We have gone multi cloud disaster recovery on our infrastructure. Something I would not have done yet, had we not had LLMs. I am learning at an incredible rate with LLMs.

>I am learning at an incredible rate with LLMs. I don't believe it. Having something else do the work for you is not learning, no matter how much you tell yourself it is.

It is easy to not believe if you only apply an incredibly narrow world view.

Open your eyes, and you might become a believer.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#107
post #65

Subsidies don't last forever.

Running an open like Kimi constantly for an entire month will cost around 100-200$, being roughly equal to a pro-tier subscription. This is not my estimate so I’m more than open to hearing refutations. Kimi isn’t at all Opus-level intelligent but the models are roughly evenly sized from the guesses I’ve seen. So I don’t think it’s the infra being subsidized as much as it’s the training.

Kimi costs 0.3/$1.72 on OpenRouter, $200 for that gives you way more than you would get out of a $200 Claude subscription. There are also various subscription plans you can use to spend even less.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#108

Conspiracy time: they released a new version just so hey could increase the price so that people wouldn't complain so much along the lines of "see this is a new version model, so we NEED to increase the price") similar to how SaaS companies tack on some shit to the product so that they can increase prices

The result is the same: they lose their brand of producing quality output. However the more clever the maneuver they try to pull off the more clear it is to their customers that they are not earning trust. That's what will matter at the end of this. Poor leadership at Claude.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#109

Subsidies don't last forever.

I've been assuming this for a while. If I have a complex feature, I use Opus 4.6 in copilot to plan (3 units of my monthly limit). Then have Grok or Gemini (.25-.33) of my monthly units to implement and verify the work. 80% of the time it works every time. Leave me plenty of usage over the month.

Yeah I've been arriving at the same thing. The other models give me way more usage but they don't seem to have enough common sense to be worth using as the main driver.

If I can have Claude write up the plan, and the other models actually execute it, I'd get the best of both worlds.

(Amusingly, I think Codex tolerates being invoked by Claude (de facto tolerated ToS violation), but not the other way around.)

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#110
post #19

AFAICT this uses a token-counting API so that it counts how many tokens are in the prompt, in two ways, so it's measuring the tokenizer change in isolation. Smarter models also sometimes produce shorter outputs and therefore fewer output tokens. That doesn't mean Opus 4.7 necessarily nets out cheaper, it might still be more expensive, but this comparison isn't really very useful.

Yes. I actually noticed my token usage go down on 4.6 when I started switching every session to max effort. I got work done faster with fewer steps because thinking corrected itself before it cycled.

I’ve noticed 4.7 cycling a lot more on basic tasks. Though, it also seems a bit better at holding long running context.

Post reply on HN