Live data from Hacker News

Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

tokens.billchambers.me

571–580 of 620 posts

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#571
post #446

Earlier quoted context omitted.

Precisely. I find Grok’s multi-agent approach very useful here. I have custom agent configured as a validator.

Do you have to use Grok? I don't anyhow that found it passed evaluations.

I find most people who use grok do so for ideological reasons

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#572

Earlier quoted context omitted.

You're absolutely wrong! You can also ask an LLM to solve that problem by spelling the word out first. And then it'll count the letters successfully. At a similar success rate to actual nine-year-olds. There's a technical explanation for why that works, but to you, it might as well be black magic. And if you could get a modern agentic LLM that somehow still fails that test? Chances are, it would solve it with no inst…

This is false. You can ask it to spell out strawberry and count the letters and it will still say 2 (it's unable to actually count the letters by the way). The only way to get a model that believes strawberry has 2 R's to consistently give the correct answer is to ask it to code the problem and return the output. In fact, asking a model not to repeat the same mistake makes it more likely to commit that mistake again,…

Have you actually tried?

The "spell out" trick, by the way, was what was added to the system prompts of frontier models back when this entire meme was first going around. It did mitigate the issue.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#573

Earlier quoted context omitted.

I agree that the argument starts from a reduction to the absurd. So at least we can agree that AI hasn't mastered the text medium , without further qualification? And what about my argument, further qualified, which is that I don't think it could even write as well as a good professional writer - not necessarily a generational one?

>AI hasn't mastered the text medium I don't know what this means and I don't know what would qualify it as having "mastered" at all. Seems like a no-true-Scotsman thing where regardless there would always be someone that it couldn't actually do a thing because this and that. >why can't they write a novel? This is what I'm disagreeing with. I think an LLM can write a novel well enough that it's recognizably a pretty m…

I don't dislike AI, I use it every day for coding and increasingly for non-technical tasks, and have also used it in enterprise workloads to great success. I am fairly optimistic about it - I think it will remove a lot of drudgery and make things economical which previously weren't.

I am just challenging the notions that "if you limit it to text, it's doing really well" or that the text contains in itself all the information that is needed to carry out a task to a certain level of quality. This applies in my experience not only to writing literature but also to certain human tasks which may appear mundane and easy to automate.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#574

Earlier quoted context omitted.

The pro plan is useless. You need at least the 5x max plan to get any real work done. That said I find the GPT plans much better value.

Yeah. But then you have to use GPT

GPT 5.4 in Codex is good enough now.

Try it.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#575

Earlier quoted context omitted.

You're absolutely wrong! You can also ask an LLM to solve that problem by spelling the word out first. And then it'll count the letters successfully. At a similar success rate to actual nine-year-olds. There's a technical explanation for why that works, but to you, it might as well be black magic. And if you could get a modern agentic LLM that somehow still fails that test? Chances are, it would solve it with no inst…

> it's so powerful that finding "anything better" is incredibly hard. We're back around to the start again. "Incredibly hard" is doing all of the heavy lifting in this statement, it's not all-powerful and there are enormous failure cases. Neither the human brain nor LLMs are a panacea for thought, but nobody in academia or otherwise is seriously comparing GPT to the human brain. They're distinct. > There's a technica…

I do know far more than you, which is a laughably low bar. If you want someone to hold your hand through it, ask an LLM.

The key words are "tokenization" and "metaknowledge", the latter being the only non-trivial part. An LLM can explain it in detail. They know more than you do too.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#576

Earlier quoted context omitted.

The cost is so small relative to the increase. The cost whining on HN is bizarre to me. Feels like everyone here is on an individual plan and has no understanding of what margins look like for actual business. Meta pays $750k+ TC and makes far more profit/eng, do you think they care about $5k/eng/mo in inference? A 1.1x increase would be so significant that it would justify the cost easily, especially when you can ju…

What? You don't think businesses do financial planning and calculations for profit margins? Do you really think they go on vibes - "welp, this AI thing seems to improve developer performance, I guess. Heck, what's an extra 5k per developer anyways, amirite". Well, maybe they really do in your neck of the woods. Explains a lot, I guess.

Yes most companies do in fact operate like this. There are tens of thousands of companies that will pay more for the best thing and call it at that, because the cost is dwarfed by what even marginal gains in quality unlock for the business.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#577

Earlier quoted context omitted.

Adaptive thinking is optional

Not when you want extended thinking - you select extended thinking and opus decides if you get it with apativenthinking. "With Opus 4.6, extended thinking was a toggle you managed: turn it on for hard stuff, off for quick stuff. If you left it on, every question paid the thinking tax whether it needed to or not. Now, with Opus 4.7, extended thinking becomes adaptive thinking. " https://claude.com/resources/tutorials/…

...are you talking about the app? Come on. The app is for quick queries. You should be using Claude Code or Cowork.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#578

Earlier quoted context omitted.

Only if you set `ENABLE_PROMPT_CACHING_1H`, which was mentioned in the release notes for a recent Claude Code release but doesn't seem to be in the official docs.

Bruh. It's getting hard to track down all these MAKE_IT_ACTUALLY_WORK settings that default to off for no reason.

That's the beginning of Googlification of feature evolution, via population statistics rather than quality.

If it increases a KPI by 5% for 95% of users but torpedos the experience for 5%? Ship it.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#579
post #565

Earlier quoted context omitted.

The link you are commenting on shows data from actual prompts from real users, and the COST of the average prompt increased 37%. I do not think synthetic benchmarks are a rebuttal to real usage data.

The cost of the input tokens, not the reasoning or output. Agree though that benchmarks aren't very helpful w.r.t. estimating real world performance or costs. What we'd need are people giving the same real world tasks to 4.6 and 4.7 and measuring time, quality and costs.

Thanks, that wasn't clear because it mentioned conversations, but it is only measuring the input tokens. So its just measuring the difference in the tokenizer.

Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7

#580
post #540
post #517

Earlier quoted context omitted.

> Today, in ~5 minutes I can do a literature review that would have taken me easily 10+ hours five years ago. And it will not yield the same outcome you would have had. Your own taste in clicking links and pre-filtering as you do your research, is no longer being done if you outsource this. I‘m guilty of this myself. But let’s not kid ourselves. I’ve had GPT Pro think 40 minutes about the ideal reverse osmosis setup…

I don't claim using AI is the same as doing it yourself. My point is that AI capabilities are much more extensive than "fancy search". By giving a metric and an example I hoped to make that point without getting into hair-splitting.

I wouldn’t call that hair-splitting. I’m saying, it’s not a real literature review, but even fancier search.
Post reply on HN