Earlier quoted context omitted.
Precisely. I find Grok’s multi-agent approach very useful here. I have custom agent configured as a validator.
Do you have to use Grok? I don't anyhow that found it passed evaluations.
Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
571–580 of 620 posts
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#572Earlier quoted context omitted.
You're absolutely wrong! You can also ask an LLM to solve that problem by spelling the word out first. And then it'll count the letters successfully. At a similar success rate to actual nine-year-olds. There's a technical explanation for why that works, but to you, it might as well be black magic. And if you could get a modern agentic LLM that somehow still fails that test? Chances are, it would solve it with no inst…
This is false. You can ask it to spell out strawberry and count the letters and it will still say 2 (it's unable to actually count the letters by the way). The only way to get a model that believes strawberry has 2 R's to consistently give the correct answer is to ask it to code the problem and return the output. In fact, asking a model not to repeat the same mistake makes it more likely to commit that mistake again,…
The "spell out" trick, by the way, was what was added to the system prompts of frontier models back when this entire meme was first going around. It did mitigate the issue.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#573Earlier quoted context omitted.
I agree that the argument starts from a reduction to the absurd. So at least we can agree that AI hasn't mastered the text medium , without further qualification? And what about my argument, further qualified, which is that I don't think it could even write as well as a good professional writer - not necessarily a generational one?
>AI hasn't mastered the text medium I don't know what this means and I don't know what would qualify it as having "mastered" at all. Seems like a no-true-Scotsman thing where regardless there would always be someone that it couldn't actually do a thing because this and that. >why can't they write a novel? This is what I'm disagreeing with. I think an LLM can write a novel well enough that it's recognizably a pretty m…
I am just challenging the notions that "if you limit it to text, it's doing really well" or that the text contains in itself all the information that is needed to carry out a task to a certain level of quality. This applies in my experience not only to writing literature but also to certain human tasks which may appear mundane and easy to automate.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#574Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#575Earlier quoted context omitted.
You're absolutely wrong! You can also ask an LLM to solve that problem by spelling the word out first. And then it'll count the letters successfully. At a similar success rate to actual nine-year-olds. There's a technical explanation for why that works, but to you, it might as well be black magic. And if you could get a modern agentic LLM that somehow still fails that test? Chances are, it would solve it with no inst…
> it's so powerful that finding "anything better" is incredibly hard. We're back around to the start again. "Incredibly hard" is doing all of the heavy lifting in this statement, it's not all-powerful and there are enormous failure cases. Neither the human brain nor LLMs are a panacea for thought, but nobody in academia or otherwise is seriously comparing GPT to the human brain. They're distinct. > There's a technica…
The key words are "tokenization" and "metaknowledge", the latter being the only non-trivial part. An LLM can explain it in detail. They know more than you do too.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#576Earlier quoted context omitted.
The cost is so small relative to the increase. The cost whining on HN is bizarre to me. Feels like everyone here is on an individual plan and has no understanding of what margins look like for actual business. Meta pays $750k+ TC and makes far more profit/eng, do you think they care about $5k/eng/mo in inference? A 1.1x increase would be so significant that it would justify the cost easily, especially when you can ju…
What? You don't think businesses do financial planning and calculations for profit margins? Do you really think they go on vibes - "welp, this AI thing seems to improve developer performance, I guess. Heck, what's an extra 5k per developer anyways, amirite". Well, maybe they really do in your neck of the woods. Explains a lot, I guess.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#577Earlier quoted context omitted.
Adaptive thinking is optional
Not when you want extended thinking - you select extended thinking and opus decides if you get it with apativenthinking. "With Opus 4.6, extended thinking was a toggle you managed: turn it on for hard stuff, off for quick stuff. If you left it on, every question paid the thinking tax whether it needed to or not. Now, with Opus 4.7, extended thinking becomes adaptive thinking. " https://claude.com/resources/tutorials/…
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#578Earlier quoted context omitted.
Only if you set `ENABLE_PROMPT_CACHING_1H`, which was mentioned in the release notes for a recent Claude Code release but doesn't seem to be in the official docs.
Bruh. It's getting hard to track down all these MAKE_IT_ACTUALLY_WORK settings that default to off for no reason.
If it increases a KPI by 5% for 95% of users but torpedos the experience for 5%? Ship it.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#579Earlier quoted context omitted.
The link you are commenting on shows data from actual prompts from real users, and the COST of the average prompt increased 37%. I do not think synthetic benchmarks are a rebuttal to real usage data.
The cost of the input tokens, not the reasoning or output. Agree though that benchmarks aren't very helpful w.r.t. estimating real world performance or costs. What we'd need are people giving the same real world tasks to 4.6 and 4.7 and measuring time, quality and costs.
Re: Anonymous request-token comparisons from Opus 4.6 and Opus 4.7
#580Earlier quoted context omitted.
> Today, in ~5 minutes I can do a literature review that would have taken me easily 10+ hours five years ago. And it will not yield the same outcome you would have had. Your own taste in clicking links and pre-filtering as you do your research, is no longer being done if you outsource this. I‘m guilty of this myself. But let’s not kid ourselves. I’ve had GPT Pro think 40 minutes about the ideal reverse osmosis setup…
I don't claim using AI is the same as doing it yourself. My point is that AI capabilities are much more extensive than "fancy search". By giving a metric and an example I hoped to make that point without getting into hair-splitting.