Folks have been saying “things are different now, the agents are now compounding success instead of error” for at least a year now, but I just don’t see it. I was lucky enough to receive a weeklong $50k per head AI training from the people saying these things, and one of their few helpful concrete recommendations was to constantly clear context all the time, to avoid things going off the rails. However, I think findi…
Tokenmaxxing is dead, long live tokenmaxxing
301–310 of 315 posts
Re: Tokenmaxxing is dead, long live tokenmaxxing
#302Earlier quoted context omitted.
No, that's not what this analogy is about. At all. See, table saws are dangerous. Famously so. One of, if not the most dangerous tools available to the general public. They spin quickly with lots of torque and pull things in faster than you can react. Pressure can also send loose pieces of wood backwards at high speed. Fast enough to pass through a person sometimes. It's like being hit with an arrow. Tablesaw acciden…
> the worst thing you can do with a table saw is experiment. And yet, that's very much exactly what happened, well predating the table saw; look into steam engines belt driving saw pit blades for logging. The table saw itself has evolved in many ways, there's a handheld angle grinder with various blades sub tree. These attachments: https://www.arbortechtools.com/au/shop-online/power-carving/... are the literal evolut…
I know GP said "available to general public", but my mind went straight to PTO after reading "One of, if not the most dangerous tools". Less common to see (especially in cities) than saws, but I think larger proportion of people understand they have to be careful near table saws than PTO shafts.
> after looking into the crazy world of farmers and military civil engineers
The whole history of aviation and space exploration is chock full of engineers, physicists and chemists doing crazy levels of experimentation.
That said, my point was different - unlike GP, I ask to consider workshops that had extensive experience with dangers of powered or high mechanical leverage hardware. It's entirely plausible and reasonable for people running those shops to say, "here is the new dangerous power tool, it's obviously pretty useful (ask your friends at $X or $Y if you don't see it), figure out how and where to best work it into our specific workflows, so it makes us most bang for the buck".
Re: Tokenmaxxing is dead, long live tokenmaxxing
#303Earlier quoted context omitted.
> They're not "3-4 trillion dollars in investments over 5 years" useful Why not? They're a general-purpose technology, in the same category as "software" or "electricity". > nor "crammed into the throat of every employee on the planet, regardless of their actual job" useful They're potentially useful for anything that can be fed into computers (VLMs lifted the "that can be expressed as text" limitation, visual and au…
LOL, it did work :-) Regarding LLMs, they are pushed too hard and too abusively by business people. Employees are being laid off and replaced with chatbots that don't do the job. Frustrating if support for McDonald's, risky if health insurance support. Also the financials don't make sense. AI companies are money pits. Money is ultimately production. We make X amount of stuff yearly, globally. We can't afford to throu…
After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on High) in parallel, and looking at their thought streams, I gave up and dug up the manual and read half of it, before the LLMs finished coming up with - wait for it - wrong answers.
Super frustrating. Doubly so, given that I use them for comparable tasks pretty often and they usually sail through them flawlessly. But this experience happens every now and then. It's only fair to report it, if only so you don't think I'm just AI boosting all the time.
Re: Tokenmaxxing is dead, long live tokenmaxxing
#304Earlier quoted context omitted.
You don't need peer reviewed studies to tell you water is wet. Peer review is a technique to get evidence from data when SNR is low. It's not "science", it's just a technique. So is "throwing shit at a wall and seeing what sticks". Don't turn techniques into rituals, and science into religion.
Vibes are not evidence, neither is a curated demo. You need actual measured evidence that has an adversarial review to actually prove something without falling to confirmation bias.
Most of the general LLM discourse in our industry is still closer to "proof of the pudding is in the eating" than to "double-blind studies on large cohorts, pAnd we're not talking about curated demos either - most of the contested value can be proven for your own specific cases with little to no expenditure of money and time, at a PoC level (it gets more expensive once you try to operationalize it and find kinks that are hard to iron out).
And that is, the article claims (and I agree), the point of last 6-12 months of tokenmaxxing policies and top-down push - it's putting pressure on people to actually go and do those PoC-s for themselves, because just giving the opportunity and permission turned out insufficient for significant part of the workforce.
--
[0] - Ironically, I remember it was the opposite around the time GPT-4 came out. Back then people talked more about specific claims and demanded measured evidence, because it was hard to get the models to reliably do something interesting. But now that the models can handle bad prompting and can understand you even when you're drunk, suddenly people are denying the general capability of LLMs and asking for randomized control trials.
(For double irony, nowadays one can just ask an LLM for randomized trials; the current SOTA models will happily design you a bespoke eval pipeline if you ask them to.)
Re: Tokenmaxxing is dead, long live tokenmaxxing
#305Earlier quoted context omitted.
Yes, this is, in fact, how adoption of table saws and other such tools looked like, while they were still new tools. The basic form and function was established and its utility proven in both testing and early adopters, but as new kind of general tool on the market, every user from "early majority" was still writing the operating playbook for their specific shop conditions and kind of work they're doing. So yes, it's…
No, that's not what this analogy is about. At all. See, table saws are dangerous. Famously so. One of, if not the most dangerous tools available to the general public. They spin quickly with lots of torque and pull things in faster than you can react. Pressure can also send loose pieces of wood backwards at high speed. Fast enough to pass through a person sometimes. It's like being hit with an arrow. Tablesaw acciden…
The danger may not be "personal injury before you can react", but it is both parts separately, as there's reports of them giving unsafe advice and also of them performing undesired tasks faster than humans can react.
https://techcrunch.com/2026/02/23/a-meta-ai-security-researc...
For everyday tragedy: https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots#Over...
For mass devastation and warcrimes: https://futurism.com/artificial-intelligence/us-military-elo...
(That said: while I regard "AI doom is marketing hype" to be a conspiracy theory when applied to OpenAI and Anthropic, public statements from this guy are absolutely a case where I'd say hyping up destructive power is the point of his job: https://www.ai.mil/About/Leadership/Bio-Page/Article/3940370...)
Re: Tokenmaxxing is dead, long live tokenmaxxing
#306Earlier quoted context omitted.
Vibes are not evidence, neither is a curated demo. You need actual measured evidence that has an adversarial review to actually prove something without falling to confirmation bias.
Proof is not binary, it depends on the claim and the constraints you put around it, and the nature of the subject of your claim. Most of the general LLM discourse in our industry is still closer to "proof of the pudding is in the eating" than to "double-blind studies on large cohorts, p And we're not talking about curated demos either - most of the contested value can be proven for your own specific cases with little…
FWIW, I think most tokenmaxxing is, to riff of what you said earlier, turning a technique into a ritual and science into religion.
This isn't specific to AI, we've had it before with pretty much everything in software (and since well before software), from "object-oriented solves every problem" to "clean code [where every function is] two, or three, or four lines long", to reporting your daily kloc, to bounties for every bug reported and/or fixed.
Humans do what doomers are afraid AI will do: make a sounds-good utility function (tokens, lines of code, bugs, dead cobras) and get surprised when it is easily gamed for something far less helpful than the vision of whoever set the goal.
Re: Tokenmaxxing is dead, long live tokenmaxxing
#307Earlier quoted context omitted.
> You don't need a peer reviewed study to tell you that a heavy rock will fall faster than a light rock. Either I don't understand gravity, or you might want to pick a different analogy...
I think GP is being sarcastic, and pointing out that 1. "heavy rock falls faster" is what common sense will tell you (I was literally told this by multiple laypeople just a few days ago when sightseeing atop a tall tower) 2. This is disproven by a trivial experiment that nobody thought worthy of trying for millenia 3. therefore we do need peer reviewed studies to confirm even "obvious" knowledge. Also, note that GP's…
Indeed, as per 2., no one is doing the experiments with rocks of different weight, and sufficient heights to easily measure time of fall. However, people have a lot of everyday experience with feathers, grains, leaves, wood, and rocks, as well as objects of various weight made of metal, paper, plastics. And in everyday experience, the heuristic actually holds out well: lighter stuff falls slower, or gets carried away by the wind.
This "heuristic" is purely empirical. You can't disprove it with peer-reviewed studies, because within its scope, it's literally the most basic, purest form of science: direct observation.
So in 1., the mistake is that of incorrect generalization. "Lighter stuff falls slower" is correct for everyday experience, it's the "therefore, heavy rock falls faster than light rock" is wrong.
Not because it doesn't fall faster, mind you - it does[0] - it's just that everyday experience is dominated by aerodynamic effects, and laypeople sometimes[1] mistakenly assign it to gravity.
Which I guess makes it a great analogy for the LLM story. Turns out everyday experience is actually valid in everyday situations. Generalizing from it is usually badly wrong, even if it sometimes arrives at correct answer for wrong reasons (and at wrong scales).
Generalization is hard.
--
[0] - Surprise. It's actually a heavy idealized particle falls at the same rate as light idealized particle. Actual matter is not an infinitely small point in space, and generates its own gravity field, so the heavy rock will land a tiny bit sooner than the lighter one, because it pulls Earth stronger towards itself - but then only if you drop the test bodies one by one (serially), and not together (in parallel, where the difference cancels out). But then it also turns out the mass canceling out for idealized particles isn't just a mathematical simplification, but a very deep truth about the universe...
[1] - Or don't. The question as phrased is, "does heavy rock fall faster than light rock"? This isn't a "specific physics theory question", it's a "real life" question. Treating a positive answer as belief on gravity is an error made by the asker.
Re: Tokenmaxxing is dead, long live tokenmaxxing
#308Earlier quoted context omitted.
I think GP is being sarcastic, and pointing out that 1. "heavy rock falls faster" is what common sense will tell you (I was literally told this by multiple laypeople just a few days ago when sightseeing atop a tall tower) 2. This is disproven by a trivial experiment that nobody thought worthy of trying for millenia 3. therefore we do need peer reviewed studies to confirm even "obvious" knowledge. Also, note that GP's…
4. And we need.... something? to realize that both "common sense" / "multiple laypeople" and peer-reviewed studies are right. Indeed, as per 2., no one is doing the experiments with rocks of different weight, and sufficient heights to easily measure time of fall. However, people have a lot of everyday experience with feathers, grains, leaves, wood, and rocks, as well as objects of various weight made of metal, paper,…
> lighter stuff falls slower, or gets carried away by the wind.
Your examples are of smaller-density or larger-surface-area objects, not lighter ones. A bedsheet is heavier than a penny.
> Actual matter is not an infinitely small point in space, and generates its own gravity field, so the heavy rock will land a tiny bit sooner than the lighter one, because it pulls Earth stronger towards itself
When you're timing how long it takes for the rock to land, you're considering the Earth as fixed and applying gravity to the rock's center of mass. The force applied between the Earth and the rock is F = G * (mEarth * mRock) / r^2
So the force that accelerates a twice-as-heavy rock is twice as large.
But the acceleration of that rock towards the earth is a = F/mRock, so in the end, if the rock is twice as heavy, its acceleration is still exactly the same as the lighter rock's.
> but then only if you drop the test bodies one by one (serially), and not together (in parallel, where the difference cancels out).
What are you talking about!?
If you want to split hairs, you could argue that if you drop them serially you're doing a minute change to the Earth's mass (which is actually so minuscule it makes no difference).
But even in your parallel universe of physics where the "heavier rock pulls the earth towards it", you're reaching a paradox similar to the one Galileo was testing for: if I link the heavy and the light rock together, they should fall slower than the heavy rock alone (because the light rock is slowing it down) but also fall faster than the heavy rock (because the total mass of the system is higher).
https://en.wikipedia.org/wiki/Galileo%27s_Leaning_Tower_of_P...
Re: Tokenmaxxing is dead, long live tokenmaxxing
#309Earlier quoted context omitted.
4. And we need.... something? to realize that both "common sense" / "multiple laypeople" and peer-reviewed studies are right. Indeed, as per 2., no one is doing the experiments with rocks of different weight, and sufficient heights to easily measure time of fall. However, people have a lot of everyday experience with feathers, grains, leaves, wood, and rocks, as well as objects of various weight made of metal, paper,…
This is possibly going to lead to a mind-blown moment for me as you reshape my entire understanding of physics. On the other hand, maybe you're slightly mistaken about newtonian physics? > lighter stuff falls slower, or gets carried away by the wind. Your examples are of smaller-density or larger-surface-area objects, not lighter ones. A bedsheet is heavier than a penny. > Actual matter is not an infinitely small poi…
Re: Tokenmaxxing is dead, long live tokenmaxxing
#310Earlier quoted context omitted.
LOL, it did work :-) Regarding LLMs, they are pushed too hard and too abusively by business people. Employees are being laid off and replaced with chatbots that don't do the job. Frustrating if support for McDonald's, risky if health insurance support. Also the financials don't make sense. AI companies are money pits. Money is ultimately production. We make X amount of stuff yearly, globally. We can't afford to throu…
FYI: I just had three SOTA LLMs + NotebookLM all fail at the simple task of explaining to me where to put a powder detergent in my particular newly bought washing machine, despite having photos of the machine and ability to find the manual (in case of NotebookLM, it literally had the manual as its only source). After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on Hig…
Was the correct answer not "use the compartment marked with two parallel vertical lines"?