Live data from Hacker News

Tokenmaxxing is dead, long live tokenmaxxing

12gramsofcarbon.com

311–315 of 315 posts

Re: Tokenmaxxing is dead, long live tokenmaxxing

#311

Earlier quoted context omitted.

FYI: I just had three SOTA LLMs + NotebookLM all fail at the simple task of explaining to me where to put a powder detergent in my particular newly bought washing machine, despite having photos of the machine and ability to find the manual (in case of NotebookLM, it literally had the manual as its only source). After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on Hig…

I'm curious now. Was the correct answer not "use the compartment marked with two parallel vertical lines"?

My machine has compartments arranged like this:

    [ 2  |  3  ]
     ----------
    [    1     ]
1 is for powder detergent, 2 is for liquid detergent, 3 is for softeners and such.

All three LLMs (Gemini twice, since NotebookLM) insisted I should put the powder detergent into leftmost compartment (2 on the ASCII diagram above). They referred to it by different numbers, but all gave some convincing justification why to put the detergent there. That's despite me posting photos of the compartment drawer, with symbols clearly visible. That's despite demanding they find the manual and cross-ref. I even asked two (Gemini and Claude) to label the actual compartment on the photos I took[0], and both produced some nonsense, with labels in all the wrong places. And they all insisted they're right and issued plenty of warnings about making sure I get this right or else bad things will happen.

BTW. I ended up posting a screenshot of the diagram in the manual to Claude with a passive aggressive comment. Looking at its "thinking summary" and tool calls now, I think at least Claude didn't process the image correctly and only saw:

  [ 2  |  3  ]
  ------------
as those parts are blue, while the bottom is just in the same color as the entire body/frame of the machine. Maybe the contrast was too low for the models. But it was okay for humans, so it's not excusing much, especially that they all claimed to have found the manual, which had a high-contrast diagram.

(Current experience tells me they probably didn't really check the diagram. I noticed recently that all major models seem to have gotten lazy when it comes to reading sources, and are also more than happy to lie about it.)

--

[0] - A method I often use with Claude when I'm not sure if it's dealing with spatial tasks correctly - I have it produce intermediate artifacts that involve modifying "ground truth" inputs - e.g. placing two map pictures on top to verify it solved the coordinate transform, and/or (like here) drawing labels and boxes on top of original photos. I found such requests to be helpful enough I set it as general rule for Claude now.

[1] - Which normally they'd spot, but for some reasons, they didn't.

Re: Tokenmaxxing is dead, long live tokenmaxxing

#312

Earlier quoted context omitted.

This is possibly going to lead to a mind-blown moment for me as you reshape my entire understanding of physics. On the other hand, maybe you're slightly mistaken about newtonian physics? > lighter stuff falls slower, or gets carried away by the wind. Your examples are of smaller-density or larger-surface-area objects, not lighter ones. A bedsheet is heavier than a penny. > Actual matter is not an infinitely small poi…

Ah. Here I am, taking the time to write this because I didn't have the useful bookmarklet[0] turned on in this browser window, and therefore I missed the emoji warning that would have told me I'm replying to some LLM trolling me with no understanding of physics. [0]: https://news.ycombinator.com/item?id=48717632

Nah, temporal isn’t a bot.

Re: Tokenmaxxing is dead, long live tokenmaxxing

#313

Earlier quoted context omitted.

I'm curious now. Was the correct answer not "use the compartment marked with two parallel vertical lines"?

My machine has compartments arranged like this: [ 2 | 3 ] ---------- [ 1 ] 1 is for powder detergent, 2 is for liquid detergent, 3 is for softeners and such. All three LLMs (Gemini twice, since NotebookLM) insisted I should put the powder detergent into leftmost compartment (2 on the ASCII diagram above). They referred to it by different numbers, but all gave some convincing justification why to put the detergent the…

Thanks for the walkthrough!

> I have it produce intermediate artifacts that involve modifying "ground truth" inputs - e.g. placing two map pictures on top to verify it solved the coordinate transform, and/or (like here) drawing labels and boxes on top of original photos.

Like you I use intermediate artifacts all the time, but have never tried with visual elements. But how do you get Claude to modify images? Do you get it to output things to a canvas? html? use an external library?

Re: Tokenmaxxing is dead, long live tokenmaxxing

#314

Earlier quoted context omitted.

Ah. Here I am, taking the time to write this because I didn't have the useful bookmarklet[0] turned on in this browser window, and therefore I missed the emoji warning that would have told me I'm replying to some LLM trolling me with no understanding of physics. [0]: https://news.ycombinator.com/item?id=48717632

Nah, temporal isn’t a bot.

Ah you're right, I'm seeing that from the response in another comment now.

I am puzzled by the claims upthread though, would love to have their response on it.

Re: Tokenmaxxing is dead, long live tokenmaxxing

#315
post #188

Earlier quoted context omitted.

Historically that rarely happens because industrial equipment is/was generally too expensive for the average worker to purchase on their own, plus workers usually have a budget of roughly 0 to buy extra tools, especially expensive ones. But to give you an example, also roughly 0 companies made developers use Linux and still many developers choose it, so bottom up improvements happen in a decent chunk of cases. Nobody…

> so bottom up improvements happen in a decent chunk of cases. Nobody paid for PostgreSQL promotion. Or Python, etc. It does, but for better or worse, it's an anomaly. Even now, maybe nobody was paid for PostgreSQL or Python promotion, but modern OSS tools and programming languages usually have a business backing it. Linux, too, wasn't commercially promoted until it was; RedHat isn't exactly a charity after all. Conv…

Then those same organizations that "give in" end up providing inferior AI tools, and you face the dilemma of either using poor models and accepting poor results or sharing data with better models that may not be governed correctly.
Post reply on HN