Earlier quoted context omitted.
FYI: I just had three SOTA LLMs + NotebookLM all fail at the simple task of explaining to me where to put a powder detergent in my particular newly bought washing machine, despite having photos of the machine and ability to find the manual (in case of NotebookLM, it literally had the manual as its only source). After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on Hig…
I'm curious now. Was the correct answer not "use the compartment marked with two parallel vertical lines"?
[ 2 | 3 ]
----------
[ 1 ]
1 is for powder detergent, 2 is for liquid detergent, 3 is for softeners and such.All three LLMs (Gemini twice, since NotebookLM) insisted I should put the powder detergent into leftmost compartment (2 on the ASCII diagram above). They referred to it by different numbers, but all gave some convincing justification why to put the detergent there. That's despite me posting photos of the compartment drawer, with symbols clearly visible. That's despite demanding they find the manual and cross-ref. I even asked two (Gemini and Claude) to label the actual compartment on the photos I took[0], and both produced some nonsense, with labels in all the wrong places. And they all insisted they're right and issued plenty of warnings about making sure I get this right or else bad things will happen.
BTW. I ended up posting a screenshot of the diagram in the manual to Claude with a passive aggressive comment. Looking at its "thinking summary" and tool calls now, I think at least Claude didn't process the image correctly and only saw:
[ 2 | 3 ]
------------
as those parts are blue, while the bottom is just in the same color as the entire body/frame of the machine. Maybe the contrast was too low for the models. But it was okay for humans, so it's not excusing much, especially that they all claimed to have found the manual, which had a high-contrast diagram.(Current experience tells me they probably didn't really check the diagram. I noticed recently that all major models seem to have gotten lazy when it comes to reading sources, and are also more than happy to lie about it.)
--
[0] - A method I often use with Claude when I'm not sure if it's dealing with spatial tasks correctly - I have it produce intermediate artifacts that involve modifying "ground truth" inputs - e.g. placing two map pictures on top to verify it solved the coordinate transform, and/or (like here) drawing labels and boxes on top of original photos. I found such requests to be helpful enough I set it as general rule for Claude now.
[1] - Which normally they'd spot, but for some reasons, they didn't.