Earlier quoted context omitted.
I don’t know much about AI, but I have this image test that everything has failed at. You basically just present an image of a maze and ask the LLM to draw a line through the most optimal path. Here’s how Nano Banana fared: https://x.com/danielvaughn/status/1971640520176029704?s=46
I tried this with Claude: ``` > [Image #1] Create a unicode "ascii-art" version of this image, with the optimal path through the maze highlighted in a solid colour. I'll create an ASCII art version of this maze with the solution path highlighted! ┌─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┬─┐ ●●│ │ │ │ │ │ │ │ │ │ │ │ │ │ ├─┤●└─┴─┐ ├───┐ │ ╔═╗ ├─────┤ ╔═══╝ │ ╔═╝ ╔═╝ │ │ ╔═╝ ├─┤ │ │●●●●●└─┤ │ │ ║ │ │ │ ║…
Gemini 3 Pro: the frontier of vision AI
91–100 of 309 posts
Re: Gemini 3 Pro: the frontier of vision AI
#92Earlier quoted context omitted.
"AI could never replace the creativity of a human" "Ok, I guess it could wipe out the economic demand for digital art, but it could never do all the autonomous tasks of a project manager" "Ok, I guess it could automate most of that away but there will always be a need for a human engineer to steer it and deal with the nuances of code" "Ok, well it could never automate blue collar work, how is it gonna wrench a pipe i…
> blue collar work I don't think it's fair to qualify this as blue collar work
Re: Gemini 3 Pro: the frontier of vision AI
#93Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…
What image are you using? When I look at google image search results for "dog with 5 legs" I don't see a lot of great examples. The first unequivocal "dog with 5 legs" was an illustration. Here was my conversation with Chat GPT. > How many legs does this dog have? "The dog in the image has four legs." > look closer. " looking closely, the drawing is a bit tricky because of the shading, but the dog actually has five v…
Re: Gemini 3 Pro: the frontier of vision AI
#94Earlier quoted context omitted.
Super interesting. I replicated this. I passed the AIs this image and asked them how many fingers were on the hands: https://media.post.rvohealth.io/wp-content/uploads/sites/3/2... Claude said there were 3 hands and 16 fingers. GPT said there are 10 fingers. Grok impressively said "There are 9 fingers visible on these two hands (the left hand is missing the tip of its ring finger)." Gemini smashed it and said 12.
I just re-ran that image through Gemini 3.0 Pro via AI Studio and it reported: I've moved on to the right hand, meticulously tagging each finger. After completing the initial count of five digits, I noticed a sixth! There appears to be an extra digit on the far right. This is an unexpected finding, and I have counted it as well. That makes a total of eleven fingers in the image. This right HERE is the issue. It's not…
Re: Gemini 3 Pro: the frontier of vision AI
#95Re: Gemini 3 Pro: the frontier of vision AI
#96Earlier quoted context omitted.
I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…
Its not that they aren’t intelligent its that they have been RL’d like crazy to not do that Its rather like as humans we are RL’d like crazy to be grossed out if we view a picture of a handsome man and beautiful woman kissing (after we are told they are brother and sister) - Ie we all have trained biases - that we are told to follow and trained on - human art is about subverting those expectations
RL has been used extensively in other areas - such as coding - to improve model behavior on out-of-distribution stuff, so I'm somewhat skeptical of handwaving away a critique of a model's sophistication by saying here it's RL's fault that it isn't doing well out-of-distribution.
If we don't start from a position of anthropomorphizing the model into a "reasoning" entity (and instead have our prior be "it is a black box that has been extensively trained to try to mimic logical reasoning") then the result seems to be "here is a case where it can't mimic reasoning well", which seems like a very realistic conclusion.
Re: Gemini 3 Pro: the frontier of vision AI
#97Earlier quoted context omitted.
"AI could never replace the creativity of a human" "Ok, I guess it could wipe out the economic demand for digital art, but it could never do all the autonomous tasks of a project manager" "Ok, I guess it could automate most of that away but there will always be a need for a human engineer to steer it and deal with the nuances of code" "Ok, well it could never automate blue collar work, how is it gonna wrench a pipe i…
> blue collar work I don't think it's fair to qualify this as blue collar work
Anything like this willl have trouble getting adopted since you'd need these to work with imperfect humans, which becomes way harder. You could bankroll a whole team of subcontractors (e.g. all trades) using that, but you would have one big liability.
The upper end of the complexity is similar to EDA in difficulty, imo. Complete with "use other layers for routing" problems.
I feel safer here than in programming. The senior guys won't be automated out any time soon, but I worry for Indian drafting firms without trade knowledge; the handholding I give them might go to an LLM soon.
Re: Gemini 3 Pro: the frontier of vision AI
#98Earlier quoted context omitted.
I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…
Isn't this proof that LLMs still don't really generalize beyond their training data?
Re: Gemini 3 Pro: the frontier of vision AI
#99Earlier quoted context omitted.
I just tried to get Gemini to produce an image of a dog with 5 legs to test this out, and it really struggled with that. It either made a normal dog, or turned the tail into a weird appendage. Then I asked both Gemini and Grok to count the legs, both kept saying 4. Gemini just refused to consider it was actually wrong. Grok seemed to have an existential crisis when I told it it was wrong, becoming convinced that I ha…
I had no trouble getting it to generate an image of a five-legged dog first try, but I really was surprised at how badly it failed in telling me the number of legs when I asked it in a new context, showing it that image. It wrote a long defense of its reasoning and when pressed, made up demonstrably false excuses of why it might be getting the wrong answer while still maintaining the wrong answer.
Re: Gemini 3 Pro: the frontier of vision AI
#100Well It is the first model to get partial-credit on an LLM image test I have. Which is counting the legs of a dog. Specifically, a dog with 5 legs. This is a wild test, because LLMs get really pushy and insistent that the dog only has 4 legs. In fact GPT5 wrote an edge detection script to see where "golden dog feet" met "bright green grass" to prove to me that there were only 4 legs. The script found 5, and GPT-5 the…
Anything that needs to overcome concepts which are disproportionately represented in the training data is going to give these models a hard time. Try generating: - A spider missing one leg - A 9-pointed star - A 5-leaf clover - A man with six fingers on his left hand and four fingers on his right You'll be lucky to get a 25% success rate. The last one is particularly ironic given how much work went into FIXING the ol…
Surprisingly, it got all of them right