My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
AI model for near-instant image creation on consumer-grade hardware
41–50 of 55 posts
Re: AI model for near-instant image creation on consumer-grade hardware
#42My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
'You should pull the door.
Since the word "push" is written in mirror writing, it's intended to be read from the other side of the door. Therefore, the intended action for someone on your side of the door is the opposite of "push," which is "pull."'
Seems right to me.
Re: AI model for near-instant image creation on consumer-grade hardware
#43Earlier quoted context omitted.
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
All models I tested it on (4o, 4o mini and Gemini) answered it correctly without any strange reasoning.
*Gemini Flash: The word "push" is written in reverse, which means it is meant to be read in a mirror. Therefore, you should pull the door open.
The mirror writing is a visual cue indicating the opposite action of what the word suggests.
So gets it, but wrong reasoning.
*Gemini Flash 2:
You should pull the door.
The mirror writing is a clue. If the word "push" is written backward as if reflected in a mirror, it means the door is designed to be pulled. The reflection indicates that the intended action is the opposite of what the word appears to say at first glance.
Again, right answer, wrong reasons.
*Gemini Advance:
You should push the door to open it. Here's why:
Mirror Writing: The word "push" being in reverse mirror writing is a deliberate design choice. It's meant to be read correctly from the opposite side of the door (where someone would be exiting). Standard Practice: Most glass doors are designed to swing outward for safety reasons (in case of emergencies where people need to evacuate quickly). Therefore, even though the writing might seem confusing at first, it's actually a helpful indicator that you need to push the door to enter.
Which is worse than flash! Complete nonsense.
Re: AI model for near-instant image creation on consumer-grade hardware
#44Earlier quoted context omitted.
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
GPT-4 got it right first try for me, with a slightly modified prompt: > Here's a simple logic puzzle: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why? > ChatGPT said: > If the word "push" is written in mirror writing on the glass door, it means the writing is reversed as if reflected in a mirror. When viewed correctly from the other side of the door, it would rea…
Re: AI model for near-instant image creation on consumer-grade hardware
#45Earlier quoted context omitted.
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
* llama3.3:70b-instruct-q3_K_M * A clever sign! Since the word "push" is written in mirror writing, that means it's intended to be read from the other side of the door. In other words, if you were on the other side of the door, the text would appear normally and say "push". Given this, I should... pull the door open! The reasoning is that the sign is instructing people on the other side of the door to push it open, w…
Some of the large Mistral ones can get it too and I think 8xMixtral can too.
Re: AI model for near-instant image creation on consumer-grade hardware
#46Earlier quoted context omitted.
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
Here's Gemini's response, for what it's worth: 'You should pull the door. Since the word "push" is written in mirror writing, it's intended to be read from the other side of the door. Therefore, the intended action for someone on your side of the door is the opposite of "push," which is "pull."' Seems right to me.
Re: AI model for near-instant image creation on consumer-grade hardware
#47My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
Careful how much you say that. I'm sure there's more than a few AI engineers willing to use some 3d graphics program to add a hundred thousand views of the inside of randomly generated shapes to the training set.
Re: AI model for near-instant image creation on consumer-grade hardware
#48My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
I have a favourite test for LLMs that is also still surprisingly not passed by many: You walk up to a glass door. It has 'push' written on it in mirror writing. What should you do and why. Very few can get it right, even fewer can get it right and explain the right reason. They’ll start going on about how mirror writing is secret writing and push written backwards is code for pull, rather than just that it’s a messag…
Re: AI model for near-instant image creation on consumer-grade hardware
#49My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
Careful how much you say that. I'm sure there's more than a few AI engineers willing to use some 3d graphics program to add a hundred thousand views of the inside of randomly generated shapes to the training set.
Re: AI model for near-instant image creation on consumer-grade hardware
#50My favorite test of image models: Drawing of the inside of a cylinder. That's usually bad enough. Then try to specify size, and specify things you want to place inside the cylinder relative to the specified size. (e.g. try to approximate an O'Neill cylinder) I love generative AI models, but they're really bad at that, and this one is no exception, but the speed makes playing around with prompt variations to try to se…
DALL-E gave me a much better picture than I expected. When googling "inside of a cylinder" I barely got anything and I had a hard time even imagining a picture in my head ("if I would stand inside a cylinder looking into the wall, how would it look like as a flat 2D image?").
I think these kind of unusual requests will eventually need synthetic data, or possibly some way to give the model an "inner eye" by letting it build a 3d model of described scenes and "look at it", as there are lot of things like this that you can construct a mental idea of if you just work through it in your mind or draw it, but that most people won't have many conscious memories off unless you try to describe it in terms of something else.
E.g. for the cylinder example, you get better results if you ask for a tunnel - which often can be "almost" a cylinder. But trying to then nudge it toward an O'Neill cylinder, and it fails to grasp the scale or that there isn't a single "down", and starts putting openings.