> I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us.
Then use a model that supports multi-modal output then, instead of using models that only support raw text output to construct an image.
> If it can't even do that well then I'm not sure why we need to talk about horses here.
Testing a model that only outputs text and forcing it to output an unsupported modality is obviously not going to work well which is what it going on here.
Would you expect Claude/GPT to replace your compiler because it can output text? That would be like expecting that horses can fly because someone saw an ancient winged-horse on a brick tablet.
> Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after.
That is because the premise of the test is ridiculous and pseudoscientific.
> (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)
Why did you say it was a "test for intelligence", when we both know it is a joke that tests for nothing?