Gemini Ultra seems better on logic than GPT4. Still messing around testing but here's a prompt Ultra nailed but GPT4 completely botched: Tabitha likes cookies but not cake. She likes mutton but not lamb, and she likes okra but not squash. Following the same rule, will she like cherries or pears https://i.imgur.com/KW6gQbc.jpeg https://i.imgur.com/OSHSvLp.png
I would have never guessed the answer. With such little data available, one can invent any arbitrary rules to fit their favorite answer. It would be more impressive to practical use cases, if a LLM simply said that it's impossible to guess without inventing their own reasoning or looking up the answer online.
In fairness though, GPT4 was objectively incorrect, it's not even internally consistent or coherent - it either thinks b & h are vowels, or that lamb and squash don't end in those letters, or has changed its mind about the rule mid-sentence, or something.