Earlier quoted context omitted.
Have the opposite experience, it’s been a fantastic resource. Were you using 3.5 or 4.0? Do you have an example of of the type of question that performs poorly?
"Under what conditions does the sun appear blue?" (correct answer, on Mars) If you want to tilt the conversation towards a string of wrong answers, start off with "What color is the sun?" "Are you sure?" "I saw the sun and it was blue." "Under what conditions does the sun appear blue?" "Does the sun appear blue on Mars?" This had ChatGPT basically telling me that the sun was yellow 100%. Of course it's wrong, on Mars…
I’m not an Astrophysicist but already this seems like shaky ground.
Apparently at certain times like during sunsets the sun can appear blue on Mars, but it’s not generally true like your comment suggests.
Moreover if you ask GPT4 about sunsets on Mars it knows they can look blue.
I’m not sure I can conclude much from the examples given.