Earlier quoted context omitted.
Its been known that most of these models hallucinate research articles frequently, perplexity.ai seems to do quite well in that regard. Not sure why that is your specific metric though when LLMs seem to be improving across a large class of other metrics.
This wasn't a "metric." It was a test to see whether or not this LLM might actually be useful to me. Just like every other LLM, the answer is a hard no: using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless. For my specific question, I would have been much better off with a plain Google Scholar search.
AKA a Metric lol.
>using this chatbot for real work is at best a huge waste of time, and at worst unconscionably reckless.
Keep living with your head in the sand if you want