Live data from Hacker News

Differences in link hallucination and source comprehension across different LLM

mikecaulfield.substack.com

1–10 of 47 posts

Re: Differences in link hallucination and source comprehension across different LLM

#4
Taken from the blog:

> Why are we talking about “graduate and PhD-level intelligence” in these systems if they can’t find and verify relevant links — even directly after a search?

This is my pet peeves, and recently OpenAI's models seem to have become very militant in how they stand by and push their obviously hallucinated sources. I'm talking about hallucinating answers, when pressed to cite sources they also hallucinate URLs that never existed, when repeatedly prompted to verify how the are hallucinating the stick to their clearly wrong output, and ultimately fall back to claiming they were right but the URL somehow changed even though it never existed ever.

In order to start talking about PhD-level intelligence, in the very least these LLMs must support PhD-level context-seeking and information verification. It is not enough to output a wall of text that reads quite fluently. You must stick to verifiable facts.

Re: Differences in link hallucination and source comprehension across different LLM

#5
If anyone is interested in a larger sample size comparing how often LLMs confabulate answers based on provided texts, I have a benchmark at https://github.com/lechmazur/confabulations/. It's always interesting to test new models with it because the results can be unintuitive compared to those from my other benchmarks.

Re: Differences in link hallucination and source comprehension across different LLM

#7
post #4

Taken from the blog: > Why are we talking about “graduate and PhD-level intelligence” in these systems if they can’t find and verify relevant links — even directly after a search? This is my pet peeves, and recently OpenAI's models seem to have become very militant in how they stand by and push their obviously hallucinated sources. I'm talking about hallucinating answers, when pressed to cite sources they also halluc…

I have search enabled 100% of the time with ChatGPT and would never go back to raw-dogging LLM citations. O3 especially has passed the threshold of “not always annoying”. Had an argument with Gemini yesterday where it was insisting on some hallucinated implementation of a function even while giving me a GitHub link to the correct source.

Re: Differences in link hallucination and source comprehension across different LLM

#8
post #5

If anyone is interested in a larger sample size comparing how often LLMs confabulate answers based on provided texts, I have a benchmark at https://github.com/lechmazur/confabulations/ . It's always interesting to test new models with it because the results can be unintuitive compared to those from my other benchmarks.

Useful benchmark. I noticed o3-high hallucinating too often for such a good model, but it is usually great with search. In my experience, Claude Opus & Sonnet 4 consistently lie, cheat, and try to hide their tracks. Maybe they are good in writing code but I don't trust them with other things.

Re: Differences in link hallucination and source comprehension across different LLM

#9
post #7
post #4

Taken from the blog: > Why are we talking about “graduate and PhD-level intelligence” in these systems if they can’t find and verify relevant links — even directly after a search? This is my pet peeves, and recently OpenAI's models seem to have become very militant in how they stand by and push their obviously hallucinated sources. I'm talking about hallucinating answers, when pressed to cite sources they also halluc…

I have search enabled 100% of the time with ChatGPT and would never go back to raw-dogging LLM citations. O3 especially has passed the threshold of “not always annoying”. Had an argument with Gemini yesterday where it was insisting on some hallucinated implementation of a function even while giving me a GitHub link to the correct source.

[flagged]

Re: Differences in link hallucination and source comprehension across different LLM

#10
post #7

Earlier quoted context omitted.

I have search enabled 100% of the time with ChatGPT and would never go back to raw-dogging LLM citations. O3 especially has passed the threshold of “not always annoying”. Had an argument with Gemini yesterday where it was insisting on some hallucinated implementation of a function even while giving me a GitHub link to the correct source.

[flagged]

Languages evolve and words get new meanings all the time.
Post reply on HN