Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
I haven't been bothered by hallucinations in premier models since early last year. Still see it in smaller local models though.
Coding, however, is solved like magic. Easier to add tests, to be fair.