AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models
1–2 of 2 posts
Re: AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models
#2I really like this. One of the most realistic evals I've seen, finally quantifying the good vibes many of us feel from the Claude models. Also lmao at CapGPT(-oss).