AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models
arxiv.org
1 thread
I really like this. One of the most realistic evals I've seen, finally quantifying the good vibes many of us feel from the Claude models. Also lmao at CapGPT(-oss).