Settings

Theme

AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models

arxiv.org

6 points by declanjackson · 1 comment

Reader

1 thread
karty1

I really like this. One of the most realistic evals I've seen, finally quantifying the good vibes many of us feel from the Claude models. Also lmao at CapGPT(-oss).

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection