GulliBench: Measuring Skepticism in Frontier Models
vetto.ai
2 threads
I liked approach to evaluating AI behavior, the fact that additional reasoning doesn’t improve performance is quite interesting!
interesting approach
I liked approach to evaluating AI behavior, the fact that additional reasoning doesn’t improve performance is quite interesting!
interesting approach