The frontier of GPQA-Dumb models
github.com
1 thread
Current highest ranking model on the GPQA-Dumb benchmark, where the lower the score the higher the score.
I encourage you to take a look at the benchmarks, boasting as low a score as 6% in some categories.
If anyone thinks they can make a worse model, I challenge you to try.