About
Question Sources
- CommonsenseQA — everyday common sense that most people find easy.
- GSM8K — grade-school arithmetic word problems.
- AGIEval (LSAT logical reasoning) — real law-school admission test items: read a short argument, spot the flaw or the assumption.
- AGIEval (AQuA-RAT) — GMAT and GRE style quantitative word problems.
- BIG-Bench Hard — logical deduction, object counting, date arithmetic, tracking things as they move, causal judgement, and truth-teller puzzles.
Methodology
The app uses a Rasch model to estimate a percentage score from a relatively small number of items. Information
Reading the range
The shaded bar is a 95% confidence range. It starts enormous and narrows with every answer. If two bars overlap heavily, the difference between them is within the margin of error.
Disclaimers
This is a proof of concept and illustration of how LLM benchmarking works in a human context, but it is not a serious test of reasoning or intelligence and has limited psychometric validty.