MMLU Benchmark - Knowledge & Reasoning
Multiple-choice questions across the 57 MMLU subjects, fetched live from the free Hugging Face datasets-server API. Each model must reason step by step before answering; the reasoning is parsed, scored for consistency, and shown per question below.
mmlu — benchmark output
Model ranking
Composite score = 0.70 × accuracy + 0.15 × reasoning consistency + 0.15 × relative speed. Per-dimension ranks are shown in parentheses (1 = best).
| # | Model | Accuracy (%) | STEM | Human. | Social | Other | Reasoning consist. (%) | Avg words | Avg time (s) | Composite | Errors |
|---|
Accuracy by model
Accuracy by category
Questions, answers & model reasoning
Every question the selected model saw, its pick vs the correct answer, and the parsed reasoning behind it.