MMLU Benchmark - Knowledge & Reasoning

Multiple-choice questions across the 57 MMLU subjects, fetched live from the free Hugging Face datasets-server API. Each model must reason step by step before answering; the reasoning is parsed, scored for consistency, and shown per question below.

Subjects
Decoding parameters

MMLU uses greedy decoding by default and a larger token budget than LAMBADA so the model has room to write its reasoning before committing to a letter. Optimal gives the most reasoning room (most reliable accuracy), Normal matches the defaults, and Best Performance caps the reasoning budget so runs finish faster.

mmlu — benchmark output
Model ranking

Composite score = 0.70 × accuracy + 0.15 × reasoning consistency + 0.15 × relative speed. Per-dimension ranks are shown in parentheses (1 = best).

# Model Accuracy (%) STEM Human. Social Other Reasoning consist. (%) Avg words Avg time (s) Composite Errors
Accuracy by model
Accuracy by category
Questions, answers & model reasoning

Every question the selected model saw, its pick vs the correct answer, and the parsed reasoning behind it.