IMDS LogoCicagolab LogoDeep Fountain Logo

EntropyMath Leaderboard

A high-entropy mathematical reasoning benchmark for LLMs

CSAT, Tokyo, and university entrance examinations

CSAT 2026 / Mathematics · English

Korean language evaluates reading comprehension and verbal reasoning. Mathematics is a separate subject, evaluated here using English problem statements.

46 problems 4 models Scores apply to this evaluation only

Jump to leaderboard
3 runs per problem · Errors included · Problem & response details available
  • Accuracy counts correct answers across all three attempts. Pass@3 counts problems with at least one correct answer. Errors and timeouts count as incorrect.
  • Token averages use only runs with reported usage; missing usage is not treated as zero.
  • Select a result cell for the recorded problem, reference answer, model response, and all three attempts. Missing responses and usage remain explicitly unreported.
  • This is the mathematics subject, in English—not the CSAT Korean language/verbal reasoning evaluation. Motif 3 was evaluated manually in web chat.
  • Kimi K3: after an initial retry of the four original timeout slots, 1 completed response was retained and the remaining 3 slots were retried via DeepInfra with provider fallbacks disabled. The same prompt, sampling settings, and 8192-token limit were used. Completed responses replace only the original error slots; any failed retries remain errors. These are not extra best-of runs. Original error and retry records are preserved separately.
  • Kimi K3: 3 latest retry requests reached the 8192-token limit without a final response and remain errors. No token-budget increase has been applied.

Model Accuracy vs Pass@3

100%
75%
50%
25%
0%
Kimi K3
Solar Pro 4
Solar Pro 4
K-EXAONE-2
K-EXAONE-2
Motif 3
Accuracy
Pass@3

Avg Tokens per Reported Run

2.4K
1.8K
1.2K
602.915579710145
0
Solar Pro 4
Solar Pro 4
K-EXAONE-2
K-EXAONE-2
Kimi K3
Avg Tokens / Problem

Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language.

Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.

Performance Legend

Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3

Leaderboard / CSAT 2026 · Mathematics · English

Change benchmark ↑

Scroll horizontally to see all problems. Select a result cell to view the problem and recorded response.

TL = token limit reached before a final answer. These runs still count as incorrect.

ModelAccPass@30123456789101112131415161718192021222324252627282930313233343536373839404142434445
API / Others
Kimi K3API · no tools3 token limits · Tokens: 135/138 runs
92.095.73/33/33/33/33/32/33/33/33/33/33/33/33/33/33/33/33/33/33/33/32/33/33/33/33/33/33/33/33/33/33/33/33/33/30/33/30/3TL ×12/33/33/33/33/31/3TL ×23/33/33/3
K-LLM Project Round 3
Solar Pro 4
Solar Pro 4API · no tools0 errors · Tokens: 138/138 runs
77.584.82/33/33/33/33/33/33/33/33/30/33/33/33/32/33/33/33/33/33/33/30/30/33/33/33/33/33/32/33/30/33/33/33/31/31/33/30/30/33/33/33/33/30/33/31/32/3
K-EXAONE-2
K-EXAONE-2API · no tools0 errors · Tokens: 138/138 runs
65.978.31/33/33/33/33/31/33/33/33/31/33/33/33/31/33/33/33/33/33/30/30/30/33/33/33/33/33/32/32/30/32/32/33/32/30/32/30/30/33/33/32/33/30/31/30/30/3
Motif 3Manual web chat · medium response level0 errors · Tokens not reported
49.373.90/32/32/31/32/30/33/33/33/31/33/33/33/31/30/32/33/32/32/31/31/30/32/30/32/33/31/31/31/31/33/33/31/33/30/30/30/30/32/33/30/30/30/31/31/32/3