EntropyMath Leaderboard
Leaderboard
Benchmarks ↑| Model | Acc | Pass@3 | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| API / Others | ||||||||||||||
Gemini-3-Pro-Preview | 100.0 | 100.0 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 |
GPT-5.2 (high) | 83.3 | 83.3 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 1/1 | 1/1 | 0/1 | 1/1 |
| K-LLM Project · 2차 평가 (Round 2) | ||||||||||||||
Solar Pro 3 | 75.0 | 83.3 | 3/3 | 3/3 | 3/3 | 3/3 | 0/3 | 2/3 | 3/3 | 0/3 | 2/3 | 3/3 | 2/3 | 3/3 |
K-EXAONE-236B-A23B | 58.3 | 75.0 | 3/3 | 2/3 | 3/3 | 3/3 | 1/3 | 0/3 | 3/3 | 0/3 | 3/3 | 0/3 | 1/3 | 2/3 |
| Local - KR | ||||||||||||||
Kanana-2-30B-Thinking-2601 | 61.1 | 83.3 | 3/3 | 3/3 | 3/3 | 0/3 | 1/3 | 3/3 | 3/3 | 0/3 | 1/3 | 2/3 | 1/3 | 2/3 |
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect
Model Accuracy vs Pass@3
100%
75%
50%
25%
0%
Gemini-3-Pro-Preview
GPT-5.2 (high)
Solar Pro 3
Kanana-2-30B-Thinking-2601
K-EXAONE-236B-A23B
Accuracy
Pass@3
Avg Token Usage (Per Problem)
20,669
15,501
10,334
5,167
0
K-EXAONE-236B-A23B
Solar Pro 3
Gemini-3-Pro-Preview
Kanana-2-30B-Thinking-2601
GPT-5.2 (high)
Avg Tokens / Problem
About this benchmark
Korean language measures verbal reasoning. Mathematics separates Python TAR from multiple-choice answer selection.
Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language. This CSAT Korean language evaluation measures reading comprehension and verbal reasoning. It is not a Korean translation of the mathematics evaluation.
Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.
Performance Legend
Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3