EntropyMath Leaderboard

CSAT 2025Korean language · Verbal reasoning

13 problems 5 models

K-LLM Project · 독파모 / EntropyMath 평가 차수 안내

K-LLM Project는 대한민국 과학기술정보통신부의 독자 AI 파운데이션 모델 프로젝트(통칭 독파모)을 뜻합니다.

Round 1·2·3은 EntropyMath의 1차·2차·3차 모델 평가를 뜻합니다. 과기정통부의 공식 단계평가 번호나 통과 여부를 뜻하지 않습니다.

프로젝트 공식 안내 공식 1차 단계평가 공식 2차 단계평가

Jump to leaderboard

Leaderboard

Benchmarks ↑
ModelAccPass@30123456789101112
API / Others
GPT-5.2 (high)
100.0100.01/11/11/11/11/11/11/11/11/11/11/11/11/1
Gemini-3-Pro-Preview
100.0100.01/11/11/11/11/11/11/11/11/11/11/11/11/1
K-LLM Project · 2차 평가 (Round 2)
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
66.784.61/30/33/31/32/33/33/32/33/30/32/33/33/3
K-LLM Project · 1차 평가 (Round 1)
EXAONE-4.0.1-32B (high)
EXAONE-4.0.1-32B (high)
53.876.91/30/33/32/31/31/33/33/31/30/30/33/33/3
Local - KR
Kanana-2-30B-Thinking-2601
Kanana-2-30B-Thinking-2601
53.869.20/30/33/30/33/32/33/33/33/30/31/32/31/3
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect

Model Accuracy vs Pass@3

100%
75%
50%
25%
0%
GPT-5.2 (high)
Gemini-3-Pro-Preview
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
EXAONE-4.0.1-32B (high)
EXAONE-4.0.1-32B (high)
Kanana-2-30B-Thinking-2601
Kanana-2-30B-Thinking-2601
Accuracy
Pass@3

Avg Token Usage (Per Problem)

20,071
15,053
10,035
5,018
0
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
Gemini-3-Pro-Preview
Kanana-2-30B-Thinking-2601
Kanana-2-30B-Thinking-2601
GPT-5.2 (high)
EXAONE-4.0.1-32B (high)
EXAONE-4.0.1-32B (high)
Avg Tokens / Problem
About this benchmark

The 2025 CSAT Korean language subject: reading comprehension and verbal reasoning, not mathematics.

Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language. This CSAT Korean language evaluation measures reading comprehension and verbal reasoning. It is not a Korean translation of the mathematics evaluation.

Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.

Performance Legend

Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3