EntropyMath Leaderboard

Open10 problems

10 problems 7 models

K-LLM Project · 독파모 / EntropyMath 평가 차수 안내

K-LLM Project는 대한민국 과학기술정보통신부의 독자 AI 파운데이션 모델 프로젝트(통칭 독파모)을 뜻합니다.

Round 1·2·3은 EntropyMath의 1차·2차·3차 모델 평가를 뜻합니다. 과기정통부의 공식 단계평가 번호나 통과 여부를 뜻하지 않습니다.

프로젝트 공식 안내 공식 1차 단계평가 공식 2차 단계평가

Jump to leaderboard

Leaderboard

Benchmarks ↑
ModelAccPass@30123456789
API / Others
Deepseek-V3.2
100.0100.03/33/33/33/33/33/33/33/33/33/3
Grok-4.1-fast
93.3100.03/33/32/33/33/33/33/33/32/33/3
GPT-5.1 (high)
93.3100.03/33/33/33/33/32/33/33/32/33/3
Claude-Opus-4.5
90.0100.03/33/33/33/33/33/32/33/31/33/3
Gemini-3-Pro-Preview
90.0100.03/32/32/33/33/33/32/33/33/33/3
K-LLM Project · 2차 평가 (Round 2)
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
86.7100.02/33/32/33/33/32/32/33/33/33/3
Solar-Open-100B
Solar-Open-100B
66.780.02/32/30/33/33/33/32/33/30/32/3
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect

Model Accuracy vs Pass@3

100%
75%
50%
25%
0%
Deepseek-V3.2
Grok-4.1-fast
GPT-5.1 (high)
Claude-Opus-4.5
Gemini-3-Pro-Preview
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
Solar-Open-100B
Solar-Open-100B
Accuracy
Pass@3

Avg Token Usage (Per Problem)

23,407
17,556
11,704
5,852
0
Solar-Open-100B
Solar-Open-100B
Grok-4.1-fast
Gemini-3-Pro-Preview
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
Deepseek-V3.2
Claude-Opus-4.5
GPT-5.1 (high)
Avg Tokens / Problem
About this benchmark

Derivative problems that test generalization beyond the original seed set.

EntropyMath is an evolutionary multi-agent system and benchmark that generates high-entropy math problems designed to systematically break current LLMs. The EntropyMath_Open_10 benchmark contains derivative problems generated from the seed set. This dataset tests the model's robustness and generalization ability by presenting variations of known problem distributions, ensuring that performance is not merely due to memorization.

Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.

Performance Legend

Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3