EntropyMath Leaderboard
Leaderboard
Benchmarks ↑| Model | Acc | Pass@3 | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|---|---|
| API / Others | |||||||
Gemini-3-Pro-Preview | 60.0 | 80.0 | 2/3 | 3/3 | 3/3 | 0/3 | 1/3 |
GPT-5.2 (high) | 33.3 | 60.0 | 3/3 | 1/3 | 1/3 | 0/3 | 0/3 |
| K-LLM Project · 2차 평가 (Round 2) | |||||||
Solar Pro 3 | 40.0 | 40.0 | 3/3 | 0/3 | 3/3 | 0/3 | 0/3 |
K-EXAONE-236B-A23B | 13.3 | 40.0 | 0/3 | 1/3 | 0/3 | 1/3 | 0/3 |
K-EXAONE-236B-A23B | 6.7 | 20.0 | 0/3 | 1/3 | 0/3 | 0/3 | 0/3 |
| Local - KR | |||||||
Kanana-2-30B-Thinking-2601 | 13.3 | 40.0 | 1/3 | 0/3 | 1/3 | 0/3 | 0/3 |
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect
Model Accuracy vs Pass@3
100%
75%
50%
25%
0%
Gemini-3-Pro-Preview
Solar Pro 3
GPT-5.2 (high)
K-EXAONE-236B-A23B
Kanana-2-30B-Thinking-2601
K-EXAONE-236B-A23B
Accuracy
Pass@3
Avg Token Usage (Per Problem)
20,777
15,583
10,388
5,194
0
K-EXAONE-236B-A23B
K-EXAONE-236B-A23B
Solar Pro 3
Gemini-3-Pro-Preview
Kanana-2-30B-Thinking-2601
GPT-5.2 (high)
Avg Tokens / Problem
About this benchmark
The JWL seed evaluation, kept separate from the standard EntropyMath seed sets.
EntropyMath is an evolutionary multi-agent system and benchmark that generates high-entropy math problems designed to systematically break current LLMs.
Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.
Performance Legend
Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3