EntropyMath Leaderboard

CSAT 2026

46 problems 4 models

K-LLM Project · 독파모 / EntropyMath 평가 차수 안내

K-LLM Project는 대한민국 과학기술정보통신부의 독자 AI 파운데이션 모델 프로젝트(통칭 독파모)을 뜻합니다.

Round 1·2·3은 EntropyMath의 1차·2차·3차 모델 평가를 뜻합니다. 과기정통부의 공식 단계평가 번호나 통과 여부를 뜻하지 않습니다.

프로젝트 공식 안내 공식 1차 단계평가 공식 2차 단계평가

Jump to leaderboard
Evaluation details · 평가 방법 Python TAR · ACC / Pass@K

2026 수능 수학 한국어 46문항(공통 22 · 확률과 통계 8 · 미적분 8 · 기하 8)을 원문 텍스트와 그림 설명으로 평가합니다. 세 선택과목을 합친 문항 정답률이며, 수능 배점 기준 100점이 아닙니다.

모델이 Python으로 계산·검산하고, 실행 결과를 다음 턴에서 읽어 풀이와 최종 답변을 이어가는 다중 턴 TAR 평가입니다. Python 사용을 지시하며, 실제 호출·실행 성공 여부를 별도로 기록합니다. 호출마다 Python 변수는 초기화되고 실행 제한은 20초입니다.

ACC · Pass@1
최종 boxed 정수(객관식은 선택지 번호)를 정답 키와 비교합니다. ACC는 전체 집계 시도의 정답률, Pass@1은 문항별 정답률의 평균이며, 문항당 시도 수가 같으면 두 값은 같습니다.
Pass@K
K회 안에 한 번 이상 정답을 낼 확률의 추정치입니다. K회 이상 집계된 문항만 포함하고 대상 문항 수를 함께 표시합니다. 1회 평가 모델의 Pass@3는 —로 표시합니다.
Answer + Exec
최종 답이 정답이고 Python 실행이 한 번 이상 성공한 시도 수입니다. 문제 셀의 색은 정답 여부, 아래 검은 표시의 위·아래 행은 회차별 Python 호출·실행 성공을 뜻합니다.

모델별 실행 조건

모델문항당 횟수호출 경로Reasoning출력 토큰 한도요청 / 문항 제한최대 턴
Muse Spark 1.31OpenRouter → MetaOn · 기본값32,768600 / 1,200초6
Solar Mini 4 Preview3Upstage APIHigh131,072600 / 1,800초6 + 1
DeepSeek V4.1 Flash1OpenRouter → DeepSeekHigh131,0721,800 / 1,800초6 + 1
Solar Pro 43OpenRouter → UpstageOn · 기본값131,072600 / 1,800초6 + 1

출력 토큰 한도는 문항 전체 누적값입니다. 6+1은 도구 사용 가능 최대 6턴 뒤 도구 없는 최종 답변 1턴입니다. Temperature·top_p는 공급자 기본값이며, 조건 차이와 개별 검토·복구 이력은 모델 옆 ⓘ에서 확인할 수 있습니다.

통신 오류(4xx·5xx 등)와 시간·출력 제한은 해당 실행의 복구 정책에 따라 재시도하며 원본을 보존합니다. 해결되지 않아 보류한 실행은 집계에서 제외하고, 답 추출 실패와 Python 조건 미충족은 별도로 기록합니다. 정답 일치·실행 성공·수학적 증명 검증은 서로 다릅니다.

Leaderboard

Benchmarks ↑
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect
Color = answer correctnessPython: top = called · bottom = execution succeeded× no successful execution · — not called · ? not reportedExecution success is not proof verification.

Answer accuracy · Korean Python TAR

100%
75%
50%
25%
0%
Muse Spark 1.3
Muse Spark 1.3
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash
Solar Pro 4
Solar Pro 4
Solar Mini 4 Preview
Solar Mini 4 Preview
Accuracy
Pass@3

Avg Tokens per Reported Run

15.5K
11.6K
7.8K
3.9K
0
Solar Pro 4
Solar Pro 4
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash
Solar Mini 4 Preview
Solar Mini 4 Preview
Muse Spark 1.3
Muse Spark 1.3
Avg Tokens / Problem
About this benchmark

Korean language measures reading comprehension and verbal reasoning. Mathematics evaluates episodic reasoning with Python TAR.

Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language.

평가 범위, 점수 정의와 모델별 실행 조건은 상단의 평가 방법에서 확인할 수 있습니다. Detailed execution traces are available for transparency. Partial-coverage Pass@K is a subset result, not a full 46-item score. Estimator: HumanEval reference implementation.

Performance Legend

Correct / scored attempts · Answer matches the key
Attempt counts vary by model; deferred records are excluded.
/ · Deferred Python episode, excluded from scoring
Answer correctness is distinct from proof verification.
46 pooled items across all electives; not a weighted 100-point exam score.
    EntropyMath Leaderboard