EntropyMath Leaderboard
Evaluation details · 평가 방법 Answer selection · ACC / Pass@1 / Pass@3 · No tools
- 객관식 33문항을 문항당 3회 평가했습니다. OpenRouter Decisions → TypeSafe의 답 선택 결과이며, 풀이 생성·Python 실행·TAR 평가는 아닙니다. 주관식 13문항은 제외했습니다. 회차별 정답: 17/33 → 15/33 → 16/33.
- ACC·평균 Pass@1 = 48/99 (48.5%). Pass@3 = 세 번 중 한 번 이상 정답인 18/33문항 (54.5%). 공통15·확통6·미적분6·기하6문항의 비가중 정답률이며 수능 100점 환산 점수가 아닙니다.
- 동일 입력을 별도 호출했습니다. 오류·재시도 없이 99회 완료했습니다. 선택지 확률과 confidence는 공급자 반환값이며 정답 보장이나 증명 검증을 뜻하지 않습니다.
Leaderboard
Benchmarks ↑Scroll for problems · Click a cell for the solution
All correctPartialIncorrect
Answer selection · ACC / Pass@3 · 33 multiple-choice items
100%
75%
50%
25%
0%
Jev 1.13
Accuracy
Pass@3
Avg Tokens per Reported Run
652
489
326
163
0
Jev 1.13
Avg Tokens / Problem
About this benchmark
Korean language measures verbal reasoning. Mathematics separates Python TAR from multiple-choice answer selection.
Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language.
객관식 33문항을 각 3회 평가한 답 선택 결과입니다. Pass@1은 세 회차 평균, Pass@3는 세 번 중 한 번 이상 정답인 문항 비율입니다. Python TAR 결과와 직접 비교하지 않습니다. 문항별 선택지와 반환 확률을 공개합니다. 풀이·코드는 생성하지 않았습니다.
Performance Legend
Correct / 3 attempts · Color = answer correctness
33 multiple-choice items · Three decisions per item · No Python