EntropyMath Leaderboard

Evaluation details · 평가 방법 Answer selection · ACC / Pass@1 / Pass@3 · No tools
  • 객관식 33문항을 문항당 3회 평가했습니다. OpenRouter Decisions → TypeSafe의 답 선택 결과이며, 풀이 생성·Python 실행·TAR 평가는 아닙니다. 주관식 13문항은 제외했습니다. 회차별 정답: 17/33 → 15/33 → 16/33.
  • ACC·평균 Pass@1 = 48/99 (48.5%). Pass@3 = 세 번 중 한 번 이상 정답인 18/33문항 (54.5%). 공통15·확통6·미적분6·기하6문항의 비가중 정답률이며 수능 100점 환산 점수가 아닙니다.
  • 동일 입력을 별도 호출했습니다. 오류·재시도 없이 99회 완료했습니다. 선택지 확률과 confidence는 공급자 반환값이며 정답 보장이나 증명 검증을 뜻하지 않습니다.

Leaderboard

Benchmarks ↑
ModelACCPass@1PASS@3Common / 공통Probability / 확통Calculus / 미적분Geometry / 기하
123456789101112131415232425262728232425262728232425262728
Jev 1.13
Jev 1.13
48.5%48/9948.5%54.5%18/330/33/33/30/30/30/30/30/33/30/31/33/30/31/31/33/30/30/30/30/33/33/33/30/30/30/33/33/33/33/33/33/33/3
Scroll for problems · Click a cell for the solution
All correctPartialIncorrect

Answer selection · ACC / Pass@3 · 33 multiple-choice items

100%
75%
50%
25%
0%
Jev 1.13
Jev 1.13
Accuracy
Pass@3

Avg Tokens per Reported Run

652
489
326
163
0
Jev 1.13
Jev 1.13
Avg Tokens / Problem
About this benchmark

Korean language measures verbal reasoning. Mathematics separates Python TAR from multiple-choice answer selection.

Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language.

객관식 33문항을 각 3회 평가한 답 선택 결과입니다. Pass@1은 세 회차 평균, Pass@3는 세 번 중 한 번 이상 정답인 문항 비율입니다. Python TAR 결과와 직접 비교하지 않습니다. 문항별 선택지와 반환 확률을 공개합니다. 풀이·코드는 생성하지 않았습니다.

Performance Legend

Correct / 3 attempts · Color = answer correctness
33 multiple-choice items · Three decisions per item · No Python