| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| MiMo-V2.6-FlashXiaomi | 1780.2 | $0.001234 | 85.1 | 28 / 0 |
| MiMo V2.6 ProXiaomi | 1728.8 | $0.003644 | 77.2 | 26 / 2 |
| Claude Opus 5.5Anthropic | 1721.8 | $0.058136 | 59.1 | 25 / 3 |
| Muse Spark 1.3 ContributorMeta | 1682.0 | $0.000733 | 79.8 | 24 / 4 |
| Claude Fable 5.1Anthropic | 1680.9 | $0.267990 | 52.8 | 23 / 5 |
| GLM 5.3 FlashZ.ai | 1668.4 | $0.001188 | 77.6 | 23 / 5 |
| GLM 5.3 PrimeZ.ai | 1665.6 | $0.099276 | 53.3 | 23 / 5 |
| DeepSeek V4.1 FlashDeepSeek | 1654.2 | $0.005439 | 69.0 | 22 / 6 |
| Hy4 previewTencent | 1635.2 | $0.021153 | 57.6 | 20 / 8 |
| Kimi K3Moonshot AI | 1618.7 | $0.058935 | 50.9 | 20 / 8 |
| Space Bunny AlphaStealth | 1568.4 | $0.000000 | 71.8 | 18 / 10 |
| GPT-6 AstraOpenAI | 1549.8 | $0.075850 | 43.5 | 16 / 12 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1536.1 | $0.150232 | 40.5 | 16 / 12 |
| DeepSeek V4 ProDeepSeek | 1504.0 | $0.004142 | 56.6 | 14 / 14 |
| Nemotron 3 Ultra (free)NVIDIA | 1495.6 | $0.000000 | 64.6 | 13 / 15 |
| Grok 4.7xAI | 1482.7 | $0.009816 | 48.4 | 14 / 14 |
| GPT-6 LunaOpenAI | 1462.0 | $0.001016 | 58.4 | 12 / 16 |
| Hy3Tencent | 1425.5 | $0.003099 | 50.5 | 10 / 18 |
| GPT-6 SolOpenAI | 1412.5 | $0.016770 | 37.6 | 9 / 19 |
| gpt-oss-120bOpenAI | 1412.4 | $0.000486 | 55.0 | 10 / 18 |
| Gemini 3.8 FlashGoogle | 1404.9 | $0.011276 | 39.8 | 9 / 19 |
| Ling 3.0 FlashinclusionAI | 1393.3 | $0.000257 | 53.8 | 8 / 20 |
| MiniMax M2.7 (Nitro)MiniMax | 1372.4 | $0.005518 | 42.0 | 7 / 21 |
| Qwen3.7 FlashQwen | 1322.4 | $0.000578 | 46.9 | 5 / 23 |
| Solar Pro 4Upstage | 1296.2 | $0.000567 | 44.9 | 4 / 24 |
| Gemini 3.1 Pro PreviewGoogle | 1292.0 | $0.039222 | 22.3 | 4 / 24 |
| Mistral Medium 3.5Mistral | 1257.4 | $0.014400 | 26.2 | 2 / 26 |
| Mercury 2.5Inception | 1247.0 | $0.000434 | 42.0 | 1 / 27 |
| gpt-oss-20b (Nitro)OpenAI | 1229.9 | $0.000699 | 40.2 | 0 / 28 |
Jev Benchmark
Compare model quality, cost, and the answers behind the scores.
Answers from OpenRouter, compared by TypeSafe's Jev. A small, fully inspectable preference experiment. Not a claim of general intelligence or verified accuracy.
Read the launch study: 29 models, 1,624 comparisons, and what stood out.
- Models
- 29
- Questions
- 6
- Answers collected
- 174 / 174
- Pairs judged
- 2436 / 2436
Overall ranking
Weighted mean of per-question Elo. A model receives an overall score only after completing every matchup across all active questions.
| Model | Weighted Elo | Mean answer cost | Value / 100 |
|---|---|---|---|
| MiMo-V2.6-FlashXiaomi | 1699.3 | $0.002860 | 76.5 |
| Claude Fable 5.1Anthropic | 1682.8 | $0.226293 | 53.2 |
| MiMo V2.6 ProXiaomi | 1681.2 | $0.010550 | 66.4 |
| Claude Opus 5.5Anthropic | 1663.9 | $0.076747 | 53.8 |
| GLM 5.3 PrimeZ.ai | 1647.8 | $0.165112 | 50.8 |
| GLM 5.3 FlashZ.ai | 1608.0 | $0.004392 | 66.4 |
| GPT-6 AstraOpenAI | 1590.5 | $0.096045 | 46.7 |
| Muse Spark 1.3 ContributorMeta | 1587.6 | $0.001163 | 70.5 |
| Hy4 previewTencent | 1581.5 | $0.027925 | 51.0 |
| Kimi K3Moonshot AI | 1562.5 | $0.082261 | 44.5 |
| DeepSeek V4.1 FlashDeepSeek | 1561.4 | $0.006077 | 59.8 |
| Space Bunny AlphaStealth | 1549.6 | $0.000000 | 70.0 |
| MiniMax M2.7 (Nitro)MiniMax | 1547.4 | $0.012704 | 53.0 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1534.8 | $0.146123 | 40.4 |
| Nemotron 3 Ultra (free)NVIDIA | 1513.8 | $0.000000 | 66.4 |
| gpt-oss-120bOpenAI | 1481.1 | $0.000647 | 61.3 |
| Grok 4.7xAI | 1477.7 | $0.036554 | 39.2 |
| Hy3Tencent | 1472.0 | $0.005701 | 51.3 |
| Gemini 3.8 FlashGoogle | 1461.1 | $0.012417 | 44.5 |
| DeepSeek V4 ProDeepSeek | 1430.1 | $0.005675 | 47.2 |
| GPT-6 LunaOpenAI | 1423.6 | $0.001223 | 54.2 |
| GPT-6 SolOpenAI | 1401.9 | $0.016871 | 36.5 |
| Solar Pro 4Upstage | 1397.7 | $0.001248 | 51.7 |
| Qwen3.7 FlashQwen | 1375.5 | $0.001225 | 49.7 |
| gpt-oss-20b (Nitro)OpenAI | 1369.9 | $0.001002 | 49.7 |
| Ling 3.0 FlashinclusionAI | 1354.2 | $0.000382 | 50.0 |
| Gemini 3.1 Pro PreviewGoogle | 1312.1 | $0.090838 | 20.7 |
| Mercury 2.5Inception | 1290.5 | $0.000450 | 44.8 |
| Mistral Medium 3.5Mistral | 1240.8 | $0.015811 | 24.5 |
Value v1 blends 70% Elo-derived quality with 30% affordability. Higher is better under this chosen tradeoff. See the formula and limitations. Mean cost uses the same question weights as Elo; unknown costs are never treated as free.
Costs are actual USD reported by OpenRouter for saved successful answers, including provider-reported reasoning usage. Zero means reported free; Unknown means no cost was reported. These totals exclude failed/retried attempts and Jev charges, and are not an account invoice. Attempts absent from retained responses and backups remain unknown.
Latest judgment: Sep 28, 2026, 17:58 UTC.
Quality and cost, together
Higher is stronger in Jev's rankings. Further left is cheaper. Solid points mark the observed frontier: no other model is at least as strong and at least as cheap, with one strict improvement. Select a point for its values.
One question at a time
| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| MiMo V2.6 ProXiaomi | 1721.7 | $0.029694 | 62.3 | 25 / 3 |
| MiMo-V2.6-FlashXiaomi | 1710.7 | $0.007339 | 71.3 | 25 / 3 |
| MiniMax M2.7 (Nitro)MiniMax | 1708.4 | $0.025938 | 62.1 | 24 / 4 |
| GLM 5.3 PrimeZ.ai | 1702.8 | $0.286540 | 54.4 | 24 / 4 |
| GLM 5.3 FlashZ.ai | 1663.0 | $0.011422 | 64.3 | 22 / 6 |
| GPT-6 AstraOpenAI | 1642.7 | $0.184510 | 50.2 | 22 / 6 |
| gpt-oss-20b (Nitro)OpenAI | 1632.1 | $0.002240 | 72.2 | 20 / 8 |
| gpt-oss-120bOpenAI | 1619.1 | $0.001453 | 72.7 | 21 / 7 |
| Hy4 previewTencent | 1615.4 | $0.055782 | 50.8 | 20 / 8 |
| Claude Fable 5.1Anthropic | 1607.6 | $0.394580 | 46.2 | 19 / 9 |
| Gemini 3.8 FlashGoogle | 1583.3 | $0.023620 | 52.2 | 20 / 8 |
| Space Bunny AlphaStealth | 1563.1 | $0.000000 | 71.3 | 18 / 10 |
| Nemotron 3 Ultra (free)NVIDIA | 1541.2 | $0.000000 | 69.1 | 16 / 12 |
| Claude Opus 5.5Anthropic | 1537.5 | $0.129412 | 40.9 | 16 / 12 |
| GPT-6 LunaOpenAI | 1502.4 | $0.002351 | 59.5 | 15 / 13 |
| Hy3Tencent | 1475.8 | $0.011622 | 46.4 | 12 / 16 |
| Kimi K3Moonshot AI | 1470.1 | $0.241173 | 33.2 | 11 / 17 |
| Gemini 3.1 Pro PreviewGoogle | 1433.0 | $0.242660 | 29.5 | 9 / 19 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1415.5 | $0.388132 | 27.4 | 10 / 18 |
| Mercury 2.5Inception | 1400.1 | $0.000798 | 53.0 | 9 / 19 |
| Qwen3.7 FlashQwen | 1396.5 | $0.001662 | 50.6 | 9 / 19 |
| Muse Spark 1.3 ContributorMeta | 1394.5 | $0.001436 | 50.9 | 8 / 20 |
| GPT-6 SolOpenAI | 1379.2 | $0.030462 | 30.7 | 8 / 20 |
| Grok 4.7xAI | 1360.0 | $0.065302 | 25.6 | 7 / 21 |
| DeepSeek V4.1 FlashDeepSeek | 1356.6 | $0.012721 | 34.5 | 6 / 22 |
| Ling 3.0 FlashinclusionAI | 1306.6 | $0.000597 | 45.6 | 5 / 23 |
| Solar Pro 4Upstage | 1274.6 | $0.001374 | 41.4 | 3 / 25 |
| DeepSeek V4 ProDeepSeek | 1265.3 | $0.004441 | 35.2 | 2 / 26 |
| Mistral Medium 3.5Mistral | 1221.3 | $0.015977 | 23.3 | 0 / 28 |
| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| GLM 5.3 PrimeZ.ai | 1782.9 | $0.532690 | 59.1 | 28 / 0 |
| MiMo-V2.6-FlashXiaomi | 1726.0 | $0.006008 | 73.8 | 26 / 2 |
| Muse Spark 1.3 ContributorMeta | 1721.4 | $0.003476 | 77.0 | 25 / 3 |
| MiMo V2.6 ProXiaomi | 1717.5 | $0.020503 | 64.3 | 25 / 3 |
| Claude Fable 5.1Anthropic | 1694.5 | $0.387750 | 53.5 | 24 / 4 |
| Claude Opus 5.5Anthropic | 1687.5 | $0.128620 | 54.4 | 23 / 5 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1594.1 | $0.265960 | 45.3 | 19 / 9 |
| MiniMax M2.7 (Nitro)MiniMax | 1593.3 | $0.029806 | 51.7 | 18 / 10 |
| gpt-oss-120bOpenAI | 1577.3 | $0.000937 | 70.1 | 18 / 10 |
| gpt-oss-20b (Nitro)OpenAI | 1564.9 | $0.002158 | 66.1 | 17 / 11 |
| Space Bunny AlphaStealth | 1563.9 | $0.000000 | 71.4 | 17 / 11 |
| Gemini 3.8 FlashGoogle | 1560.5 | $0.022731 | 50.2 | 18 / 10 |
| GPT-6 AstraOpenAI | 1549.0 | $0.160860 | 41.7 | 16 / 12 |
| Hy4 previewTencent | 1536.9 | $0.055824 | 43.3 | 16 / 12 |
| Kimi K3Moonshot AI | 1521.8 | $0.114538 | 39.6 | 15 / 13 |
| Hy3Tencent | 1518.3 | $0.015662 | 48.5 | 16 / 12 |
| GLM 5.3 FlashZ.ai | 1503.6 | $0.011531 | 49.3 | 15 / 13 |
| Nemotron 3 Ultra (free)NVIDIA | 1490.6 | $0.000000 | 64.1 | 14 / 14 |
| Ling 3.0 FlashinclusionAI | 1439.9 | $0.000866 | 56.6 | 10 / 18 |
| GPT-6 SolOpenAI | 1375.8 | $0.027982 | 30.9 | 7 / 21 |
| GPT-6 LunaOpenAI | 1372.5 | $0.002207 | 47.3 | 8 / 20 |
| Solar Pro 4Upstage | 1361.0 | $0.001447 | 47.9 | 7 / 21 |
| Gemini 3.1 Pro PreviewGoogle | 1355.1 | $0.182064 | 22.8 | 7 / 21 |
| DeepSeek V4.1 FlashDeepSeek | 1352.7 | $0.011758 | 34.8 | 6 / 22 |
| DeepSeek V4 ProDeepSeek | 1324.4 | $0.004741 | 39.0 | 5 / 23 |
| Grok 4.7xAI | 1277.7 | $0.120054 | 17.5 | 3 / 25 |
| Mercury 2.5Inception | 1266.0 | $0.000709 | 42.5 | 2 / 26 |
| Qwen3.7 FlashQwen | 1243.0 | $0.003241 | 35.6 | 1 / 27 |
| Mistral Medium 3.5Mistral | 1227.8 | $0.047226 | 17.3 | 0 / 28 |
| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| GLM 5.3 PrimeZ.ai | 1724.4 | $0.050310 | 59.9 | 26 / 2 |
| Claude Fable 5.1Anthropic | 1724.2 | $0.081930 | 58.2 | 25 / 3 |
| Muse Spark 1.3 ContributorMeta | 1723.9 | $0.000505 | 83.4 | 26 / 2 |
| GPT-6 AstraOpenAI | 1710.5 | $0.049070 | 59.0 | 24 / 4 |
| Kimi K3Moonshot AI | 1699.8 | $0.050908 | 58.1 | 23 / 5 |
| Claude Opus 5.5Anthropic | 1666.0 | $0.035472 | 57.2 | 23 / 5 |
| DeepSeek V4.1 FlashDeepSeek | 1653.3 | $0.002109 | 74.3 | 22 / 6 |
| GLM 5.3 FlashZ.ai | 1635.3 | $0.000877 | 75.6 | 20 / 8 |
| Hy4 previewTencent | 1633.9 | $0.013976 | 60.4 | 21 / 7 |
| MiMo V2.6 ProXiaomi | 1625.6 | $0.001057 | 74.3 | 21 / 7 |
| Space Bunny AlphaStealth | 1611.9 | $0.000000 | 75.9 | 20 / 8 |
| Solar Pro 4Upstage | 1558.7 | $0.000387 | 69.7 | 18 / 10 |
| MiMo-V2.6-FlashXiaomi | 1529.8 | $0.000173 | 67.5 | 15 / 13 |
| Nemotron 3 Ultra (free)NVIDIA | 1513.2 | $0.000000 | 66.3 | 15 / 13 |
| gpt-oss-120bOpenAI | 1489.7 | $0.000298 | 63.1 | 14 / 14 |
| MiniMax M2.7 (Nitro)MiniMax | 1466.0 | $0.003556 | 53.7 | 12 / 16 |
| GPT-6 LunaOpenAI | 1454.8 | $0.000593 | 58.8 | 11 / 17 |
| DeepSeek V4 ProDeepSeek | 1436.0 | $0.001896 | 53.8 | 10 / 18 |
| Hy3Tencent | 1416.6 | $0.002122 | 51.5 | 9 / 19 |
| GPT-6 SolOpenAI | 1409.3 | $0.008394 | 42.4 | 10 / 18 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1408.0 | $0.023564 | 34.9 | 9 / 19 |
| Grok 4.7xAI | 1386.2 | $0.006072 | 42.6 | 9 / 19 |
| Gemini 3.8 FlashGoogle | 1384.0 | $0.007327 | 41.0 | 8 / 20 |
| Qwen3.7 FlashQwen | 1334.8 | $0.000313 | 48.6 | 5 / 23 |
| Gemini 3.1 Pro PreviewGoogle | 1293.3 | $0.028020 | 24.2 | 4 / 24 |
| gpt-oss-20b (Nitro)OpenAI | 1277.6 | $0.000409 | 44.0 | 3 / 25 |
| Mistral Medium 3.5Mistral | 1260.5 | $0.005984 | 32.9 | 2 / 26 |
| Mercury 2.5Inception | 1250.6 | $0.000401 | 42.3 | 1 / 27 |
| Ling 3.0 FlashinclusionAI | 1221.9 | $0.000157 | 41.3 | 0 / 28 |
| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| Grok 4.7xAI | 1756.5 | $0.012107 | 70.6 | 27 / 1 |
| MiMo-V2.6-FlashXiaomi | 1731.6 | $0.001612 | 81.2 | 26 / 2 |
| MiMo V2.6 ProXiaomi | 1707.1 | $0.005753 | 72.7 | 24 / 4 |
| DeepSeek V4.1 FlashDeepSeek | 1690.0 | $0.002854 | 75.8 | 24 / 4 |
| Claude Fable 5.1Anthropic | 1686.7 | $0.142860 | 54.1 | 23 / 5 |
| Muse Spark 1.3 ContributorMeta | 1637.0 | $0.000528 | 76.6 | 22 / 6 |
| Nemotron 3 Ultra (free)NVIDIA | 1633.4 | $0.000000 | 77.8 | 20 / 8 |
| GPT-6 AstraOpenAI | 1618.7 | $0.061170 | 50.7 | 20 / 8 |
| Claude Opus 5.5Anthropic | 1613.6 | $0.065984 | 50.0 | 20 / 8 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1613.4 | $0.030976 | 53.4 | 19 / 9 |
| GLM 5.3 FlashZ.ai | 1572.6 | $0.000817 | 69.9 | 19 / 9 |
| Hy4 previewTencent | 1549.2 | $0.012338 | 53.4 | 17 / 11 |
| Kimi K3Moonshot AI | 1544.9 | $0.013937 | 52.0 | 16 / 12 |
| GLM 5.3 PrimeZ.ai | 1504.7 | $0.012455 | 48.8 | 15 / 13 |
| Hy3Tencent | 1499.6 | $0.000939 | 62.4 | 14 / 14 |
| Gemini 3.8 FlashGoogle | 1496.6 | $0.005426 | 54.1 | 14 / 14 |
| Space Bunny AlphaStealth | 1472.5 | $0.000000 | 62.2 | 12 / 16 |
| GPT-6 SolOpenAI | 1468.1 | $0.010014 | 46.8 | 12 / 16 |
| MiniMax M2.7 (Nitro)MiniMax | 1450.2 | $0.005990 | 48.8 | 11 / 17 |
| DeepSeek V4 ProDeepSeek | 1415.8 | $0.013057 | 39.7 | 9 / 19 |
| Qwen3.7 FlashQwen | 1406.7 | $0.000725 | 53.8 | 9 / 19 |
| gpt-oss-120bOpenAI | 1388.1 | $0.000493 | 52.7 | 9 / 19 |
| GPT-6 LunaOpenAI | 1385.8 | $0.000700 | 51.9 | 8 / 20 |
| Mercury 2.5Inception | 1318.2 | $0.000194 | 47.6 | 5 / 23 |
| Ling 3.0 FlashinclusionAI | 1310.7 | $0.000259 | 46.9 | 4 / 24 |
| Solar Pro 4Upstage | 1293.2 | $0.001797 | 41.8 | 4 / 24 |
| Gemini 3.1 Pro PreviewGoogle | 1259.3 | $0.026318 | 22.3 | 2 / 26 |
| Mistral Medium 3.5Mistral | 1246.7 | $0.006441 | 31.5 | 1 / 27 |
| gpt-oss-20b (Nitro)OpenAI | 1229.2 | $0.000298 | 41.3 | 0 / 28 |
| Model | Elo | Answer cost | Value / 100 | W / L |
|---|---|---|---|---|
| Claude Opus 5.5Anthropic | 1757.3 | $0.042860 | 62.7 | 27 / 1 |
| MiMo-V2.6-FlashXiaomi | 1717.3 | $0.000793 | 82.2 | 25 / 3 |
| Claude Fable 5.1Anthropic | 1702.6 | $0.082650 | 56.6 | 24 / 4 |
| MiniMax M2.7 (Nitro)MiniMax | 1694.2 | $0.005416 | 72.2 | 24 / 4 |
| DeepSeek V4.1 FlashDeepSeek | 1661.7 | $0.001583 | 76.1 | 23 / 5 |
| Qwen3.8 Max PrimeAlibaba / Qwen | 1641.5 | $0.017872 | 59.3 | 22 / 6 |
| DeepSeek V4 ProDeepSeek | 1635.2 | $0.005775 | 67.0 | 20 / 8 |
| GLM 5.3 FlashZ.ai | 1605.0 | $0.000521 | 73.8 | 19 / 9 |
| Grok 4.7xAI | 1603.0 | $0.005973 | 63.9 | 20 / 8 |
| Solar Pro 4Upstage | 1602.3 | $0.001917 | 70.2 | 20 / 8 |
| MiMo V2.6 ProXiaomi | 1586.4 | $0.002650 | 67.2 | 19 / 9 |
| Qwen3.7 FlashQwen | 1549.5 | $0.000829 | 67.7 | 15 / 13 |
| Kimi K3Moonshot AI | 1519.7 | $0.014076 | 49.4 | 15 / 13 |
| Hy4 previewTencent | 1518.2 | $0.008479 | 53.1 | 14 / 14 |
| Space Bunny AlphaStealth | 1517.9 | $0.000000 | 66.8 | 15 / 13 |
| GLM 5.3 PrimeZ.ai | 1506.3 | $0.009402 | 51.1 | 14 / 14 |
| Hy3Tencent | 1496.0 | $0.000765 | 62.5 | 14 / 14 |
| GPT-6 AstraOpenAI | 1472.5 | $0.044810 | 37.7 | 13 / 15 |
| Ling 3.0 FlashinclusionAI | 1452.6 | $0.000155 | 59.8 | 12 / 16 |
| Nemotron 3 Ultra (free)NVIDIA | 1408.7 | $0.000000 | 56.0 | 9 / 19 |
| gpt-oss-120bOpenAI | 1399.7 | $0.000218 | 54.5 | 9 / 19 |
| Muse Spark 1.3 ContributorMeta | 1367.2 | $0.000302 | 51.4 | 7 / 21 |
| GPT-6 SolOpenAI | 1366.4 | $0.007602 | 39.2 | 7 / 21 |
| GPT-6 LunaOpenAI | 1364.3 | $0.000475 | 50.6 | 7 / 21 |
| Gemini 3.8 FlashGoogle | 1337.1 | $0.004123 | 40.9 | 6 / 22 |
| gpt-oss-20b (Nitro)OpenAI | 1285.5 | $0.000210 | 45.2 | 3 / 25 |
| Mercury 2.5Inception | 1261.0 | $0.000164 | 43.6 | 2 / 26 |
| Gemini 3.1 Pro PreviewGoogle | 1239.7 | $0.026746 | 21.0 | 1 / 27 |
| Mistral Medium 3.5Mistral | 1231.1 | $0.004838 | 32.5 | 0 / 28 |
Methodology
Value score v1
Quality q = 1 / (1 + 10((1500 − Elo) / 400)). Affordability a = 1 / (1 + cost / $0.01). Value = 100 × (0.7q + 0.3a).
The 70/30 blend and one-cent reference are explicit design choices, not learned truths. Quality is an Elo-implied preference against a 1,500-rated reference, not factual accuracy. Free answers receive finite scores; unknown costs and incomplete scopes are not scored. Overall cost is the question-weighted mean per saved answer. Free-tier availability and prices can change.
One answer. Every opponent.
Each model answers each question once, without tools or web access. The same prompt goes to every model; provider-default sampling settings apply. Reasoning effort can be explicitly configured on a new question version; existing inputs are frozen after use. Answers start with each model's configured output budget. Failed, truncated answers may be retried with an explicitly raised allowance, recorded on each answer. Truncated or failed answers are not scored.
Each unordered pair is judged once. Model labels are omitted from the judge's input, and A/B presentation is assigned by a stable hash. Answers may still reveal their origin. Probabilities and confidence are saved alongside each decision.
Reproducible Elo
Every question starts all models at 1,500. Each stored win/loss is replayed once with K = 32 in a stable, hashed model-ID order. That makes the same dataset reproducible, not order-independent. Adding opponents can change existing ratings without rejudging their old matches.
Question weights default to one. The overall score is the weighted arithmetic mean of complete per-question ratings. Pending models have no overall rank; per-question results remain provisional until the round robin is complete.
A preference experiment
Jev selects a winner, not a verified truth. This small sample measures one judge's preferences on these prompts, not general model ability. Code is not executed and mathematical proofs are not independently verified. A single presentation per pair does not eliminate position bias.
Confidence describes Jev's decision distribution; it is not the probability that the answer is factually correct. No repeated trials, error bars, or statistical-significance claims are implied.
Use this as a source
The full JSON dataset includes prompts, rubrics, original answers, generation settings, costs, attempt history, and every pairwise decision. The agent-readable methodology explains the schema and scoring. Save an export with its timestamp and SHA-256 checksum when citing a report: this live benchmark changes as questions and models are added.