Jev Benchmark

Compare model quality, cost, and the answers behind the scores.

Answers from OpenRouter, compared by TypeSafe's Jev. A small, fully inspectable preference experiment. Not a claim of general intelligence or verified accuracy.

Read the launch study: 29 models, 1,624 comparisons, and what stood out.

Models
29
Questions
6
Answers collected
174 / 174
Pairs judged
2436 / 2436

Overall ranking

How scoring works ↓

Weighted mean of per-question Elo. A model receives an overall score only after completing every matchup across all active questions.

Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelWeighted EloMean answer costValue / 100
MiMo-V2.6-FlashXiaomi 1699.3 $0.002860 76.5
Claude Fable 5.1Anthropic 1682.8 $0.226293 53.2
MiMo V2.6 ProXiaomi 1681.2 $0.010550 66.4
Claude Opus 5.5Anthropic 1663.9 $0.076747 53.8
GLM 5.3 PrimeZ.ai 1647.8 $0.165112 50.8

Value v1 blends 70% Elo-derived quality with 30% affordability. Higher is better under this chosen tradeoff. See the formula and limitations. Mean cost uses the same question weights as Elo; unknown costs are never treated as free.

Costs are actual USD reported by OpenRouter for saved successful answers, including provider-reported reasoning usage. Zero means reported free; Unknown means no cost was reported. These totals exclude failed/retried attempts and Jev charges, and are not an account invoice. Attempts absent from retained responses and backups remain unknown.

Latest judgment: Sep 28, 2026, 17:58 UTC.

Quality and cost, together

Higher is stronger in Jev's rankings. Further left is cheaper. Solid points mark the observed frontier: no other model is at least as strong and at least as cheap, with one strict improvement. Select a point for its values.

17501150Weighted Elo $0$0.001$0.01$0.1 Mean answer cost (USD, log1p scale; includes $0) MiMo-V2.6-Flash: 1699.3 Elo; $0.002859698333333333333333333333Claude Fable 5.1: 1682.8 Elo; $0.2262933333333333333333333333MiMo V2.6 Pro: 1681.2 Elo; $0.010550345Claude Opus 5.5: 1663.9 Elo; $0.07674733333333333333333333333GLM 5.3 Prime: 1647.8 Elo; $0.1651120GLM 5.3 Flash: 1608.0 Elo; $0.004392485GPT-6 Astra: 1590.5 Elo; $0.096045Muse Spark 1.3 Contributor: 1587.6 Elo; $0.001163283333333333333333333333Hy4 preview: 1581.5 Elo; $0.02792524333333333333333333333Kimi K3: 1562.5 Elo; $0.08226114853333333333333333333DeepSeek V4.1 Flash: 1561.4 Elo; $0.006077301666666666666666666667Space Bunny Alpha: 1549.6 Elo; $0MiniMax M2.7 (Nitro): 1547.4 Elo; $0.0127039Qwen3.8 Max Prime: 1534.8 Elo; $0.1461226666666666666666666667Nemotron 3 Ultra (free): 1513.8 Elo; $0gpt-oss-120b: 1481.1 Elo; $0.0006473833333333333333333333333Grok 4.7: 1477.7 Elo; $0.03655413333333333333333333333Hy3: 1472.0 Elo; $0.005701395333333333333333333333Gemini 3.8 Flash: 1461.1 Elo; $0.01241700DeepSeek V4 Pro: 1430.1 Elo; $0.005675268373333333333333333333GPT-6 Luna: 1423.6 Elo; $0.00122345GPT-6 Sol: 1401.9 Elo; $0.01687066666666666666666666667Solar Pro 4: 1397.7 Elo; $0.00124800Qwen3.7 Flash: 1375.5 Elo; $0.001224756666666666666666666667gpt-oss-20b (Nitro): 1369.9 Elo; $0.001002208333333333333333333333Ling 3.0 Flash: 1354.2 Elo; $0.000381808Gemini 3.1 Pro Preview: 1312.1 Elo; $0.09083833333333333333333333333Mercury 2.5: 1290.5 Elo; $0.00044982Mistral Medium 3.5: 1240.8 Elo; $0.01581075
Hover, tap, or keyboard-focus a model. All plotted values are also available in the overall table and JSON.

One question at a time

Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
MiMo-V2.6-FlashXiaomi 1780.2 $0.001234 85.1 28 / 0
MiMo V2.6 ProXiaomi 1728.8 $0.003644 77.2 26 / 2
Claude Opus 5.5Anthropic 1721.8 $0.058136 59.1 25 / 3
Muse Spark 1.3 ContributorMeta 1682.0 $0.000733 79.8 24 / 4
Claude Fable 5.1Anthropic 1680.9 $0.267990 52.8 23 / 5
Read the prompt, answers, and judgments →
Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
MiMo V2.6 ProXiaomi 1721.7 $0.029694 62.3 25 / 3
MiMo-V2.6-FlashXiaomi 1710.7 $0.007339 71.3 25 / 3
MiniMax M2.7 (Nitro)MiniMax 1708.4 $0.025938 62.1 24 / 4
GLM 5.3 PrimeZ.ai 1702.8 $0.286540 54.4 24 / 4
GLM 5.3 FlashZ.ai 1663.0 $0.011422 64.3 22 / 6
Read the prompt, answers, and judgments →
Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
GLM 5.3 PrimeZ.ai 1782.9 $0.532690 59.1 28 / 0
MiMo-V2.6-FlashXiaomi 1726.0 $0.006008 73.8 26 / 2
Muse Spark 1.3 ContributorMeta 1721.4 $0.003476 77.0 25 / 3
MiMo V2.6 ProXiaomi 1717.5 $0.020503 64.3 25 / 3
Claude Fable 5.1Anthropic 1694.5 $0.387750 53.5 24 / 4
Read the prompt, answers, and judgments →
Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
GLM 5.3 PrimeZ.ai 1724.4 $0.050310 59.9 26 / 2
Claude Fable 5.1Anthropic 1724.2 $0.081930 58.2 25 / 3
Muse Spark 1.3 ContributorMeta 1723.9 $0.000505 83.4 26 / 2
GPT-6 AstraOpenAI 1710.5 $0.049070 59.0 24 / 4
Kimi K3Moonshot AI 1699.8 $0.050908 58.1 23 / 5
Read the prompt, answers, and judgments →
Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
Grok 4.7xAI 1756.5 $0.012107 70.6 27 / 1
MiMo-V2.6-FlashXiaomi 1731.6 $0.001612 81.2 26 / 2
MiMo V2.6 ProXiaomi 1707.1 $0.005753 72.7 24 / 4
DeepSeek V4.1 FlashDeepSeek 1690.0 $0.002854 75.8 24 / 4
Claude Fable 5.1Anthropic 1686.7 $0.142860 54.1 23 / 5
Read the prompt, answers, and judgments →
Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
Claude Opus 5.5Anthropic 1757.3 $0.042860 62.7 27 / 1
MiMo-V2.6-FlashXiaomi 1717.3 $0.000793 82.2 25 / 3
Claude Fable 5.1Anthropic 1702.6 $0.082650 56.6 24 / 4
MiniMax M2.7 (Nitro)MiniMax 1694.2 $0.005416 72.2 24 / 4
DeepSeek V4.1 FlashDeepSeek 1661.7 $0.001583 76.1 23 / 5
Read the prompt, answers, and judgments →

Methodology

Value score v1

Quality q = 1 / (1 + 10((1500 − Elo) / 400)). Affordability a = 1 / (1 + cost / $0.01). Value = 100 × (0.7q + 0.3a).

The 70/30 blend and one-cent reference are explicit design choices, not learned truths. Quality is an Elo-implied preference against a 1,500-rated reference, not factual accuracy. Free answers receive finite scores; unknown costs and incomplete scopes are not scored. Overall cost is the question-weighted mean per saved answer. Free-tier availability and prices can change.

One answer. Every opponent.

Each model answers each question once, without tools or web access. The same prompt goes to every model; provider-default sampling settings apply. Reasoning effort can be explicitly configured on a new question version; existing inputs are frozen after use. Answers start with each model's configured output budget. Failed, truncated answers may be retried with an explicitly raised allowance, recorded on each answer. Truncated or failed answers are not scored.

Each unordered pair is judged once. Model labels are omitted from the judge's input, and A/B presentation is assigned by a stable hash. Answers may still reveal their origin. Probabilities and confidence are saved alongside each decision.

Reproducible Elo

Every question starts all models at 1,500. Each stored win/loss is replayed once with K = 32 in a stable, hashed model-ID order. That makes the same dataset reproducible, not order-independent. Adding opponents can change existing ratings without rejudging their old matches.

Question weights default to one. The overall score is the weighted arithmetic mean of complete per-question ratings. Pending models have no overall rank; per-question results remain provisional until the round robin is complete.

A preference experiment

Jev selects a winner, not a verified truth. This small sample measures one judge's preferences on these prompts, not general model ability. Code is not executed and mathematical proofs are not independently verified. A single presentation per pair does not eliminate position bias.

Confidence describes Jev's decision distribution; it is not the probability that the answer is factually correct. No repeated trials, error bars, or statistical-significance claims are implied.

Use this as a source

The full JSON dataset includes prompts, rubrics, original answers, generation settings, costs, attempt history, and every pairwise decision. The agent-readable methodology explains the schema and scoring. Save an export with its timestamp and SHA-256 checksum when citing a report: this live benchmark changes as questions and models are added.