MiMo-V2.6-Flash won every writing matchup in our new Jev AI benchmark. On the personal-advice question, the same model finished thirteenth. That difference is a good reason to look past the overall leaderboard.

We built Jev Benchmark to see which answers TypeSafe's Jev prefers when different AI models face the same questions. OpenRouter supplies the answers; Jev compares them in pairs; our Django app saves the decisions and calculates rankings. You can read the prompts, full answers, and individual judgments yourself.

This is a preference experiment, not a verdict on the best AI model. Jev is the judge, not a contestant. We have not independently validated the code or mathematical proofs.

How we used Jev to benchmark AI models

The original idea was small: ten models, three questions, and one comparison for every pair of answers. The completed run grew to 29 models and four questions spanning writing, programming, mathematics, and personal advice.

We chose tasks with enough constraints to make a polished but incomplete answer distinguishable from a useful one. Each question has its own rubric, so the judge gets more guidance than "pick whichever sounds better."

  • Writing: advise a four-engineer Django team on whether to split into microservices before an enterprise launch. The answer needs a recommendation, an opposing argument, a six-week plan, and a calculation. The company in the prompt is hypothetical.
  • Programming: implement offline dynamic connectivity in Python, using a segment tree and a rollback disjoint-set union. Repeated edges, removals, memory use, and differential tests all matter.
  • Mathematics: derive the expected stopping time and variance when sums of uniform random variables first exceed two, then generalize the expectation to other thresholds.
  • Personal advice: respond to a friendship strained by caregiving, canceled plans, and late-night messages. The answer must balance compassion with boundaries and acknowledge the person's own contribution to the conflict.
One answer is reused against every opponent
  1. Generate29 models answer each prompt through OpenRouter.
  2. CompareJev chooses A or B for each pair, using the question's rubric.
  3. ReplaySaved decisions produce per-question and overall Elo ratings.

TypeSafe describes Jev as a model built for structured decisions rather than text generation. That fits this job: we need a choice between A and B, not another essay about both answers. Our benchmark pins jev-1.13.0 and stores the choice, probabilities, and confidence returned by the API.

Every model receives the same question, without tools or web access. We omit model labels from the judging request and assign A/B order using a stable hash. The judge sees both answers, the original question, and the rubric. We also instruct it to treat answer text as untrusted material and not reward length alone. Those instructions are precautions, not proof that all bias disappears.

For 29 models, a round robin contains 29 × 28 ÷ 2 = 406 pairs per question. Four questions make 1,624 comparisons. Each model participates in 28 matchups per question, or 112 in total.

The important engineering choice was to save the work. Adding a thirtieth model would require four new saved answers and 116 new pairwise decisions, assuming the existing questions stay unchanged. It would not require regenerating the old answers or rejudging their old pairs. The app's missing-work runner can resume incomplete work while retaining completed results.

What stood out in the results

1. GLM led overall, but different tasks had different winners

GLM 5.3 Prime finished first overall at 1,718.9 Elo in the September 28 snapshot. MiMo V2.6 Pro followed at 1,698.4, with MiMo-V2.6-Flash third at 1,686.7. All four questions had equal weight.

Overall leaders, September 28, 2026
RankModelWeighted Elo
1GLM 5.3 Prime1,718.9
2MiMo V2.6 Pro1,698.4
3MiMo-V2.6-Flash1,686.7
4Claude Fable 5.11,676.8
5Claude Opus 5.51,653.2
6Muse Spark 1.3 Contributor1,630.4

The task leaders tell a more useful story. MiMo-V2.6-Flash led the writing question, MiMo V2.6 Pro led the programming question, and GLM 5.3 Prime led both mathematics and personal advice.

Even a first-place finish needs context. On personal advice, GLM scored 1,724.4, Claude Fable 5.1 scored 1,724.2, and Muse Spark 1.3 Contributor scored 1,723.9. A half-point separates all three. We have no repeated trials or uncertainty intervals that would justify treating that order as a reliable distinction.

2. MiMo Flash's overall score hides a task-specific gap

MiMo-V2.6-Flash was first in writing, second in programming, and second in mathematics. It then fell to thirteenth on personal advice. The chart shows its wins, with its Elo rank alongside each task.

Same model, different tasks: MiMo-V2.6-Flash
Writing28 / 28 wins · rank 1
Programming25 / 28 wins · rank 2
Mathematics26 / 28 wins · rank 2
Personal advice15 / 28 wins · rank 13

Bar scale: 0 to 28 wins. One saved answer per task, judged once against each opponent. Source: the four public question pages, September 28, 2026.

That makes the overall table a starting point, not a model-selection policy. If your application writes technical articles, you would inspect a different slice of these results than if it helps people think through difficult conversations.

It also raises a question worth testing: does this pattern persist across many writing and advice prompts? Four questions cannot answer that. They give us a concrete lead for the next experiment.

3. An unbeaten answer is still not a verified answer

There were two clean sweeps: MiMo-V2.6-Flash went 28-0 on writing, and GLM 5.3 Prime went 28-0 on the mathematics question. Those are strong preferences within this field of answers. They are not independent checks of the recommendations or derivations.

The distinction is especially important for programming. The prompt requests executable code and tests, but this benchmark does not run that code. A judge can prefer an explanation that contains a bug. Before adopting a solution, we would want the tests executed and the edge cases checked.

The cheaper answers were worth a closer look

The cost column adds another useful dimension. OpenRouter reported $0.01475271 for MiMo-V2.6-Flash's four saved answers, compared with $1.13225 for Claude Fable 5.1's four saved answers. Flash placed third overall; Fable placed fourth. For these saved responses, Fable's total was about 77 times higher.

That is enough to make Flash worth testing on a relevant workload. It is not enough to call it 77 times better value in production. These totals depend on the prompts, generated output, reasoning usage, and provider pricing for this run. They exclude failed or retried attempts and Jev's judging charges. They are not the experiment's full bill.

Generation settings were not identical budgets across models either. The app records the requested output allowance on each answer, and some truncated attempts needed a higher allowance before producing a complete answer. Provider-default sampling applied. We therefore describe these as the results of our configured workflow, not a controlled equal-compute comparison.

How to read the Elo rankings without overreading them

Elo turns a sequence of head-to-head wins and losses into a relative rating. Here, each model starts each question at 1,500. Every saved match is replayed once with K = 32, which controls the size of rating updates. The overall score is the weighted arithmetic mean of completed per-question ratings.

There are a few consequences worth keeping next to the leaderboard:

  • Elo is not an accuracy percentage. A score of 1,700 does not mean an answer is 17% better than one rated 1,450. These ratings are relative to this set of opponents.
  • Replay order matters. We use a stable hashed model-ID order so the same stored dataset produces the same result. That is reproducibility, not order independence. New opponents can change existing ratings without changing any old judgment.
  • One judge is one perspective. Jev is asked to choose a winner even when answers are close; there is no tie option. We have not compared its choices with a human panel here.
  • Blinding has limits. Model labels are omitted, but answer style or self-identification can still reveal origin. We do not repeat each pair with A/B positions swapped.
  • Confidence is not correctness. Jev's confidence describes its decision distribution. It does not establish the probability that a proof, program, or piece of advice is right.

And the 1,624 comparisons are not 1,624 independent examples of model ability. They reuse just 116 saved answers across four prompts. More pairings make the tournament complete; they do not make the task sample broad.

The saved judgments also do not contain a written explanation of why Jev picked a winner. We can inspect which answer it selected and how its probabilities were distributed, but we cannot turn that into a claim that it preferred a particular sentence, tone, or algorithm. To investigate those causes, we would need a separate experiment that changes one feature at a time.

For a practical comparison, choose a task that resembles your workload, read the top answers, and check what they missed against the rubric. Then test the promising models on your own examples. A complete tournament is useful evidence; relevance to your actual job still has to be established.

What we would test next

The next useful step is more independent prompts, not more precision after the decimal point. We would also repeat generations, swap answer positions, run the programming tests, check the mathematical solutions, and ask humans to review a sample of disagreements. Those checks are future work, not features we claim to have completed.

For now, the benchmark gives us something concrete to inspect: the same questions answered by many models, with every head-to-head decision available. Start with the personal-advice results if you want to see where the overall ordering stops being a useful shortcut.

Explore Jev Benchmark, read a few answers, and decide where you agree with its judge. If you need an evaluation workflow built around your own product's tasks, talk to LVTD.