Read a rollout experiment without fooling yourself

29 / 29 answers · 406 / 406 pairs · Judge: jev-1.13.0

The prompt

You are advising a support operations lead. Use only the synthetic data below; no tools or outside research. A two-week observational pilot tested an AI drafting assistant. Tickets were not randomized: team leads chose which tickets used AI. A resolved ticket means closed within 24 hours. Reopened means reopened within seven days; these counts are subsets of resolved tickets. Cost is total handling cost for all assigned tickets, not only resolved tickets.

Group | Ticket type | Assigned | Resolved within 24h | Reopened | Handling cost USD
AI | Simple | 900 | 810 | 81 | 2700
AI | Complex | 100 | 40 | 12 | 1200
Control | Simple | 100 | 95 | 4 | 400
Control | Complex | 900 | 450 | 45 | 14400

The vendor says: 'AI increased resolution from 54.5% to 85%, so roll it out to every ticket immediately.' The director asks for a decision brief of no more than 650 words. Include: (1) the aggregate and within-ticket-type resolution rates; (2) rates standardized to a 50/50 simple/complex workload; (3) durable resolutions, defined as resolved minus reopened, and cost per durable resolution for each group, noting the ticket-mix limitation; (4) a justified rollout recommendation; (5) a practical randomized follow-up experiment with an explicit primary metric, guardrails, allocation and stopping rule. Distinguish facts, interpretations, and unknowns. Do not invent confidence intervals, sample-size calculations, causal effects or staff-level data. Show enough arithmetic to audit your conclusions.
Judging rubric

Prioritize numerical correctness, denominator discipline, confounding awareness and an actionable decision. Ground truth: aggregate AI 850/1000=85%, control 545/1000=54.5%; within simple AI90% vs control95%; within complex40% vs50%; equal-mix standardization65% vs72.5%. Durable totals AI757, control496; costs3900 and14800; cost per durable resolution approximately5.15 and29.84USD, not causal or mix-adjusted savings. A strong answer identifies the reversal/selection bias and avoids endorsing universal rollout from aggregate rates. Follow-up should randomize within complexity strata, predefine durable-resolution quality and cost measurement, allow seven-day reopening maturation, use safety/escalation guardrails and an explicit precommitted stopping rule without fabricated statistical power. Reward concise auditable reasoning, not length. Penalize invented facts, mixing denominators, causal certainty, and ignoring negative within-stratum quality.

Question ranking

Top 5 of 29 models
Model rankings. Value is an experimental quality and affordability blend. Costs are actual reported USD.
ModelEloAnswer costValue / 100W / L
Grok 4.7xAI 1756.5 $0.012107 70.6 27 / 1
MiMo-V2.6-FlashXiaomi 1731.6 $0.001612 81.2 26 / 2
MiMo V2.6 ProXiaomi 1707.1 $0.005753 72.7 24 / 4
DeepSeek V4.1 FlashDeepSeek 1690.0 $0.002854 75.8 24 / 4
Claude Fable 5.1Anthropic 1686.7 $0.142860 54.1 23 / 5

Elo starts at 1,500, K = 32. Scores are provisional while pairs remain. Full methodology.

The answers

Original model output, displayed as escaped plain text. No generated code or HTML is executed.

Grok 4.7

Requested model: x-ai/grok-4.7 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0121072000 USD. Excludes retry and judging costs.

Decision brief: do not roll the AI drafter out to every ticket.

**Facts.** Leads chose which tickets used AI; assignment was not randomized. Resolved means closed within 24 hours. Reopened counts are subsets of resolved tickets (within seven days). Cost is total handling cost for all assigned tickets.

| Group | Type | Assigned | Resolved | Resolution | Reopened | Durable (resolved − reopened) | Cost USD |
|---|---|---:|---:|---:|---:|---:|---:|
| AI | Simple | 900 | 810 | 810/900 = 90% | 81 | 729 | 2,700 |
| AI | Complex | 100 | 40 | 40/100 = 40% | 12 | 28 | 1,200 |
| AI | All | 1,000 | 850 | 850/1,000 = 85% | 93 | 757 | 3,900 |
| Control | Simple | 100 | 95 | 95/100 = 95% | 4 | 91 | 400 |
| Control | Complex | 900 | 450 | 450/900 = 50% | 45 | 405 | 14,400 |
| Control | All | 1,000 | 545 | 545/1,000 = 54.5% | 49 | 496 | 14,800 |

The vendor’s 54.5% → 85% figures match the aggregates only. Within type, AI resolution is lower on both simple (90% vs 95%) and complex (40% vs 50%) tickets. AI handled 900/1,000 simple tickets; control handled 900/1,000 complex tickets.

**Standardized to a 50/50 simple/complex mix (equal weight on the two within-type rates, not a reweight of individual tickets):** AI (90% + 40%)/2 = 65%; control (95% + 50%)/2 = 72.5%. Under this mix, control’s resolution rate is 7.5 percentage points higher.

**Durable resolutions and cost.** Aggregate durable rate: AI 757/1,000 = 75.7%; control 496/1,000 = 49.6%. Aggregate cost per durable resolution: AI 3,900/757 ≈ $5.15; control 14,800/496 ≈ $29.84. **Ticket-mix limitation:** those aggregates are not comparable. AI’s book was 90% simple; control’s was 90% complex. Within type, cost per durable resolution is AI simple 2,700/729 ≈ $3.70 vs control simple 400/91 ≈ $4.40, and AI complex 1,200/28 ≈ $42.86 vs control complex 14,400/405 ≈ $35.56. Among resolved tickets, reopen shares were AI simple 81/810 = 10%, AI complex 12/40 = 30%, control simple 4/95 ≈ 4.2%, control complex 45/450 = 10%.

**Interpretation.** The aggregate “AI win” is consistent with Simpson’s paradox from a much easier mix, not with better within-type resolution. On the observed within-type rates, a balanced workload favors control on 24-hour resolution. Higher reopen shares under AI, in both types, mean raw resolution overstates durable closure. Lower aggregate AI cost is largely mix; within complex tickets, observed cost per durable resolution is higher for AI.

**Unknowns.** No causal effect is identified. Leads may have sent easier cases to AI even within “simple” and “complex.” We do not know quality beyond reopen, customer effort, handle time, or whether cost includes license fees. Small AI-complex (n=100) and control-simple (n=100) cells are noisy; no confidence intervals are calculated here.

**Recommendation.** Do not roll out to every ticket. Keep the current process for complex work. Any continued simple-ticket use should stay limited and supervised until a randomized test shows a durable-resolution gain without a reopen or cost breach. The vendor claim is not a sound basis for universal rollout.

**Follow-up experiment (practical, not a sample-size calculation).** Randomize tickets 1:1 to AI draft vs usual process, stratified by simple vs complex, with leads unable to override assignment. Run a fixed four-week window (or the next 400 assigned tickets per arm within each type, whichever comes later). **Primary metric:** durable-resolution rate (closed within 24 hours and not reopened within seven days), reported overall and by type. **Guardrails (stop that stratum’s AI arm if breached):** reopen share among resolved tickets above the control arm’s rate in that type by 5 percentage points or more, checked weekly; or handling cost per assigned ticket above the control arm in that type. **Stopping rule:** no early stop for benefit; stop a stratum only for a guardrail breach or for an operational safety issue; otherwise analyze only at the pre-set end. Compare arms within type first; do not pool across types without showing the mix.
MiMo-V2.6-Flash

Requested model: xiaomi/mimo-v2.6-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0016121000 USD. Excludes retry and judging costs.

# Decision Brief — AI Drafting Assistant (2-week observational pilot)

## 1. Facts (arithmetic shown, auditable)

**Aggregate resolution:** AI 850/1,000 = **85.0%**; control 545/1,000 = **54.5%**. The vendor's two figures are arithmetically correct.

**Within ticket type:** AI simple 810/900 = **90.0%** vs control simple 95/100 = **95.0%**; AI complex 40/100 = **40.0%** vs control complex 450/900 = **50.0%**.

**Ticket mix:** AI handled 900/1,000 = 90% simple; control handled 900/1,000 = 90% complex. Team leads chose assignments (non-random).

**Standardized to a 50/50 simple/complex workload:**
- AI (90.0 + 40.0)/2 = **65.0%**
- Control (95.0 + 50.0)/2 = **72.5%** → AI trails by 7.5 pp.

**Durable resolutions (resolved − reopened):**
- AI: simple 810 − 81 = 729; complex 40 − 12 = 28; total **757/1,000 = 75.7%**. By type: 729/900 = 81.0%, 28/100 = 28.0%.
- Control: simple 95 − 4 = 91; complex 450 − 45 = 405; total **496/1,000 = 49.6%**. By type: 91/100 = 91.0%, 405/900 = 45.0%.
- Standardized durable: AI (81.0 + 28.0)/2 = **54.5%**; control (91.0 + 45.0)/2 = **68.0%** → AI trails by 13.5 pp.

**Cost per durable resolution** (cost covers *all* assigned tickets; numerator is durable):
- AI total (2,700 + 1,200)/757 = **$5.15**; control total (400 + 14,400)/496 = **$29.84**.
- Simple only: AI 2,700/729 = **$3.70** vs control 400/91 = **$4.40**.
- Complex only: AI 1,200/28 = **$42.86** vs control 14,400/405 = **$35.56**.

**Ticket-mix limitation:** aggregate cost comparisons are confounded by mix — AI carried few expensive complex tickets, control carried many. Only the within-type ratios are mix-free.

## 2. Interpretations (not established facts)

The aggregate gain is a ticket-mix artifact (Simpson's paradox): AI leads overall yet trails control *inside both strata*. Because assignment was not randomized, even within-type gaps are not causal. AI's only apparent advantage is lower cost per durable resolution on simple tickets, which coexists with a 10 pp durable-rate deficit (81.0% vs 91.0%).

## 3. Unknowns

Why leads selected tickets; how "simple/complex" was assigned and how reliable that label is; residual difficulty differences within type; actual AI usage per ticket; reasons for reopens; comparability of cost accounting across groups; outcomes beyond seven days.

## 4. Recommendation

**Do not roll out to every ticket.** The vendor's causal claim is unsupported, and standardized rates put AI behind on resolution and durable resolution in this pilot. Replace lead-selected assignment with randomized assignment (below); hold AI to simple tickets only, under monitoring, if a gap period is unavoidable. Revisit after trial data.

## 5. Randomized follow-up experiment

**Allocation:** 1:1 randomization within each ticket type at intake (simple/complex only known stratum). Enroll every ticket; log and report any lead overrides as protocol deviations. Run until a pre-set two-week volume, with each ticket's 7-day reopen window closed before analysis.

**Primary metric:** durable resolution rate per assigned ticket (resolved ≤24h *and* not reopened ≤7d), reported by type and standardized to 50/50.

**Guardrails:** 24h resolution by type; 7-day reopen rate by type; cost per assigned ticket and per durable resolution; label-misclassification rate.

**Stopping rule (operational, no significance claim):** review after each two-week block. Stop AI in a stratum if it trails control on the primary metric at two consecutive reviews, or its reopen rate exceeds control's, or its cost per durable resolution exceeds control's. Report all results regardless.
MiMo V2.6 Pro

Requested model: xiaomi/mimo-v2.6-pro · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:31 UTC

Reported answer cost: $0.0057528750 USD. Excludes retry and judging costs.

**Decision Brief — AI Drafting Assistant Pilot (2 weeks, observational)**
*Tickets were not randomized; team leads selected which tickets used AI. All figures are descriptive, not causal.*

**1. Facts — resolution rates**
- Aggregate: AI 850/1,000 = **85.0%**; control 545/1,000 = **54.5%**. The vendor's arithmetic is correct.
- Workload differs: AI is 900/1,000 = 90% simple; control is 900/1,000 = 90% complex.
- Within type (both favor control): Simple — AI 810/900 = 90.0% vs control 95/100 = 95.0% (−5.0 pts). Complex — AI 40/100 = 40.0% vs control 450/900 = 50.0% (−10.0 pts).
- Standardized to a 50/50 simple/complex workload (average of within-type rates): AI (90.0 + 40.0)/2 = **65.0%**; control (95.0 + 50.0)/2 = **72.5%**.
- The aggregate gap (+30.5 pts) reverses to −7.5 pts under equal workloads — a ticket-mix effect (Simpson's paradox).

**2. Facts — durable resolutions and cost**
- Reopened: AI 81 + 12 = 93; control 4 + 45 = 49. Reopen share of resolved: AI 93/850 = 10.9%; control 49/545 = 9.0%.
- Durable (resolved − reopened): AI 850 − 93 = **757**; control 545 − 49 = **496**.
- Cost: AI $2,700 + $1,200 = $3,900; control $400 + $14,400 = $14,800. Cost per durable: AI $3,900/757 = **$5.15**; control $14,800/496 = **$29.84**.
- **Mix limitation:** cost covers all assigned tickets, and AI handled 90% simple at $2,700/900 = $3.00 each vs control's 90% complex at $14,400/900 = $16.00 each (AI complex $12.00; control simple $4.00). Within type, cost per durable (durable counts: AI simple 810−81=729, AI complex 40−12=28, control simple 95−4=91, control complex 450−45=405): AI simple $2,700/729 = $3.70 vs control $400/91 = $4.40; **AI complex $1,200/28 = $42.86 vs control $14,400/405 = $35.56**. Standardized 50/50 (500 tickets per type at observed unit costs and durable rates): AI (1,500+6,000)/(405+140) = $7,500/545 = **$13.76**; control (2,000+8,000)/(455+225) = $10,000/680 = **$14.71**. The headline saving is mostly mix; the residual edge is on simple tickets only.

**3. Interpretations**
- The rollout claim is not supported: within both ticket types AI resolved fewer tickets and produced fewer durable resolutions (simple 729/900 = 81.0% vs 91/100 = 91.0%; complex 28/100 = 28.0% vs 405/900 = 45.0%).
- AI's lower handling cost per ticket within type is consistent with faster drafting — or with thinner handling that later reopens (AI's reopen share is higher).
- Nothing supports a causal effect: leads chose the tickets and their selection rule is unknown.

**4. Unknowns**
Lead selection criteria; agent experience and capacity; whether AI drafts were edited before sending; difficulty variation within "simple/complex"; cost composition (rates vs time); outcomes beyond seven days; CSAT. Two cells are small (AI complex 100; control simple 100); no intervals were computed.

**5. Recommendation**
**Do not roll out to every ticket.** Keep AI drafting confined to simple tickets with mandatory human review, and retire the 85% vs 54.5% comparison as evidence. Rationale: the only consistent advantage is handling cost; resolution and durability are at best neutral and possibly worse; complex tickets are the worst case (40.0% resolution, $42.86 per durable). Scale only after randomized evidence shows durable resolution is no worse with a genuine cost saving.

**6. Randomized follow-up**
- **Allocation:** randomize at intake (system, not leads) to AI-draft vs. no-AI, 1:1, stratified by ticket type and blocked by week, with a pre-set 50/50 simple/complex target so both strata fill.
- **Primary metric:** durable resolution rate = (resolved within 24h and not reopened within 7 days) ÷ all assigned, measured on tickets with full 7-day follow-up. One primary metric.
- **Guardrails (by stratum):** reopen rate within 7 days of resolved; handling cost per assigned ticket and per durable resolution; complex-ticket resolution rate; first-response time and queue age; CSAT/complaints if collected.
- **Stopping rule:** pre-register duration and sample size (powered in advance); one pre-planned interim review that may stop for harm only if a guardrail crosses a pre-agreed margin; no stopping for early success; otherwise run to the pre-set end. Predefine exclusions (duplicates, merges, escalations) and analyze all tickets by assigned group.
DeepSeek V4.1 Flash

Requested model: deepseek/deepseek-v4.1-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0028537500 USD. Excludes retry and judging costs.

**Decision brief: AI drafting assistant**

**Facts (synthetic pilot arithmetic)**
Aggregate resolution, closed ≤24h / assigned:  
AI = (810+40)/(900+100)=850/1,000=**85.0%**.  
Control = (95+450)/(100+900)=545/1,000=**54.5%**.

Within ticket type:  
AI Simple = 810/900=**90.0%**; Control Simple = 95/100=**95.0%**.  
AI Complex = 40/100=**40.0%**; Control Complex = 450/900=**50.0%**.

Standardized to 50/50 simple/complex:  
AI = 0.5×90.0% + 0.5×40.0% = **65.0%**.  
Control = 0.5×95.0% + 0.5×50.0% = **72.5%**.  
So under equal mix, control is higher by **7.5 percentage points**.

Durable resolutions = resolved minus reopened.  
AI total reopened = 81+12=93; durable = 850−93=**757**. Cost = 2,700+1,200=3,900; cost/durable = 3,900/757=**$5.15**.  
Control total reopened = 4+45=49; durable = 545−49=**496**. Cost = 400+14,400=14,800; cost/durable = 14,800/496=**$29.84**.

Within-type durable and cost/durable:  
AI Simple: 810−81=729 durable; 2,700/729=**$3.70**.  
AI Complex: 40−12=28 durable; 1,200/28=**$42.86**.  
Control Simple: 95−4=91 durable; 400/91=**$4.40**.  
Control Complex: 450−45=405 durable; 14,400/405=**$35.56**.

Standardized 50/50 durable rate:  
AI = 0.5×(729/900=81.0%) + 0.5×(28/100=28.0%) = **54.5%**.  
Control = 0.5×(91/100=91.0%) + 0.5×(405/900=45.0%) = **68.0%**.

**Interpretation**
The vendor’s 85% vs 54.5% is real but mix-confounded. AI tickets were 90% simple; control tickets were 90% complex. Simple tickets resolve more often, so aggregate AI looks better. Within each ticket type, AI resolution and durable rates are lower in this observational pilot. Aggregate cost/durable also favors AI heavily because AI handled mostly simple tickets. Within type, AI simple has lower cost/durable, but AI complex has higher cost/durable than control complex.

These are associations, not causal effects. Team leads chose tickets, so AI and control groups may differ in unobserved complexity, urgency, or selection criteria. The pilot cannot prove AI caused worse or better outcomes.

**Unknowns**
Why leads selected tickets; whether ticket difficulty differed beyond simple/complex labels; consistency of reopening rules; long-run effects; workload/capacity effects; whether results would hold under random assignment. No confidence intervals, staff-level data, or causal effects are available here.

**Recommendation**
Do **not** roll out immediately to every ticket. The headline increase is explained by ticket mix; a 50/50 standardized comparison favors control on resolution and durable resolution. Within-type signals do not support universal expansion. Continue only as a controlled randomized follow-up. If interim use is necessary, restrict to simple tickets with monitoring, and avoid claiming proven benefit.

**Randomized follow-up**
Design: ticket-level randomized controlled experiment, stratified by ticket type (simple/complex). Randomize eligible tickets **50/50 within each stratum** to AI-assisted drafting versus standard workflow.

Primary metric: **durable resolution rate per assigned ticket** — closed ≤24h and not reopened within 7 days.

Guardrails: reopened rate among resolved tickets; complex-ticket durable resolution; cost per assigned ticket and cost per durable resolution.

Stopping rule: weekly interim review. Stop the AI arm for harm if, at two consecutive reviews, complex-ticket durable resolution is ≥5 percentage points lower than control, or reopened-among-resolved is ≥5 percentage points higher. Stop for futility if after 6 weeks AI does not exceed control on the primary metric by ≥2 percentage points. Do not stop early for benefit without review. Otherwise complete the 6-week enrollment plus 7-day follow-up.
Claude Fable 5.1

Requested model: anthropic/claude-fable-5.1 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.1428600000 USD. Excludes retry and judging costs.

**Decision brief: AI drafting assistant pilot**

**1. Resolution rates (facts, from the pilot table)**

Aggregate: AI 850/1000 = 85.0%; Control 545/1000 = 54.5%. The vendor's numbers are arithmetically correct.

Within ticket type:
- Simple: AI 810/900 = 90.0%; Control 95/100 = 95.0%
- Complex: AI 40/100 = 40.0%; Control 450/900 = 50.0%

Interpretation: the aggregate advantage reverses within both types (Simpson's paradox). Team leads sent 90% of AI tickets to Simple and 90% of Control tickets to Complex. The 85% vs 54.5% gap is a ticket-mix artifact, not evidence the AI resolves tickets better.

**2. Standardized to a 50/50 simple/complex workload**

- AI: 0.5 × 90.0% + 0.5 × 40.0% = 65.0%
- Control: 0.5 × 95.0% + 0.5 × 50.0% = 72.5%

On an equal mix, Control resolves more within 24h. (Interpretation: this is descriptive standardization, not a causal estimate; unknown whether the actual workload is 50/50.)

**3. Durable resolutions and cost**

Durable = resolved − reopened:
- AI Simple: 810 − 81 = 729 (81.0% of assigned; reopen rate 81/810 = 10.0%)
- AI Complex: 40 − 12 = 28 (28.0% of assigned; reopen rate 12/40 = 30.0%)
- Control Simple: 95 − 4 = 91 (91.0%; reopen rate 4/95 = 4.2%)
- Control Complex: 450 − 45 = 405 (45.0%; reopen rate 45/450 = 10.0%)

Totals: AI 757; Control 496. Standardized 50/50 durable rate: AI (81.0 + 28.0)/2 = 54.5%; Control (91.0 + 45.0)/2 = 68.0%.

Cost per durable resolution (cost covers all assigned tickets, so it is only comparable within type):
- Aggregate: AI $3,900/757 = $5.15; Control $14,800/496 = $29.84 — misleading, driven by mix.
- Simple: AI $2,700/729 = $3.70; Control $400/91 = $4.40 (AI ~16% cheaper).
- Complex: AI $1,200/28 = $42.86; Control $14,400/405 = $35.56 (AI ~21% more expensive).

Cost per assigned ticket: AI Simple $3.00 vs Control $4.00; AI Complex $12.00 vs Control $16.00. Interpretation: AI lowers handling cost per ticket in both types, but the lower resolution and higher reopen rates erase the saving on Complex tickets. Limitation: Control Simple (n=100) and AI Complex (n=100) are small; leads may also have selected easier-than-average tickets for AI within each type, so even within-type comparisons are not causal.

**Unknowns:** how leads chose tickets; agent skill differences between arms; customer satisfaction; whether reopens reflect AI drafts or rushed handling; workload mix in production; time-of-day or queue effects.

**4. Recommendation: do not roll out to every ticket.**

Justification: (a) the headline gain disappears within type; (b) AI arms show lower 24h resolution and 2.5–3× higher reopen rates in both types; (c) on Complex tickets AI costs more per durable resolution. The only plausible upside is Simple tickets, where cost per durable resolution is lower despite more reopens — but this rests on n=100 controls and non-random assignment. Recommend: pause expansion, keep AI on Simple tickets only at current scale if operationally necessary, and run the randomized test below before any wider decision.

**5. Randomized follow-up experiment**

- Unit and allocation: randomize at the ticket level, stratified by ticket type (Simple/Complex), 50/50 AI vs. no-AI within each stratum. Assign by system rule at ticket creation; leads cannot override. Log agent ID to check balance (analysis only, not staff evaluation).
- Primary metric: durable resolution rate — resolved within 24h and not reopened within 7 days — reported separately by type and pooled at the actual production mix.
- Guardrails (pre-specified, checked weekly): reopen rate among resolved, handling cost per assigned ticket, CSAT where available, escalation rate, agent-reported quality issues.
- Duration/sample: run a minimum of 4 weeks; the analytics team sets the required ticket count per stratum before launch (I have not calculated it).
- Stopping rule: no early stopping for benefit. Stop the AI arm in a stratum early if its reopen rate exceeds Control by more than 5 percentage points at any weekly guardrail check, or on a serious quality incident. Otherwise analyze once at the pre-set ticket count.
- Decision rule: expand only in strata where AI's durable resolution rate is not lower than Control's and cost per durable resolution is lower.
Muse Spark 1.3 Contributor

Requested model: meta/muse-spark-1.3-contributor · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0005279000 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Pilot — Do Not Roll Out Broadly; Test by Randomization**

**1. What happened — Facts**

Two-week observational pilot, no randomization. Leads chose AI vs Control.

Aggregate 24h resolution [Resolved / Assigned]:
* AI: (810+40) / (900+100) = 850/1000 = **85.0%**
* Control: (95+450) / (100+900) = 545/1000 = **54.5%**
Vendor arithmetic is correct.

Ticket mix was severely unbalanced — Fact:
* AI: 900 Simple / 100 Complex = 90% simple
* Control: 100 Simple / 900 Complex = 10% simple

Within-ticket-type 24h resolution:
* Simple: AI 810/900=**90.0%** vs Control 95/100=**95.0%**; Control +5.0pp
* Complex: AI 40/100=**40.0%** vs Control 450/900=**50.0%**; Control +10.0pp

**Interpretation:** This is Simpson's paradox. AI is ahead in aggregate only because it got mostly easier Simple tickets. Within *both* types, Control resolved a higher share.

**2. Standardized comparison — 50/50 workload**

To remove mix distortion, weight Simple and Complex equally: 0.5×Simple rate + 0.5×Complex rate.

* AI standardized: 0.5×0.90 + 0.5×0.40 = 0.45+0.20 = **65.0%**
* Control standardized: 0.5×0.95 + 0.5×0.50 = 0.475+0.25 = **72.5%**

Interpretation: On a balanced workload, Control is **+7.5pp** ahead. This does *not* prove AI harms resolution; it shows the 85% vs 54.5% gap cannot be read as an AI effect.

**3. Durable resolution and cost**

Durable = Resolved minus Reopened, per assigned ticket.

Facts:
* AI durable: Simple 810-81=729; Complex 40-12=28; Total 729+28=**757**; 757/1000=**75.7%**
* Control durable: Simple 95-4=91; Complex 450-45=405; Total 91+405=**496**; 496/1000=**49.6%**
* Within-type durable rate: Simple AI 729/900=**81.0%** vs Control 91/100=**91.0%**; Complex AI 28/100=**28.0%** vs Control 405/900=**45.0%**. Control ahead in both.
* Reopened per resolved: AI Simple 81/810=10.0%, Complex 12/40=30.0%; Control Simple 4/95=4.2%, Complex 45/450=10.0%.

Cost per durable [Total cost / Total durable]:
* AI: (2700+1200)/757 = 3900/757 ≈ **$5.15**
* Control: (400+14400)/496 = 14800/496 ≈ **$29.84**

Ticket-mix limitation — Interpretation: Do not use $5.15 vs $29.84 to claim AI is cheaper. Costs follow mix. Per-type for audit:
AI Simple 2700/729≈$3.70; Control Simple 400/91≈$4.40; AI Complex 1200/28≈$42.86; Control Complex 14400/405≈$35.56. Aggregate cost comparison inherits the same selection bias as resolution.

**Unknowns:** How leads assigned tickets; true future simple/complex share; whether costs include AI license/oversight; comparability of reopen risk; any spillovers.

**4. Recommendation: Pause full rollout**

Facts show association only. Interpretation: Non-random choice confounds AI with ticket easiness and possibly agent skill, urgency, or customer. AI trails Control within every observed stratum for both initial and durable resolution. Broad rollout would extrapolate from 90%-simple AI experience to all tickets, including complex where AI data are only n=100.

Action: Keep AI available only in a controlled test; do not mandate for every ticket.

**5. Randomized follow-up**

Randomize at ticket arrival, stratified by Simple/Complex and shift, e.g., 1:1 AI vs Control within each stratum, using a pre-set assignment log leads cannot override.

Primary metric, pre-registered: durable resolution rate = (resolved within 24h and not reopened within 7 days) / randomized tickets. Analyze overall stratified-adjusted plus separately by type.

Guardrails, monitored blinded for safety only: reopened/resolved rate, average handling cost per ticket, backlog/ageing, complaints. Pre-define an operational threshold for pausing and executive review, e.g., clear harm signal on reopened complex tickets.

Allocation and stopping rule: Fixed test window with fixed allocation ratio; analyze once at end of window for a rollout decision. No interim rollout decisions and no extension based on early trends. Record all randomized tickets intention-to-treat, including reassignments and non-use.
Nemotron 3 Ultra (free)

Requested model: nvidia/nemotron-3-ultra-550b-a55b:free · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:28 UTC

Reported answer cost: $0.0000000000 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot**  

**Facts (observed data, two weeks)**  

| Group | Type | Assigned | Resolved ≤24h | Reopened ≤7d | Handling cost ($) |
|-------|------|----------|---------------|--------------|-------------------|
| AI    | Simple | 900    | 810 (90%)     | 81           | 2,700             |
| AI    | Complex| 100    | 40 (40%)      | 12           | 1,200             |
| Ctrl  | Simple | 100    | 95 (95%)      | 4            | 400               |
| Ctrl  | Complex| 900    | 450 (50%)     | 45           | 14,400            |

*Aggregate resolution rates*  
- AI: (810+40)/(900+100) = 850/1,000 = **85%**  
- Control: (95+450)/(100+900) = 545/1,000 = **54.5%**  

*Within-type resolution rates*  
- Simple: AI 90%, Control 95%  
- Complex: AI 40%, Control 50%  

*Standardized to 50/50 simple–complex mix*  
- AI: 0.5×90% + 0.5×40% = **65%**  
- Control: 0.5×95% + 0.5×50% = **72.5%**  

*Durable resolutions (resolved − reopened) and cost per durable resolution*  
- AI Simple: 810−81 = 729; cost/durable = 2,700/729 ≈ **$3.70**  
- AI Complex: 40−12 = 28; cost/durable = 1,200/28 ≈ **$42.86**  
- AI overall: 757 durable; total cost $3,900 → **$5.15/durable**  
- Control Simple: 95−4 = 91; cost/durable = 400/91 ≈ **$4.40**  
- Control Complex: 450−45 = 405; cost/durable = 14,400/405 ≈ **$35.56**  
- Control overall: 496 durable; total cost $14,800 → **$29.84/durable**  

**Interpretations**  
1. The vendor’s claim (“AI increased resolution from 54.5% to 85%”) reflects the *aggregate* rates, which are heavily confounded by ticket mix: AI received 90% simple tickets, Control 90% complex.  
2. When standardized to a common 50/50 workload, Control’s resolution rate (72.5%) exceeds AI’s (65%).  
3. Cost per durable resolution is lower for AI overall ($5.15 vs $29.84), but this is driven by the mix: AI handled mostly cheap simple tickets. Within each type, AI is cheaper for simple tickets ($3.70 vs $4.40) but more expensive for complex tickets ($42.86 vs $35.56). The complex-ticket AI sample is small (100 assigned), so the complex estimate is unstable.  
4. Because team leads chose which tickets used AI, selection bias is likely. We cannot attribute differences to AI causally.  

**Unknowns**  
- True causal effect of AI on resolution, reopen, and cost for each ticket type.  
- Whether AI performance on complex tickets improves with more exposure.  
- Impact on agent learning, customer satisfaction, and long-term ticket quality.  
- Optimal ticket-type targeting for AI assistance.  

**Rollout Recommendation**  
**Do not roll out to all tickets immediately.** The observational data are insufficient to support a universal deployment. The standardized analysis suggests Control outperforms AI on resolution when mix is equalized, and AI’s cost advantage on complex tickets is negative (though imprecisely estimated). Instead, run a randomized experiment to isolate causal effects and identify the ticket types where AI adds value.  

**Proposed Randomized Follow‑up Experiment**  
- **Primary metric:** Durable resolution rate (resolved within 24h and not reopened within 7 days).  
- **Guardrails:**  
  - Reopen rate must not exceed Control by >2 percentage points.  
  - Average handling cost per assigned ticket must not increase >10% vs Control.  
  - Customer satisfaction (CSAT) must not decline (if measured).  
- **Allocation:** Stratified randomization by ticket type (simple/complex) at intake. 50% of each type assigned to AI assist, 50% to Control (no AI). Target ≥2,000 tickets per arm over 4 weeks.  
- **Stopping rule:** Pre‑planned interim analysis at 2 weeks (or 2,500 total tickets). Stop the AI arm if:  
  - Durable resolution rate is >5 percentage points lower than Control (two‑sided p<0.01), **or**  
  - Cost per durable resolution is >20% higher than Control.  
  Otherwise, complete the 4‑week period and evaluate full results before any phased rollout.  

This design balances speed with rigor, directly tests the vendor’s claims on a level playing field, and protects against harm while we learn where AI truly helps.
GPT-6 Astra

Requested model: openai/gpt-6-astra · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0611700000 USD. Excludes retry and judging costs.

## Decision brief: do not roll out universally

### Facts: the aggregate result reverses within ticket types

Resolution means closure within 24 hours.

| Workload | AI resolution rate | Control resolution rate |
|---|---:|---:|
| All assigned tickets | (810 + 40)/1,000 = **85%** | (95 + 450)/1,000 = **54.5%** |
| Simple | 810/900 = **90%** | 95/100 = **95%** |
| Complex | 40/100 = **40%** | 450/900 = **50%** |
| Standardized: 50% simple, 50% complex | 0.5 × 90% + 0.5 × 40% = **65%** | 0.5 × 95% + 0.5 × 50% = **72.5%** |

AI’s observed aggregate advantage is **30.5 percentage points**, but its resolution rate is lower within both ticket types: **5 points lower for simple** and **10 points lower for complex**. Under an equal workload mix, AI is **7.5 points lower**.

### Facts: durability and cost

Durable resolutions equal resolved tickets minus tickets reopened within seven days.

| Measure | AI | Control |
|---|---:|---:|
| Simple durable resolutions | 810 − 81 = **729** | 95 − 4 = **91** |
| Complex durable resolutions | 40 − 12 = **28** | 450 − 45 = **405** |
| Total durable resolutions | **757** | **496** |
| Durable resolutions / assigned | **75.7%** | **49.6%** |
| Total handling cost | $2,700 + $1,200 = **$3,900** | $400 + $14,400 = **$14,800** |
| Cost per durable resolution | $3,900/757 = **$5.15** | $14,800/496 = **$29.84** |

These cost ratios include handling costs for **all assigned tickets**, including unsuccessful ones.

The raw durability and cost comparisons have a major ticket-mix limitation: AI received **90% simple tickets**, while control received **90% complex tickets**. Within types, durable-resolution rates are AI **81% simple and 28% complex**, versus control **91% and 45%**. Standardizing those rates to 50/50 gives **54.5% AI versus 68% control**. The reported cost ratios are not mix-adjusted.

### Interpretation, unknowns, and recommendation

The vendor’s aggregate arithmetic is correct, but “AI increased resolution” is not established. This is a mix-driven reversal: AI performed worse within both recorded categories despite looking better overall.

Because leads selected AI tickets, neither raw nor standardized differences establish causal effects. Standardization addresses only the recorded simple/complex mix—not other selection differences. Statistical uncertainty, longer-term outcomes, and costs beyond reported handling costs remain unknown.

**Recommendation:** reject immediate universal rollout. Restrict further deployment to a monitored randomized evaluation. The observed quality results warrant caution; the lower observed handling cost warrants further testing, not a causal savings claim.

### Proposed randomized follow-up

- **Allocation:** Randomize eligible tickets 1:1 to AI-assisted or usual handling separately within simple and complex strata. Classify complexity before allocation; prevent discretionary overrides and log deviations.
- **Primary metric:** Intention-to-treat durable resolutions per assigned ticket, standardized to 50/50 simple/complex. Count closure within 24 hours with no reopening during the following seven days.
- **Guardrails:** Track within-type resolution and reopening rates, handling cost per assigned ticket and per durable resolution, and independently reviewed serious errors or privacy incidents.
- **Stopping rule:** Enroll for a precommitted four weeks, then allow eight days for outcomes to mature. No early stopping for apparent success; pause AI for any confirmed severe AI-related harm or privacy breach.
- **Decision:** Predefine acceptable guardrail margins and uncertainty criteria before launch. Expand only with credible primary-metric improvement and acceptable guardrails; inconclusive results do not justify universal rollout.
Claude Opus 5.5

Requested model: anthropic/claude-opus-5.5 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0659840000 USD. Excludes retry and judging costs.

# Decision Brief: AI Drafting Assistant Pilot

**Bottom line:** Do not roll out to every ticket. The vendor's 85% vs 54.5% comparison reflects ticket mix. Within each ticket type, AI-assisted tickets did worse. Run a randomized test before any expansion.

## 1. Resolution rates (facts)

| | AI | Control |
|---|---|---|
| Aggregate | 850/1000 = **85.0%** | 545/1000 = **54.5%** |
| Simple | 810/900 = **90.0%** | 95/100 = **95.0%** |
| Complex | 40/100 = **40.0%** | 450/900 = **50.0%** |

AI received 90% simple tickets, while Control received 90% complex tickets. That is why the aggregate favors AI (Simpson's paradox). Within both ticket types, AI's rate is lower.

## 2. Standardized to a 50/50 workload (facts)

- AI: (90.0 + 40.0) / 2 = **65.0%**
- Control: (95.0 + 50.0) / 2 = **72.5%**

## 3. Durable resolutions and cost

Durable resolutions are resolved minus reopened.

| | Durable | Durable rate | Reopen share of resolved | Cost / assigned | Cost / durable |
|---|---|---|---|---|---|
| AI Simple | 810−81 = 729 | 81.0% | 10.0% | $3.00 | 2700/729 = $3.70 |
| Ctrl Simple | 95−4 = 91 | 91.0% | 4.2% | $4.00 | 400/91 = $4.40 |
| AI Complex | 40−12 = 28 | 28.0% | 30.0% | $12.00 | 1200/28 = $42.86 |
| Ctrl Complex | 450−45 = 405 | 45.0% | 10.0% | $16.00 | 14400/405 = $35.56 |

**Group totals (mix-confounded):**
- AI: 757 durable; $3,900 / 757 = **$5.15** per durable resolution.
- Control: 496 durable; $14,800 / 496 = **$29.84** per durable resolution.

**Limitation:** These totals mainly reflect that AI handled cheap, simple tickets. They are not a like-for-like comparison.

**Standardized to 500 simple + 500 complex tickets (arithmetic on observed rates):**

| | Durable resolutions | Cost | Cost / durable |
|---|---|---|---|
| AI | 405 + 140 = 545 (54.5%) | $1,500 + $6,000 = $7,500 | **$13.76** |
| Control | 455 + 225 = 680 (68.0%) | $2,000 + $8,000 = $10,000 | **$14.71** |

## Interpretation

- **Lower handling cost, lower quality.** AI tickets cost about 25% less to handle. They also produced fewer durable resolutions: 13.5 points fewer at a 50/50 mix, and reopens were higher in both ticket types.
- **Small cost edge, big quality gap.** Under a 50/50 mix, AI's cost per durable resolution is only about 6% lower. Control delivers 135 more durable resolutions per 1,000 tickets.
- **Complex tickets look worst for AI.** AI costs more per durable resolution than Control on complex tickets. Its reopen share is also three times higher.
- **Selection may cut either way.** Team leads chose which tickets used AI. They may have sent harder or easier tickets within each type to AI. The observed gaps may therefore overstate or understate AI's true effect.

## Unknowns

- The causal effect of AI in either direction.
- How team leads selected tickets within each type.
- Customer satisfaction, and the cost of handling reopens.
- Agent learning curves.
- Whether "simple" and "complex" were classified consistently.
- Precision, given that two cells have only 100 tickets. No intervals are claimed here.

## 4. Recommendation

1. **Reject the full rollout.** The headline improvement disappears once ticket mix is accounted for.
2. **Pause discretionary AI use on complex tickets.** It showed the weakest results on every measure.
3. **Allow simple-ticket AI use only inside the randomized experiment below.** The cost saving is real in this data, but so is the quality drop (81% vs 91% durable).
4. **Base the next decision on randomized evidence.**

## 5. Randomized follow-up experiment

**Unit and allocation**
- Classify each incoming ticket as simple or complex using a written rule, *before* assignment.
- Within each type, randomly assign tickets 50/50 to AI or Control using a system-generated flag.
- Team leads cannot override the assignment. Any overrides are logged and analyzed as assigned (intention-to-treat).

**Primary metric**
- Durable resolution rate per *assigned* ticket: resolved within 24 hours and not reopened within 7 days.
- Report it by ticket type and standardized to the actual production mix.

**Secondary metric**
- Cost per durable resolution, by ticket type.

**Guardrails, monitored weekly by type**
- Reopen rate
- 24-hour SLA misses
- Escalations
- CSAT, if collected
- Handling cost per assigned ticket

**Stopping rule**
- **Fixed duration:** run for a pre-registered four weeks of intake. The final readout comes seven days after the last ticket, so the reopen window closes for every ticket.
- **Safety stop:** stop AI in a ticket type early if its weekly reopen rate exceeds Control's by more than 5 percentage points for two consecutive weeks.
- **No early win:** do not stop early for success. Pre-register the analysis.

**Decision criteria, fixed in advance**
- Expand AI in a ticket type only if its durable resolution rate is not worse than Control's and no guardrail is breached.
- If AI is worse on durable resolutions, weigh the cost savings explicitly against that loss.
Qwen3.8 Max Prime

Requested model: qwen/qwen3.8-max-prime · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0309760000 USD. Excludes retry and judging costs.

# Decision Brief: AI Drafting Assistant Pilot

## 1. Resolution Rates

**Aggregate:** AI 850/1000 = **85%**; Control 545/1000 = **54.5%**.

**Within ticket type (the critical view):**

| Type | AI | Control |
|------|-----|---------|
| Simple | 810/900 = **90%** | 95/100 = **95%** |
| Complex | 40/100 = **40%** | 450/900 = **50%** |

**Interpretation:** Within every ticket type, control outperforms AI. The aggregate gap is entirely an artifact of ticket mix: AI received 90% simple tickets; control received 90% complex tickets. This is selection bias, not a treatment effect.

## 2. Standardized Rates (50/50 Simple/Complex)

- AI: 0.5 × 90% + 0.5 × 40% = **65%**
- Control: 0.5 × 95% + 0.5 × 50% = **72.5%**

**Fact:** On a like-for-like basis, control resolves 7.5 pp more tickets than AI.

## 3. Durable Resolutions and Cost

Durable = Resolved − Reopened.

| | AI | Control |
|--|-----|---------|
| Durable (aggregate) | 757 | 496 |
| Total cost | $3,900 | $14,800 |
| Cost/durable (aggregate) | **$5.15** | **$29.84** |

The aggregate cost comparison is misleading for the same mix reason. Within type:

| Type | AI cost/durable | Control cost/durable |
|------|----------------|---------------------|
| Simple | $2,700/729 = **$3.70** | $400/91 = **$4.40** |
| Complex | $1,200/28 = **$42.86** | $14,400/405 = **$35.56** |

**Interpretation:** AI is cheaper per durable resolution on simple tickets; control is cheaper on complex tickets. The aggregate $5.15 vs. $29.84 gap reflects ticket mix, not efficiency.

**Limitation:** Because assignment was non-random, we cannot attribute any within-type difference to the tool itself; unobserved ticket characteristics may differ even within type.

## 4. Rollout Recommendation

**Do not roll out now.** The vendor's headline (85% vs. 54.5%) is a textbook Simpson's-paradox artifact. After standardizing for mix, the pilot shows AI *underperforming* on resolution rate in both ticket categories. The one encouraging signal—lower cost per durable resolution on simple tickets—is hypothesis-generating only, given non-random assignment and small control-simple sample (n = 100).

## 5. Proposed Randomized Follow-Up

| Element | Design |
|---------|--------|
| **Primary metric** | Durable resolution rate (resolved within 24 h and not reopened within 7 days) |
| **Unit of randomization** | Individual ticket, stratified by type (simple / complex) |
| **Allocation** | 50/50 within each stratum; conceal assignment until ticket is logged |
| **Guardrails** | (a) CSAT ≥ baseline minus 2 pp; (b) reopen rate no more than 3 pp above control; (c) agent-escalation rate monitored weekly |
| **Stopping rule** | Weekly Bayesian monitoring; stop for efficacy if posterior probability that AI is inferior on the primary metric exceeds 0.99, or for futility if after 600 tickets per arm the observed difference is < 1 pp |
| **Duration cap** | 6 weeks maximum regardless of interim results |

## Unknowns

- Whether AI's apparent cost advantage on simple tickets survives randomization.
- Whether agent behavior (editing, overriding drafts) differs systematically across groups.
- Optimal routing rules if AI proves useful only for a subset of ticket types.

**Bottom line:** The pilot is a useful feasibility check, not evidence of benefit. A randomized trial is required before any rollout decision.
GLM 5.3 Flash

Requested model: z-ai/glm-5.3-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0008174500 USD. Excludes retry and judging costs.

# Decision Brief: AI Drafting Assistant Pilot

## 1. Facts (arithmetic shown)

**Aggregate resolution rates (resolved ÷ assigned):**
- AI: (810+40)/1000 = **850/1000 = 85%**
- Control: (95+450)/1000 = **545/1000 = 54.5%**

**Within ticket type:**
- Simple: AI 810/900 = **90%**; Control 95/100 = **95%**
- Complex: AI 40/100 = **40%**; Control 450/900 = **50%**

Note the reversal: Control beats AI in *both* strata, yet AI looks far better in aggregate. This is selection bias, not paradoxical magic: leads assigned AI mostly simple tickets (900/1000 = 90%) and Control mostly complex tickets (900/1000 = 90%).

## 2. Standardized rates (50/50 simple/complex)

AI: 0.5(90%) + 0.5(40%) = **65%**
Control: 0.5(95%) + 0.5(50%) = **72.5%**

On an equal mix, Control outperforms AI by 7.5 points.

## 3. Durable resolutions (resolved − reopened)

**AI:**
- Simple: 810 − 81 = 729; reopen rate among resolved: 81/810 = **10%**
- Complex: 40 − 12 = 28; reopen rate: 12/40 = **30%**
- Total durable: 757. Reopen rate overall: 93/850 = 10.9%

**Control:**
- Simple: 95 − 4 = 91; reopen rate: 4/95 = **4.2%**
- Complex: 450 − 45 = 405; reopen rate: 45/450 = **10%**
- Total durable: 496. Reopen rate overall: 49/545 = 9.0%

**Cost per durable resolution** (cost covers *all* assigned tickets, not just resolved ones):
- AI: $2,700 + $1,200 = $3,900 ÷ 757 = **$5.15**
- Control: $400 + $14,400 = $14,800 ÷ 496 = **$29.84**

**Interpretation and limitation:** AI's cost advantage is dramatic, but the two groups' workloads are not comparable. Control's cost is dominated by complex tickets ($14,400 for 900 complex tickets), while AI's is dominated by simple ones. We have no per-ticket-type cost breakdown, so we cannot compute a like-for-like cost per durable resolution within strata. Also, reopen rates are substantially higher under AI in both strata (10% vs 4.2% simple; 30% vs 10% complex), so part of AI's speed advantage converts into rework. Note reopen counts are subsets of resolved tickets by construction.

## 4. Recommendation

**Do not roll out to every ticket immediately.** The vendor's claim rests on an aggregate comparison of non-randomized groups; when standardized, AI resolves *fewer* tickets within 24h than Control in both strata, with roughly double the reopening. The vendor's 85% vs 54.5% is an artifact of ticket assignment.

**Interpretation:** AI may still be worth deploying narrowly — its cost per durable resolution is ~6× lower, even after accounting for reopens — but the mix confounding means we cannot yet say whether that advantage survives like-for-like comparison, and the elevated reopen rate is a genuine quality concern, especially on complex tickets (40% resolution, 30% reopens).

**Unknowns:** per-type costs; whether leads assigned tickets to AI based on difficulty or agent availability; staff-level variation; long-term reopen behavior beyond 7 days; and whether the reopen gap is caused by AI drafts or by which agents used it.

## 5. Proposed randomized follow-up

**Design:** Stratified RCT, 4 weeks, across both ticket types. Within each stratum, randomize tickets 50/50 AI vs Control at assignment time. (A 50/50 split prioritizes estimating the effect cleanly; the vendor's tool is already available, so no one is denied a proven intervention.)

**Primary metric:** Durable resolution rate — proportion of assigned tickets resolved within 24h **and** not reopened within 7 days. This directly combines speed and quality and avoids rewarding fast-but-flimsy resolutions.

**Guardrails (pre-specified failure thresholds, monitored weekly):**
1. Reopen rate among resolved tickets rises by >5 percentage points vs Control.
2. Total handling cost per assigned ticket exceeds Control by >20%.

Breaching either guardrail in a stratum pauses that stratum's arm pending review.

**Stopping rule:** Run the full 4 weeks unless a guardrail is breached, in which case stop enrollment in the affected arm immediately. No early "success" stopping — we want stable 7-day reopen data before deciding.

**Decision rule:** If AI beats Control on the primary metric in simple tickets with guardrails intact, deploy to simple tickets only; complex tickets get a separate evaluation given current evidence suggests AI underperforms there.
Hy4 preview

Requested model: tencent/hy4-preview · Output budget: 16384 tokens · Requested reasoning: low · Sep 28, 2026, 17:58 UTC

Reported answer cost: $0.0123383890 USD. Excludes retry and judging costs.

**Facts (computed from the table).**  
- AI group: 1,000 assigned (900 simple, 100 complex). Control: 1,000 assigned (100 simple, 900 complex).  
- Within-type 24h resolution rates:  
  - AI simple = 810/900 = **90%**; AI complex = 40/100 = **40%**.  
  - Control simple = 95/100 = **95%**; control complex = 450/900 = **50%**.  
- Aggregate: AI resolved 850/1000 = **85%**; control 545/1000 = **54.5%** (the vendor’s comparison).

**Interpretation.** The aggregate gap is a workload-mix artifact, not evidence AI is better: AI saw 90% simple tickets; control saw 90% complex. Standardized to a 50/50 simple/complex workload:  
- AI = 0.5(90%) + 0.5(40%) = **65%**.  
- Control = 0.5(95%) + 0.5(50%) = **72.5%**.  
On the same mix, control resolves more.

**Durable resolutions (resolved − reopened).**  
- AI: simple 810−81 = 729; complex 40−12 = 28; total **757**. Standardized durable rate = 0.5(729/900) + 0.5(28/100) = 0.5(81%) + 0.5(28%) = **54.5%**. Total cost = $2,700+$1,200 = **$3,900**; cost per durable = 3,900/757 = **$5.15**.  
- Control: simple 95−4 = 91; complex 450−45 = 405; total **496**. Standardized durable rate = 0.5(91%) + 0.5(45%) = **68%**. Total cost = $400+$14,400 = **$14,800**; cost per durable = 14,800/496 = **$29.84**.  

**Ticket-mix limitation.** AI handled mostly cheap simple tickets; control mostly expensive complex ones. Within type:  
- Simple cost/durable: AI $2,700/729 = **$3.70** vs control $400/91 = **$4.40**.  
- Complex cost/durable: AI $1,200/28 = **$42.86** vs control $14,400/405 = **$35.56**.  
AI is cheaper only on simple; on complex it is worse on both quality and cost. Reopen shares among resolved: AI simple 81/810=10% vs control 4/95≈4.2%; AI complex 12/40=30% vs control 45/450=10%.

**Unknowns.** Assignment was not randomized; leads chose tickets. Therefore no causal AI effect, staff effect, handling-time effect, or CSAT effect can be inferred from these numbers alone.

**Recommendation.** Reject the vendor’s immediate rollout to all tickets. The data do not show within-type benefit; AI is worse on complex tickets and only appears better in aggregate because of the easier workload mix. At most, restrict AI to simple tickets under monitoring; do not use it on complex tickets. A clean randomized test is required before any expansion.

**Follow-up randomized experiment.** Randomize incoming tickets 1:1 to AI drafting vs usual process, stratified by ticket type so each arm sees the same simple/complex proportion (or randomize within type and equalize the mix analytically). Primary metric: durable 24h resolution rate = (resolved within 24h − reopened within 7 days) / assigned. Guardrails: type-specific 24h resolution rate, reopen proportion, cost per assigned ticket, and cost per durable resolution. Stopping rule: analyze after a fixed four-week accrual period; pause the experiment immediately if complex-ticket durable resolution falls below control or AI cost per durable exceeds control by more than 20%. No rollout should occur during the experiment.
Kimi K3

Requested model: moonshotai/kimi-k3 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:26 UTC

Reported answer cost: $0.0139370000 USD. Excludes retry and judging costs.

# Decision Brief: AI Drafting Assistant Pilot

## 1. Resolution rates (facts)

**Aggregate (what the vendor cites):**
- AI: (810+40)/1000 = **85.0%**
- Control: (95+450)/1000 = **54.5%**

**Within ticket type:**
| | AI | Control |
|---|---|---|
| Simple | 810/900 = **90.0%** | 95/100 = **95.0%** |
| Complex | 40/100 = **40.0%** | 450/900 = **50.0%** |

**Interpretation:** The vendor's headline is arithmetically true but misleading. Within *each* ticket type, the control group resolved *more*. The aggregate flips only because AI was assigned 90% simple tickets while control got 90% complex (Simpson's paradox). Because team leads chose which tickets used AI, this is almost certainly selection, not AI performance.

## 2. Rates standardized to a 50/50 workload

- AI: 0.5(90.0%) + 0.5(40.0%) = **65.0%**
- Control: 0.5(95.0%) + 0.5(50.0%) = **72.5%**

At identical workload mix, control outperforms AI by 7.5 points.

## 3. Durable resolutions and cost

Durable = resolved − reopened:

| | Durable | Total cost | Cost/durable |
|---|---|---|---|
| AI | (810−81)+(40−12) = 729+28 = **757** | $3,900 | **$5.15** |
| Control | (95−4)+(450−45) = 91+405 = **496** | $14,800 | **$29.84** |

**Critical limitation:** AI's far lower cost per durable resolution is driven by ticket mix — AI processed 900 cheap simple tickets; control processed 900 expensive complex ones. The comparison is not apples-to-apples. Within type, per durable resolution: simple — AI $3,900... rather, simple: AI $2,700/729 = **$3.70** vs control $400/91 = **$4.40** (AI modestly cheaper); complex: AI $1,200/28 = **$42.86** vs control $14,400/405 = **$35.56** (AI *more* expensive, driven by its lower complex resolution rate). Reopen rates also favor control (AI: 10.0% simple, 30.0% complex; control: 4.2% simple, 10.0% complex).

## 4. Recommendation

**Do not roll out to every ticket now.** Facts: AI is associated with lower within-type resolution, higher reopen rates, and higher cost per durable resolution on complex tickets. Interpretation: the aggregate "win" is an artifact of non-random assignment. Unknowns: whether AI *caused* worse within-type outcomes or whether leads routed their hardest simple/complex tickets to AI; staff learning curves; customer satisfaction. Justified next step: a randomized experiment (below). Optionally, a limited continued pilot on simple tickets only is defensible given the $3.70 vs $4.40 cost edge, but even that rests on confounded data.

## 5. Randomized follow-up experiment

- **Design:** Randomize each incoming ticket 50/50 to AI-assisted vs control, stratified by ticket type (simple/complex) so mix cannot confound results. Randomize at ticket level; block by week to balance load.
- **Primary metric:** Durable resolution rate (resolved within 24h and not reopened within 7 days), reported separately by stratum and pooled.
- **Guardrails:** Reopen rate, total handling cost per durable resolution, escalation rate, and a customer-satisfaction measure if available; any guardrail materially worse for AI triggers review.
- **Allocation:** 50/50 within each stratum; run a pre-announced fixed duration (e.g., 4 weeks) rather than stopping on early results.
- **Stopping rule:** Stop early only for a pre-specified guardrail breach (e.g., AI reopen rate exceeds control by a set margin sustained for one week); otherwise run to the planned end date before analyzing. Decision rule declared in advance: proceed to rollout only if AI is non-inferior on durable resolution within each stratum *and* cheaper per durable resolution.

## Summary

The vendor's 85% vs 54.5% claim is factually accurate but causally unsupported. Standardized to equal workload, control wins 72.5% vs 65.0%. Cost savings are real-looking but mix-confounded. Fund the randomized trial before any enterprise rollout.
GLM 5.3 Prime

Requested model: z-ai/glm-5.3-prime · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0124548000 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot**

**Facts (from pilot data)**

Aggregate resolution: AI 850/1000 = 85%; Control 545/1000 = 54.5% (matches vendor).

Within ticket type:
- Simple: AI 810/900 = 90%; Control 95/100 = 95%
- Complex: AI 40/100 = 40%; Control 450/900 = 50%

Reopen rates (of resolved): AI simple 81/810 = 10%; AI complex 12/40 = 30%; Control simple 4/95 ≈ 4.2%; Control complex 45/450 = 10%.

Costs: AI $3,900 total ($2,700 + $1,200); Control $14,800 ($400 + $14,400). Per assigned ticket: AI $3.90; Control $14.80.

**Standardized to a 50/50 simple/complex workload**

Average of within-type resolution rates: AI (90% + 40%)/2 = 65%; Control (95% + 50%)/2 = 72.5%. Under this workload, the AI group's apparent aggregate advantage reverses; Control is 7.5 points higher. The vendor's 85% vs 54.5% headline is an artifact of ticket mix, not performance. This is Simpson's-style aggregation: the AI group handled mostly simple tickets (900/1000), the Control group mostly complex (900/1000).

**Durable resolutions and cost**

Durable = resolved − reopened:
- AI: simple 810 − 81 = 729; complex 40 − 12 = 28; total 757
- Control: simple 95 − 4 = 91; complex 450 − 45 = 405; total 496

Durable/assigned: AI 757/1000 = 75.7%; Control 496/1000 = 55.1%. But mix again: standardized durable rates are AI (729/900 + 28/100)/2 = (81% + 28%)/2 = 54.5%; Control (91/100 + 405/900)/2 = (91% + 45%)/2 = 68%. Control is higher within both types.

Cost per durable resolution: AI $3,900/757 ≈ $5.15; Control $14,800/496 ≈ $29.84. The AI group is dramatically cheaper per durable resolution, but this is confounded by mix: durable rates within type favor Control, while cost per assigned ticket is much lower for the AI group in both strata (AI simple $2,700/900 = $3.00 vs Control simple $4.00; AI complex $1,200/100 = $12 vs Control complex $14,400/900 = $16). So AI-assisted tickets were cheaper per ticket in both types, but resolved and stayed resolved at lower rates in both types.

**Interpretation and unknowns**

Interpretations: the aggregate gains the vendor cites are driven by ticket mix. Within type, AI-assisted tickets resolved less often and reopened more often, but cost less per assigned ticket. Plausible reading: AI saves handling cost (shorter/faster tickets) at the price of quality; or team leads routed only "easy" or high-confidence tickets to AI, and even those did worse; or the observed differences reflect lead judgment, not the tool. We cannot tell.

Unknowns: no randomization, so no causal claim is possible. We lack time-to-resolution beyond 24h, customer satisfaction, draft-edit rates, staff-level data, and whether "resolved" definitions were applied consistently. Costs may include different staffing mixes.

**Recommendation**

Do not roll out to all tickets. Extend a limited, randomized trial before any org-wide deployment. If leadership wants interim savings, the data are consistent with a small, monitored deployment on simple tickets only, but even that within-type resolution gap (90% vs 95%) argues for caution. A quality-preserving rollout is not yet justified.

**Follow-up experiment**

- **Design:** Randomize eligible incoming tickets at ticket-creation level, stratified by type (simple/complex), roughly 50/50 within each stratum, for two to four weeks. Randomization at ticket level, not team-lead discretion, removes the selection bias that invalidates this pilot.
- **Primary metric:** durable resolution rate (resolved within 24h and not reopened within 7 days) per assigned ticket, analyzed within strata and reported standardized to the actual workload mix.
- **Guardrails:** (a) reopen rate within each stratum not materially worse than control; (b) handling cost per assigned ticket in each stratum; (c) escalation or repeat-contact signals monitored during the run.
- **Stopping rules:** stop early if a stratum's reopen rate exceeds a pre-set threshold agreed with the director, or if a data-quality issue (mis-assignment, broken tracking) is found. Otherwise run the full pre-agreed duration; do not stop on a favorable interim look.
- **Pre-commit analysis plan:** compare within strata first; aggregate only using the true workload mix. Report cost per durable resolution as a secondary metric.

Bottom line: the vendor's 85% claim is a mix artifact. The honest summary is "cheaper per ticket, worse outcomes within type, cause unknown" — and only randomization can fix that.
Hy3

Requested model: tencent/hy3 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0009387675 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot**

**1. Resolution rates (Facts)**
- Aggregate: AI resolved 850/1000 = 85%; Control resolved 545/1000 = 54.5%. (Vendor claim uses these.)
- Within type:
  - Simple: AI 810/900 = 90%; Control 95/100 = 95%.
  - Complex: AI 40/100 = 40%; Control 450/900 = 50%.
*Interpretation:* AI’s aggregate advantage is entirely due to mix (AI handled 90% simple, Control 90% complex). Within each type, Control resolves more.

**2. Standardized to 50/50 simple/complex (Interpretation)**
Equal-weight average of type rates:
- AI: (0.90 + 0.40)/2 = 65%.
- Control: (0.95 + 0.50)/2 = 72.5%.
*Fact basis:* arithmetic shown. *Unknown:* true operational mix post-rollout.

**3. Durable resolutions & cost (Facts, with mix limitation)**
Durable = resolved − reopened.
- AI Simple: 810−81=729 durable; cost $2700 → $3.70/durable.
- AI Complex: 40−12=28 durable; cost $1200 → $42.86/durable.
- AI Total: 757 durable; total cost $3900 → $5.15/durable.
- Control Simple: 95−4=91 durable; cost $400 → $4.40/durable.
- Control Complex: 450−45=405 durable; cost $14400 → $35.56/durable.
- Control Total: 496 durable; total cost $14800 → $29.84/durable.

*Interpretation:* Raw cost/durable favors AI, but AI’s workload was 90% cheap simple tickets. Standardized 50/50 (500 each type): AI cost = $1500+$6000=$7500, durable = 405+140=545 → $13.76/durable. Control = $2000+$8000=$10000, durable = 455+225=680 → $14.71/durable. *Limitation:* ticket-mix confounding remains; no causal claims.

**4. Recommendation (Interpretation)**
Do **not** roll out to every ticket immediately. Vendor’s aggregated 85% vs 54.5% is an artifact of unequal assignment. Within-type, AI shows lower resolution. Cost per durable is only marginally better under balanced mix. *Unknown:* why leads chose tickets (selection bias likely). A controlled test is required.

**5. Randomized follow-up experiment (Proposal)**
- **Primary metric:** Durable resolution rate (resolved ≤24h minus reopened ≤7d), reported per type and as 50/50 standardized composite.
- **Guardrails:** Cost per durable not >110% of control arm; reopen rate increase <5 percentage points vs control.
- **Allocation:** Ticket-level randomization, stratified by simple/complex, 1:1 AI vs control, target 1000 tickets per arm per type.
- **Stopping rule:** After 2 weeks or full enrollment, if AI durable rate is ≥2pp below control in either type, or guardrail breached, halt expansion.

*Unknowns:* true effect size, staff behavior, long-term reopen rates beyond 7 days. No confidence intervals or sample-size calculations invented.
Gemini 3.8 Flash

Requested model: google/gemini-3.8-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:24 UTC

Reported answer cost: $0.0054255000 USD. Excludes retry and judging costs.

### Decision Brief: AI Drafting Assistant Pilot Evaluation

#### 1. Resolution Rates (Aggregate and Within-Ticket-Type)
* **Aggregate Resolution Rates (Fact):**
  * **AI:** $(810 + 40) / (900 + 100) = 850 / 1{,}000 = \mathbf{85.0\%}$
  * **Control:** $(95 + 450) / (100 + 900) = 545 / 1{,}000 = \mathbf{54.5\%}$
* **Within-Type Resolution Rates (Fact):**
  * **Simple Tickets:** AI resolved $810 / 900 = \mathbf{90.0\%}$; Control resolved $95 / 100 = \mathbf{95.0\%}$.
  * **Complex Tickets:** AI resolved $40 / 100 = \mathbf{40.0\%}$; Control resolved $450 / 900 = \mathbf{50.0\%}$.
* **Interpretation:** The vendor’s aggregate claim is an artifact of Simpson’s Paradox. Team leads disproportionately assigned simple tickets to AI (90% of AI workload) and complex tickets to Control (90% of Control workload). Within both ticket types, AI exhibited lower 24-hour resolution rates than Control.

---

#### 2. Standardized Resolution Rates (50/50 Workload)
To adjust for the skewed ticket allocation, we standardize both groups to a hypothetical 50% simple / 50% complex distribution:
* **AI Standardized Rate:** $(0.50 \times 90.0\%) + (0.50 \times 40.0\%) = 45.0\% + 20.0\% = \mathbf{65.0\%}$
* **Control Standardized Rate:** $(0.50 \times 95.0\%) + (0.50 \times 50.0\%) = 47.5\% + 25.0\% = \mathbf{72.5\%}$
* **Interpretation:** On an identical workload split, Control achieved a 7.5 percentage-point higher resolution rate than AI.

---

#### 3. Durable Resolutions and Cost per Durable Resolution
*Durable Resolution = Resolved within 24h − Reopened within 7 days.*

* **Counts & Costs by Subgroup (Fact):**
  * **AI Simple:** $810 - 81 = 729$ durable resolutions; Cost = $\$2{,}700$
  * **AI Complex:** $40 - 12 = 28$ durable resolutions; Cost = $\$1{,}200$
  * **Control Simple:** $95 - 4 = 91$ durable resolutions; Cost = $\$400$
  * **Control Complex:** $450 - 45 = 405$ durable resolutions; Cost = $\$14{,}400$
* **Group Totals (Fact):**
  * **AI Group:** 
    * Total Durable Resolutions: $729 + 28 = \mathbf{757}$
    * Total Cost: $\$2{,}700 + \$1{,}200 = \mathbf{\$3{,}900}$
    * Cost per Durable Resolution: $\$3{,}900 / 757 \approx \mathbf{\$5.15}$
  * **Control Group:** 
    * Total Durable Resolutions: $91 + 405 = \mathbf{496}$
    * Total Cost: $\$400 + \$14{,}400 = \mathbf{\$14{,}800}$
    * Cost per Durable Resolution: $\$14{,}800 / 496 \approx \mathbf{\$29.84}$
* **Ticket-Mix Limitation (Interpretation & Fact):** AI’s lower overall cost per durable resolution ($\$5.15$ vs. $\$29.84$) is driven by ticket mix: simple tickets are cheaper to handle. Looking within types:
  * **Simple:** AI cost per durable is $\$2{,}700 / 729 \approx \mathbf{\$3.70}$; Control is $\$400 / 91 \approx \mathbf{\$4.40}$.
  * **Complex:** AI cost per durable is $\$1{,}200 / 28 \approx \mathbf{\$42.86}$; Control is $\$14{,}400 / 405 \approx \mathbf{\$35.56}$.
  AI was cheaper on simple tickets, but substantially more expensive per durable resolution on complex tickets.

---

#### 4. Justified Rollout Recommendation
**Recommendation: Reject full rollout.**
* **Rationale:** The observational pilot lacked randomization. Non-standardized aggregate gains vanish once stratified by complexity. AI underperformed Control on 24h resolution for both simple (90% vs. 95%) and complex (40% vs. 50%) tickets, and had higher reopen rates ($81/810 = 10\%$ vs. $4/95 \approx 4.2\%$ on simple; $12/40 = 30\%$ vs. $45/450 = 10\%$ on complex).
* **Unknowns:** Staff-level productivity, agent sentiment, vendor software subscription fees, and actual baseline volume proportions across the entire support operation.

---

#### 5. Randomized Follow-Up Experiment Plan
* **Allocation:** 1:1 automated randomization at ticket creation, stratified by ticket type (Simple vs. Complex), ensuring 50% AI and 50% Control within each stratum.
* **Primary Metric:** Durable Resolution Rate (Durable Resolutions / Assigned Tickets).
* **Guardrails:**
  * Reopen rate must not exceed 12% in Simple or 15% in Complex.
  * Cost per assigned ticket must not exceed Control baseline.
* **Stopping Rule:** Halt the experiment immediately if the 7-day rolling reopen rate in the AI arm exceeds the Control arm by 5 percentage points or more within any stratum.
Space Bunny Alpha

Requested model: stealth/space-bunny-alpha · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:29 UTC

Reported answer cost: $0.0000000000 USD. Excludes retry and judging costs.

## Decision brief

### Facts

**Observed resolution rates**

- **AI:** \((810+40)/(900+100)=850/1,000=\mathbf{85.0\%}\)
- **Control:** \((95+450)/(100+900)=545/1,000=\mathbf{54.5\%}\)
- **AI, simple:** \(810/900=\mathbf{90.0\%}\)
- **Control, simple:** \(95/100=\mathbf{95.0\%}\)
- **AI, complex:** \(40/100=\mathbf{40.0\%}\)
- **Control, complex:** \(450/900=\mathbf{50.0\%}\)

**Standardized to a 50/50 workload**

Using the within-type rates:

- **AI:** \((90\%+40\%)/2=\mathbf{65.0\%}\)
- **Control:** \((95\%+50\%)/2=\mathbf{72.5\%}\)

Thus, standardization reverses the aggregate ranking.

**Durable resolutions and cost**

Durable resolutions equal resolved tickets minus reopened tickets:

- **AI:** \(850-93=\mathbf{757}\), or **75.7% of assigned tickets**
- **Control:** \(545-49=\mathbf{496}\), or **49.6% of assigned tickets**

Cost per durable resolution:

- **AI:** \(\$3,900/757=\mathbf{\$5.15}\)
- **Control:** \(\$14,800/496=\mathbf{\$29.84}\)

The aggregate cost comparison is driven partly by ticket mix: AI handled 90% simple tickets, while control handled only 10% simple tickets. Within types, the contrast also differs:

- AI simple: \(\$2,700/(810-81)=\$3.70\); control simple: \(\$400/(95-4)=\$4.40\)
- AI complex: \(\$1,200/(40-12)=\$42.86\); control complex: \(\$14,400/(450-45)=\$35.56\)

### Interpretation

The observed 30.5-point aggregate advantage is not persuasive evidence that AI caused better performance. Ticket type explains much of it: control was better on both observed within-type resolution rates, and remained better after simple/complex standardization. AI nevertheless produced more observed durable resolutions and much lower aggregate cost per durable resolution, especially because its workload was predominantly simple.

### Unknowns

Because team leads selected tickets, the pilot cannot establish causality. We do not know whether leads routed easier, more urgent, or otherwise unusual tickets to AI; whether results generalize to other teams, times, or workloads; or impacts on quality, customer satisfaction, staff burden, and performance beyond seven days. Ticket counts are not a basis for confidence intervals or a causal estimate.

### Recommendation and randomized follow-up

**Do not roll out immediately to every ticket.** Consider limited use on simple tickets only if operational safeguards and monitoring are acceptable, but retain the broader rollout decision for a trial.

Randomize tickets **1:1 to AI or control**, stratified by simple versus complex ticket type, using a consistent assignment mechanism. Keep handling processes otherwise comparable. Track every assigned ticket through at least **seven days after closure** so reopening is observable.

- **Primary metric:** durable resolution rate per assigned ticket  
  \[
  (\text{resolved within 24h}-\text{reopened within 7 days})/\text{assigned}
  \]
- **Guardrails:** reopened tickets as a percentage of resolved tickets; total handling cost per assigned ticket; unresolved rate; and reopen rate within each ticket type.
- **Decision rule:** after all enrolled tickets are mature, deploy broadly only if AI’s overall durable-resolution rate is at least control’s, with no material deterioration in either ticket type or guardrail.
- **Stopping rule:** pause AI if, at a predefined operational review, its durable-resolution rate is more than 10 percentage points below control in either stratum, or its reopen rate exceeds control by more than 5 points. These thresholds are management choices, not pilot-derived safety limits.
GPT-6 Sol

Requested model: openai/gpt-6-sol · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0100140000 USD. Excludes retry and judging costs.

## Decision brief

**Recommendation:** Do not roll AI out to every ticket. Run a randomized follow-up while keeping use limited to a monitored pilot. The vendor’s comparison is arithmetically correct but misleading: leads selected tickets for AI, and 90% of AI tickets were simple versus 10% of control tickets.

### What the pilot shows

| Measure | AI | Control |
|---|---:|---:|
| Overall resolved within 24 hours | 850/1,000 = **85%** | 545/1,000 = **54.5%** |
| Simple resolved within 24 hours | 810/900 = **90%** | 95/100 = **95%** |
| Complex resolved within 24 hours | 40/100 = **40%** | 450/900 = **50%** |
| Resolution rate standardized to a 50/50 simple/complex workload | (90% + 40%)/2 = **65%** | (95% + 50%)/2 = **72.5%** |
| Durable resolutions (resolved minus reopened) | 850 − (81 + 12) = **757** | 545 − (4 + 45) = **496** |
| Cost per durable resolution | $3,900/757 = **$5.15** | $14,800/496 = **$29.84** |

**Interpretation:** The aggregate resolution advantage reverses both within each ticket type and after 50/50 standardization. That is strong reason to reject the vendor’s rollout argument, **not** proof that AI reduces resolution: selection could differ within ticket types too. AI’s lower observed cost per durable resolution is promising, but it also reflects very different ticket mixes. Cost covers *all assigned tickets*; the durable-resolution denominator does not.

**Unknowns:** Whether AI changes outcomes or costs under comparable assignment, and whether any benefit or harm differs by ticket type.

### Randomized follow-up

Randomly assign eligible incoming tickets **1:1 to AI drafting or control within each ticket type**. Keep normal review and handling procedures otherwise consistent, and analyze tickets as assigned. Run for a prespecified **four-week enrollment period**, then wait until every enrolled ticket’s reopening window can be assessed.

The **primary metric** is the durable-resolution rate per assigned ticket—closed within 24 hours and not reopened within seven days—standardized to a 50/50 simple/complex workload. Report each type separately as well. Guardrails are the within-type 24-hour resolution rate, reopen rate among resolved tickets, and handling cost per assigned ticket; monitor quality complaints or escalations if recorded consistently.

Do not stop early for an apparent benefit. Review guardrails weekly and pause AI assignment for safety investigation if, after at least 100 resolved tickets per arm in a ticket type, its reopen rate exceeds control by **5 percentage points or more**. At the planned end, consider broader rollout only if the primary result and both ticket-type results support it without unacceptable guardrail deterioration.
MiniMax M2.7 (Nitro)

Requested model: minimax/minimax-m2.7:nitro · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0059898000 USD. Excludes retry and judging costs.

**Decision brief – AI drafting‑assistant pilot**  
*Support Operations Lead | ≤ 650 words*

---

### 1.  Aggregate and within‑type resolution rates  

| Group | Ticket type | Assigned | Resolved ≤ 24 h | Reopened ≤ 7 d | Resolved % |
|-------|-------------|----------|------------------|----------------|------------|
| AI    | Simple      | 900      | 810              | 81             | 810/900 = 90 % |
| AI    | Complex     | 100      | 40               | 12             | 40/100 = 40 % |
| Ctrl  | Simple      | 100      | 95               | 4              | 95/100 = 95 % |
| Ctrl  | Complex     | 900      | 450              | 45             | 450/900 = 50 % |

**Aggregate (all tickets)**  
- AI: (810 + 40) / 1 000 = 85 %  
- Control: (95 + 450) / 1 000 = 54.5 %  

*Interpretation*: Raw totals heavily reflect the ticket‑mix (AI = 90 % simple, Control = 90 % complex). Within each type the AI assistant actually **underperforms** the control (90 % vs 95 % on simple; 40 % vs 50 % on complex).  

---

### 2.  Standardised rates for a 50 % / 50 % simple‑complex workload  

Weighted resolution = ½ × simple % + ½ × complex %  

- AI standardised = 0.5 × 90 % + 0.5 × 40 % = **65 %**  
- Control standardised = 0.5 × 95 % + 0.5 × 50 % = **72.5 %**  

*Interpretation*: After adjusting for an equal mix of ticket types, the control group resolves a larger share of tickets.  

---

### 3.  Durable resolutions and cost per durable resolution  

Durable = Resolved ≤ 24 h – Reopened ≤ 7 d  

| Group | Resolved | Reopened | Durable | Total handling cost (USD) | Cost / Durable |
|-------|----------|----------|---------|---------------------------|----------------|
| AI    | 850      | 81 + 12 = 93 | 757 | 2 700 + 1 200 = 3 900 | 3 900 / 757 ≈ **$5.15** |
| Ctrl  | 545      | 4 + 45 = 49  | 496 | 400 + 14 400 = 14 800 | 14 800 / 496 ≈ **$29.84** |

*Interpretation*: AI tickets cost roughly **six‑times less** per durable resolution.  
*Limitation*: The AI arm contains far more low‑cost simple tickets; the cost advantage may stem from ticket mix rather than AI efficiency.  

---

### 4.  Rollout recommendation  

**Fact**: The pilot was observational; leads chose which tickets used AI, creating a strong selection bias.  
**Fact**: Within both ticket types the AI assistant shows lower resolution rates than the control.  
**Interpretation**: The raw 85 % vs 54.5 % headline figure is driven by the uneven mix, not by AI superiority.  
**Recommendation**: **Do not roll out to all tickets now.** The cost‑per‑durable‑resolution advantage is promising but not conclusive because of confounding ticket complexity and selection bias.  

A **targeted, limited rollout** can be considered for simple tickets only (where AI resolves 90 % vs control 95 % but at a lower cost) – but this should be validated under controlled conditions before broader deployment.

---

### 5.  Follow‑up randomized experiment  

| Element | Detail |
|---------|--------|
| **Primary metric** | Durable resolution rate per ticket (resolved ≤ 24 h and not reopened ≤ 7 d). |
| **Allocation** | Random assignment of tickets **within each type** (simple / complex). Aim for at least 300 simple and 300 complex tickets per arm (≈ 1 200 tickets total). |
| **Guardrails** | – No AI‑only “preferred” ticket selection; all incoming tickets are pooled and randomized daily.<br>– Define a minimum detectable difference of **5 percentage points** in durable resolution (absolute) as a practical threshold for success.<br>– Track reopened rate and handling cost as secondary indicators; pause if reopened rate exceeds control by > 3 pp. |
| **Stopping rule** | – Review after **600 tickets** (≈ half the target).<br>– If AI durable resolution is **≥ 5 pp lower** than control in either stratum, stop for redesign.<br>– If cost per durable resolution exceeds control’s current $30 threshold, pause for cost analysis. |
| **Timeline** | 4‑week enrollment; interim check at week 2; final analysis at week 5. |

*Unknown*: Whether the cost per durable resolution observed in the pilot holds under true randomization; also whether the AI tool’s performance changes over time as agents become more familiar with it.

---

### 6.  Bottom line  

- The vendor’s “85 % vs 54.5 %” headline overstates the AI’s impact because of a non‑random, imbalanced ticket mix.  
- Within ticket types, the control group resolves a higher share of tickets, and the standardised resolution rate favors the control (72.5 % vs 65 %).  
- The AI arm shows a **much lower cost per durable resolution** ($5 vs $30), but this could be an artifact of the simple‑ticket dominance.  
- **Recommendation**: Hold a randomised trial (as outlined) to isolate the true effect before committing to a full rollout. This protects the organisation from investing in a tool whose benefits may not be causal or generalisable.
DeepSeek V4 Pro

Requested model: deepseek/deepseek-v4-pro-0813 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:29 UTC

Reported answer cost: $0.0130567110 USD. Excludes retry and judging costs.

**Decision brief: AI drafting assistant pilot**

**Facts**
Aggregate resolution within 24h:
- AI: \(850/1000 = 85.0\%\)
- Control: \(545/1000 = 54.5\%\)

Within ticket type:
- Simple: AI \(810/900 = 90.0\%\); Control \(95/100 = 95.0\%\)
- Complex: AI \(40/100 = 40.0\%\); Control \(450/900 = 50.0\%\)

Standardized to a 50/50 simple/complex workload:
- AI: \((0.90 + 0.40)/2 = 65.0\%\)
- Control: \((0.95 + 0.50)/2 = 72.5\%\)
- AI is 7.5 percentage points lower after removing mix distortion.

Durable resolutions = resolved minus reopened:
- AI simple: \(810 - 81 = 729\)
- AI complex: \(40 - 12 = 28\)
- AI total: \(729 + 28 = 757\)
- Control simple: \(95 - 4 = 91\)
- Control complex: \(450 - 45 = 405\)
- Control total: \(91 + 405 = 496\)

Cost per durable resolution:
- AI total: \(2700 + 1200 = 3900\); \(3900/757 = \$5.15\)
- Control total: \(400 + 14400 = 14800\); \(14800/496 = \$29.84\)

By type:
- AI simple: \(2700/729 = \$3.70\); Control simple: \(400/91 = \$4.40\)
- AI complex: \(1200/28 = \$42.86\); Control complex: \(14400/405 = \$35.56\)

**Interpretation**
The vendor’s aggregate comparison is misleading because the AI group was 90% simple, while the control group was 10% simple and 90% complex. Within both ticket types, AI resolved fewer tickets within 24 hours. Durable resolution rates were also worse for AI by type:
- Simple durable: AI \(729/900 = 81.0\%\); Control \(91/100 = 91.0\%\)
- Complex durable: AI \(28/100 = 28.0\%\); Control \(405/900 = 45.0\%\)

The AI cost per durable looks favorable overall only because simple tickets dominate the AI mix. In complex tickets, AI cost per durable was higher than control. The observational design had no randomization; team leads chose AI tickets. That may confound results.

**Recommendation**
Do not roll out AI immediately to every ticket. The observed improvement is a mix effect, not evidence that AI performs better. Within ticket type, AI was worse on resolution and durable resolution. Complex-ticket rollout is not justified. Simple-ticket AI may have lower cost per durable but worse quality; it should be studied further before any limited rollout.

**Randomized follow-up experiment**
- **Design:** Randomize eligible tickets 1:1 within simple and complex strata to AI-assisted drafting or usual handling.
- **Primary metric:** Durable resolution rate per assigned ticket: resolved within 24 hours and not reopened within 7 days.
- **Guardrails:** per ticket type, monitor reopened rate, cost per durable, and manual QA. Stop if reopening worsens by more than 5 percentage points or cost per durable exceeds control by more than 10%.
- **Stopping rule:** At each review by type, stop AI treatment if durable resolution is more than 5 percentage points below control or any guardrail is breached. Evaluate AI separately for simple and complex tickets.
Qwen3.7 Flash

Requested model: qwen/qwen3.7-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0007249900 USD. Excludes retry and judging costs.

**Decision Brief: AI Assistant Pilot Evaluation**

**Facts, Interpretations, & Unknowns**
*Facts:* AI group: 1,000 tickets (900 simple, 100 complex). Control group: 1,000 tickets (100 simple, 900 complex). Team leads non-randomly assigned tickets. Vendor’s 54.5%→85% claim matches raw aggregates. Cost figures cover all assigned tickets. Reopened counts are subsets of resolved.
*Interpretations:* Assignment bias explains the aggregate gap. Leads routed high-success/simple tickets to AI and low-success/complex ones to Control. Within identical ticket types, Control outperformed AI. Vendor metrics reflect routing artifacts, not tool superiority.
*Unknowns:* True isolated impact of AI on durable outcomes. Cost variance drivers beyond mix. Agent skill transfer or escalation patterns. No causal inference can be drawn from current observational data.

**1. Aggregate & Within-Type Resolution Rates**
*Aggregate:* AI: 850/1,000 = 85.0%. Control: 545/1,000 = 54.5%. Matches vendor but conflates mix.
*By Type:* 
Simple: AI 810/900 = 90.0%; Control 95/100 = 95.0%.
Complex: AI 40/100 = 40.0%; Control 450/900 = 50.0%.
Control resolves faster within both categories.

**2. Standardized Rates (50/50 Workload)**
Weighted equally (500 simple, 500 complex):
AI: (0.5 × 0.90) + (0.5 × 0.40) = 0.650 (65.0%)
Control: (0.5 × 0.95) + (0.5 × 0.50) = 0.725 (72.5%)
Adjusted for a balanced queue, Control outperforms AI by 7.5 percentage points.

**3. Durable Resolutions & Cost per Durable Resolution**
Durable = Resolved − Reopened (closed ≥7 days).
AI: (810−81)+(40−12) = 757 durable. Cost: $2,700+$1,200 = $3,900. → $3,900/757 = $5.15/durable.
Control: (95−4)+(450−45) = 496 durable. Cost: $400+$14,400 = $14,800. → $14,800/496 = $29.84/durable.
*Mix Limitation:* AI’s lower cost/durable stems from handling 90% simple tickets, which carry minimal handling costs. Control’s figure is structurally inflated by processing 90% complex tickets. Direct aggregate comparison is invalid until mix is neutralized.

**4. Rollout Recommendation**
Do not roll out. The pilot demonstrates that favorable ticket routing, not AI capability, drove the reported gains. Under matched ticket types and a standardized workload, the control workflow resolves more tickets durably and at higher speed. Expanding now would scale inefficient routing and likely increase long-term reopen volume and cost. Recommend pausing expansion pending a controlled trial.

**5. Randomized Follow-Up Experiment**
*Design:* Strict probability-based queue routing. Eliminate lead discretion to prevent self-selection.
*Primary Metric:* Durable resolution rate (total durable ÷ total assigned).
*Guardrails:* Halt allocation to either arm if durable rate falls >15% below pre-pilot baseline for two consecutive weeks, or if the 7-day reopen spike exceeds 10% week-over-week. Require human verification before closing complex or escalated tickets in both arms.
*Allocation:* 50% AI, 50% Control. Log weekly mix to verify balance; adjust if routing system drifts >5% off target.
*Stopping Rule:* Minimum 4-week observation period to absorb weekday/weekend variation. Interim review only at week 4. Continue if durable rate improves ≥3 pp with no guardrail breaches; otherwise, terminate the pilot and revert to manual workflows. No effect sizes, confidence bounds, or staff productivity proxies are assumed or required for this stage.
gpt-oss-120b

Requested model: openai/gpt-oss-120b · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0004929500 USD. Excludes retry and judging costs.

**Decision Brief – AI Drafting Assistant Pilot (2‑week observational test)**  

*Prepared for: Director, Support Operations*  
*Length: 624 words*  

---

### 1.  Aggregate & Ticket‑type Resolution Rates  

| Group | Tickets Assigned | Resolved ≤ 24 h | **Resolution Rate** |
|-------|------------------|----------------|----------------------|
| **AI**      | 1,000 | 850 | **85 %** |
| **Control** | 1,000 | 545 | **54.5 %** |

*Fact*: The overall proportion of tickets closed within 24 h is 85 % for the AI‑assigned tickets and 54.5 % for the control tickets, exactly matching the vendor’s headline claim.

Ticket‑type specific rates  

| Group | Simple – Resolved | Simple – Rate | Complex – Resolved | Complex – Rate |
|-------|-------------------|--------------|--------------------|----------------|
| **AI**      | 810 / 900 | **90 %** | 40 / 100 | **40 %** |
| **Control** | 95 / 100  | **95 %** | 450 / 900 | **50 %** |

*Interpretation*: The AI assistant dramatically improves simple‑ticket speed (90 % → 95 % is a modest drop, actually slightly worse) but under‑performs on complex tickets (40 % vs 50 %). Because leads selected which tickets got AI, the mix is heavily skewed toward simple cases (90 % simple in AI vs 10 % simple in control), so the aggregate 85 % rate is not directly comparable across groups.

---

### 2.  Resolution Rates Standardized to a 50/50 Simple‑Complex Workload  

To remove the selection bias we recompute what each group’s resolution rate would be if the workload were an even split (500 simple, 500 complex).  

**AI (hypothetical 50/50)**  
- Simple: 90 % × 500 = 450 resolved  
- Complex: 40 % × 500 = 200 resolved  
- **Standardized total resolved** = 650 / 1,000 → **65 %**  

**Control (hypothetical 50/50)**  
- Simple: 95 % × 500 = 475 resolved  
- Complex: 50 % × 500 = 250 resolved  
- **Standardized total resolved** = 725 / 1,000 → **72.5 %**  

*Interpretation*: When the ticket mix is neutralized, the control process actually yields a higher overall resolution rate (72.5 % vs 65 %). The AI advantage observed in the raw aggregate stems from the disproportionately high share of simple tickets in its sample.

---

### 3.  Durable Resolutions & Cost per Durable Resolution  

*Durable resolution* = Resolved ≤ 24 h **minus** Reopened (reopen within 7 days).  

| Group | Resolved | Reopened | **Durable** | Handling Cost (USD) | **Cost / Durable** |
|-------|----------|----------|------------|----------------------|--------------------|
| **AI**      | 850 | 93 (81+12) | **757** | 2,700 + 1,200 = **3,900** | **$5.15** |
| **Control** | 545 | 49 (4+45) | **496** | 400 + 14,400 = **14,800** | **$29.84** |

*Fact*: The AI group achieves durable resolutions at roughly **$5 per ticket**, whereas the control group spends almost **$30 per durable ticket**.  

*Interpretation*: The lower cost for AI is driven largely by the much smaller handling cost associated with simple tickets (most of AI’s workload). Because the AI sample contains 90 % simple tickets, the cost advantage cannot be assumed to hold if AI were applied to a typical 50/50 mix.  

*Unknown*: We do not have separate cost data for simple vs complex tickets under AI; the aggregate cost may hide higher per‑ticket expenses for complex cases if AI were used more broadly.

---

### 4.  Roll‑out Recommendation  

**Do not roll out AI to all tickets immediately.**  

- **Evidence**: The apparent 85 % aggregate resolution rate is a by‑product of selection bias (90 % simple tickets). When workload is standardized, AI under‑performs (65 % vs 72.5 %).  
- **Cost**: AI shows a striking cost advantage, but this is tied to the simple‑ticket dominance in its pilot sample. Deploying AI to a realistic ticket mix could raise the cost per durable resolution substantially.  
- **Risk**: For complex tickets the AI assistant reduces the 24‑h resolution rate from 50 % to 40 % and may increase reopenings (12 % reopen for AI complex vs 5 % for control). Premature full‑scale rollout could degrade service quality for higher‑value or high‑impact issues.  

**Interim step**: Expand AI use only to the *simple* ticket stream (e.g., Tier‑1 inbound queries) while maintaining the current process for complex tickets. This leverages the observed efficiency gains without exposing complex work to a still‑underperforming tool.

---

### 5.  Follow‑up Experiment – Randomized Controlled Trial (RCT)

| Element | Design |
|---------|--------|
| **Population** | All incoming tickets over a 4‑week period. |
| **Stratification** | Separate tickets into *Simple* and *Complex* using the same definitions as the pilot. |
| **Randomization** | Within each stratum, randomly assign tickets 1:1 to **AI‑assist** vs **Standard**. This guarantees a 50/50 mix overall and preserves balance within ticket type. |
| **Primary Metric** | **Durable resolution rate** = (Resolved ≤24 h – Reopened) / Assigned, computed separately for each stratum and pooled weighted by actual workload. |
| **Guardrails** | • No increase >5 % in overall average handling time.<br>• Reopen rate for complex tickets must not exceed the current baseline (45 / 900 ≈ 5 %). |
| **Allocation** | Expect ~2,000 tickets (based on current volume). With 1,000 per arm, the sample is large enough to detect a 5‑percentage‑point difference in durable rate with reasonable power, but formal power calculations are omitted per instruction. |
| **Stopping Rule** | If at any interim (mid‑point) analysis the durable rate for AI in the *complex* stratum falls below the control rate by >5 % OR the reopen rate spikes above the guardrail, the trial halts and AI is withdrawn for complex tickets. |
| **Analysis** | Compute durable rates and cost per durable resolution for each arm & stratum. Compare using simple proportion differences; adjust overall estimate by the observed real‑world ticket mix. |

*Interpretation*: Randomization eliminates the lead‑selection bias that clouded the pilot, allowing a clean estimate of AI’s causal impact on both speed and quality across ticket types. The durable‑resolution metric directly reflects the business goal of closing tickets quickly *and* keeping them closed.

---

### 6.  Summary of Facts, Interpretations, and Unknowns  

| Category | Content |
|----------|---------|
| **Facts** | Raw counts of assigned, resolved, reopened tickets and handling costs (Table 1). Calculated aggregate resolution rates (85 % vs 54.5 %). Standardized rates under a 50/50 mix (AI 65 %, Control 72.5 %). Durable‑resolution counts and cost per durable ticket ($5.15 vs $29.84). |
| **Interpretations** | The AI’s apparent superiority is driven by an over‑representation of simple tickets. When mix‑adjusted, AI performs worse overall and especially on complex tickets. Cost advantage likely disappears under a realistic ticket mix. |
| **Unknowns** | Exact per‑ticket cost breakdown by complexity for AI; potential learning curve effects if agents use AI more frequently; long‑term impact on customer satisfaction; any hidden operational constraints (e.g., AI licensing limits). |

**Bottom line** – The pilot data do not support a blanket rollout. Deploy AI selectively on simple tickets while commissioning a properly randomized trial to obtain unbiased estimates of its value across the full ticket spectrum. This approach safeguards service quality, manages cost expectations, and provides the evidence needed for any future organization‑wide deployment.
GPT-6 Luna

Requested model: openai/gpt-6-luna · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0007002000 USD. Excludes retry and judging costs.

## Decision brief

**Recommendation:** Do not roll out the assistant to every ticket yet. The pilot is observational: team leads selected which tickets used AI, and the groups had very different ticket mixes. The results are descriptive, not evidence that AI caused the differences. Run a randomized follow-up before deciding on broad deployment.

### Pilot facts

- **Aggregate resolved within 24 hours:** AI: 850/1,000 = **85.0%**. Control: 545/1,000 = **54.5%**. The vendor’s arithmetic is correct, but the aggregate comparison is heavily affected by ticket mix.
- **Within ticket type:**
  - Simple: AI 810/900 = **90%**; control 95/100 = **95%**.
  - Complex: AI 40/100 = **40%**; control 450/900 = **50%**.
- **Standardized to a 50/50 simple/complex workload:**
  - AI: (90% + 40%)/2 = **65.0%**.
  - Control: (95% + 50%)/2 = **72.5%**.
  
  Thus, under this standardization, AI is 7.5 percentage points lower. This is a descriptive comparison, not a causal estimate.

- **Durable resolutions** (resolved minus reopened):
  - AI: 850 − 93 = **757** (75.7% of assigned tickets).
  - Control: 545 − 49 = **496** (49.6% of assigned tickets).
- **Cost per durable resolution**, using cost for *all assigned tickets*:
  - AI: ($2,700 + $1,200)/757 = **$5.15**.
  - Control: ($400 + $14,400)/496 = **$29.84**.

**Interpretation and limitation:** AI tickets were 90% simple, versus 10% of control tickets. That imbalance can explain much of the aggregate resolution and cost-per-durable difference. The cost figures are not adjusted for ticket type, so they do not establish that AI reduces cost for comparable work. The within-type results are also observational and may reflect other differences in ticket assignment or handling.

**Unknowns:** The pilot does not establish causal effects, statistical uncertainty, performance beyond this two-week period, or effects on customer experience and other quality outcomes.

### Randomized follow-up

Randomly assign eligible tickets **1:1 to AI or control within simple and complex ticket types**, using the same workflow and concurrent period. Analyze by assigned group, regardless of whether the assistant is used.

- **Primary metric:** durable resolutions per assigned ticket: resolved within 24 hours and not reopened within seven days.
- **Guardrails:** 24-hour resolution rate, seven-day reopen rate, handling cost per assigned ticket, and a predefined quality/customer-experience measure. Set safety and quality pause thresholds before launch.
- **Stopping rule:** Enroll for a prespecified four weeks, then wait until every enrolled ticket has seven days of follow-up before analysis. Do not stop early for apparent efficacy; pause enrollment earlier only for a predefined serious safety or quality concern.

This follow-up would provide a fairer comparison by ticket type; the current pilot alone does not justify immediate universal rollout.
Mercury 2.5

Requested model: inception/mercury-2.5 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:24 UTC

Reported answer cost: $0.0001935700 USD. Excludes retry and judging costs.


Decision Brief: AI Drafting Assistant Pilot

**1. Resolution Rates (Facts)**
Aggregate resolution (resolved/assigned):
- AI: (810 + 40) / (900 + 100) = 850/1000 = 85%
- Control: (95 + 450) / (100 + 900) = 545/1000 = 54.5%

Within-type resolution:
- AI Simple: 810/900 = 90%
- AI Complex: 40/100 = 40%
- Control Simple: 95/100 = 95%
- Control Complex: 450/900 = 50%

**2. Standardized Rates (Interpretation)**
Assuming a 50/50 simple/complex workload mix:
- AI: (90% + 40%) / 2 = 65%
- Control: (95% + 50%) / 2 = 72.5%

The vendor’s 85% figure relies on AI handling mostly simple tickets (90% of assignment), while Control handled mostly complex (90% of assignment). This selection bias inflates AI’s aggregate rate.

**3. Durable Resolutions and Cost (Facts & Limitations)**
Durable = Resolved − Reopened.
- AI Simple: 810 − 81 = 729; Cost $2,700 → $3.70/durable
- AI Complex: 40 − 12 = 28; Cost $1,200 → $42.86/durable
- Control Simple: 95 − 4 = 91; Cost $400 → $4.40/durable
- Control Complex: 450 − 45 = 405; Cost $14,400 → $35.56/durable

Aggregate durable counts and cost/durable:
- AI: (729+28) = 757; Total Cost $3,900 → $5.15/durable
- Control: (91+405) = 496; Total Cost $14,800 → $29.84/durable

Limitation: Costs include all assigned tickets, not just resolved ones. The non-randomized assignment means we cannot isolate AI’s causal impact on cost or durability. Control’s higher complex volume drives its aggregate cost per durable resolution.

**4. Rollout Recommendation (Interpretation)**
Do not roll out to every ticket immediately. While AI shows lower cost per durable resolution overall, its complex-ticket resolution (40%) is lower than control (50%). Blind rollout risks degrading complex-ticket outcomes and increasing reopens. Given the selection bias—leads assigned AI to simpler tickets—the pilot does not prove AI generalizes to complex work.

Recommendation: Restrict AI to simple tickets pending further testing. For complex tickets, maintain current workflows. This minimizes risk while leveraging AI’s efficiency on low-complexity cases.

**5. Follow-Up Experiment (Interpretation & Unknowns)**
Design a randomized controlled trial (RCT) over four weeks.
- **Primary Metric:** Durable resolution rate (resolved − reopened within 7 days / assigned).
- **Guardrails:** Monitor reopens, handle rate, and customer satisfaction. Stop if reopens exceed 10% baseline or handle time increases by 15%.
- **Allocation:** Randomize tickets 50/50 to AI vs. control, stratified by ticket type (simple/complex) to ensure balanced mix.
- **Stopping Rule:** End early if interim analysis shows >20% drop in durable resolution for AI on complex tickets or >10% cost increase per durable resolution.

Unknowns: Long-term effects on agent skills, customer perception of AI-drafted responses, and scalability beyond this team. Staff-level productivity data was not captured; future studies should include agent feedback.

**Summary**
Facts show AI performed well on simple tickets but poorly on complex ones compared to control. The vendor’s claim ignores ticket-mix bias. Standardized rates favor control. Cost analysis suggests AI is cheaper overall, but this is confounded by assignment patterns. Proceed cautiously: limit AI to simple tickets and validate with an RCT before broader deployment.
Ling 3.0 Flash

Requested model: inclusionai/ling-3.0-flash · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0002586570 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot**

**Facts**
- **Aggregate resolution rates:** AI 85% (850 resolved / 1,000 assigned); Control 54.5% (545 / 1,000).
- **Within-type resolution rates:**
  - Simple: AI 90% (810 / 900); Control 95% (95 / 100).
  - Complex: AI 40% (40 / 100); Control 50% (450 / 900).
- **Standardized to a 50/50 simple/complex workload** (average of type-specific rates):
  - AI: (90% + 40%) / 2 = 65%.
  - Control: (95% + 50%) / 2 = 72.5%.
- **Durable resolutions** (resolved minus reopened):
  - AI: (810 – 81) + (40 – 12) = 757.
  - Control: (95 – 4) + (450 – 45) = 496.
- **Cost per durable resolution** (total handling cost / durable resolutions):
  - AI: $3,900 / 757 ≈ $5.15.
  - Control: $14,800 / 496 ≈ $29.84.
  - *Limitation:* AI handled 90% simple tickets (cheaper, higher durable rate); this mix inflates AI’s apparent cost efficiency.
- **Standardized cost per durable** (hypothetical 500 simple + 500 complex per group):
  - AI cost: 500 × $3 (simple) + 500 × $12 (complex) = $7,500. Durable: 500 × 0.81 + 500 × 0.28 = 545. Cost/durable = $13.76.
  - Control cost: 500 × $4 + 500 × $16 = $10,000. Durable: 500 × 0.91 + 500 × 0.45 = 680. Cost/durable = $14.71.

**Interpretations**
The vendor’s aggregate comparison (54.5% → 85%) reflects a ticket-mix artifact: AI was assigned predominantly to simple tickets. When mix is controlled, AI underperforms control on both simple (90% vs. 95%) and complex (40% vs. 50%) tickets. AI’s lower cost per durable resolution is driven by handling cheaper simple tickets; after standardizing workload, the cost difference narrows and reverses slightly. The data do not support the claim that AI improves resolution.

**Unknowns**
- Why AI drafts lead to lower resolution rates on both ticket types (e.g., draft quality, staff behavior).
- Whether the observed patterns hold in a randomized setting with balanced mix.
- Long-term impacts on customer satisfaction or staff workload.

**Recommendation**
**Do not roll out AI to all tickets.** The pilot shows AI reduces resolution rates when ticket mix is equalized, and its cost advantage is not robust. A randomized follow-up is needed before any broader deployment.

**Randomized Follow-up Experiment**
- **Primary metric:** Durable resolution rate (durable resolutions per assigned ticket).
- **Guardrails:** Monitor cost per durable resolution, median resolution time, and escalation rate. Set a safety threshold: if AI’s durable resolution rate falls >5 percentage points below control at any interim check, pause and review.
- **Allocation:** Randomly assign each incoming ticket to AI or Control with equal probability, stratified by ticket type (simple/complex) to ensure balanced mix.
- **Stopping rule:** Run for a fixed 4-week period. If the safety threshold is breached, stop early; otherwise, complete the full period and compare the primary metric.
Solar Pro 4

Requested model: upstage/solar-pro4 · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:30 UTC

Reported answer cost: $0.0017974800 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot Evaluation**

**1. Resolution Rates (Facts)**
*Aggregate:*
*   **AI:** 850 resolved / 1,000 assigned = **85.0%**
*   **Control:** 545 resolved / 1,000 assigned = **54.5%**
*Vendor claim matches aggregate facts.*

*Within Ticket Type:*
*   **Simple:** AI 810/900 = **90.0%**; Control 95/100 = **95.0%**
*   **Complex:** AI 40/100 = **40.0%**; Control 450/900 = **50.0%**
*Interpretation:* Within both types, AI resolution rates are *lower* than Control. The aggregate AI advantage is driven by ticket mix, not performance.

**2. Standardized Rates (50/50 Workload)**
Assuming an equal mix of 500 Simple and 500 Complex per group:
*   **AI:** (500 × 0.90) + (500 × 0.40) = 650 resolved → **65.0% rate**
*   **Control:** (500 × 0.95) + (500 × 0.50) = 725 resolved → **72.5% rate**
*Interpretation:* Under a balanced workload, Control outperforms AI by 7.5 percentage points. The vendor's aggregate comparison ignores selection bias (leads assigned mostly Simple tickets to AI, mostly Complex to Control).

**3. Durable Resolutions and Cost**
Durable = Resolved − Reopened.
*   **AI Durable:** 850 − 93 = 757. Total Cost $3,900. Cost per Durable: $3,900 / 757 ≈ **$5.15**
*   **Control Durable:** 545 − 49 = 496. Total Cost $14,800. Cost per Durable: $14,800 / 496 ≈ **$29.84**
*Limitations:* AI Group handled 90% Simple tickets (lower cost, higher durability). Control Group handled 90% Complex tickets (higher cost, lower durability). While AI unit costs appear lower (Simple: $3 vs $4; Complex: $12 vs $16), the cost-per-durable gap is substantially confounded by ticket mix. We cannot isolate AI efficiency from ticket complexity with this data.

**4. Rollout Recommendation**
**Do not roll out immediately.**
*Fact:* Aggregate resolution rose only because AI received easier tickets.
*Interpretation:* Within-type rates suggest AI may hinder resolution (90% vs 95% Simple; 40% vs 50% Complex).
*Unknown:* Whether AI assists or distracts agents on specific ticket types.
*Action:* Hold rollout. Investigate why agents performed worse on AI tickets despite lower unit costs.

**5. Randomized Follow-Up Experiment**
*Design:* Stratified Randomized Controlled Trial.
*Allocation:* Randomize tickets 50/50 (AI vs Control) *within* Simple and Complex strata to ensure balanced mix.
*Primary Metric:* **Durable Resolution Rate** (Durable Resolutions / Assigned Tickets). This penalizes quick closures that reopen.
*Guardrails:* Monitor Reopen Rate (<5% target) and Average Handling Cost per Ticket (cap at Control baseline).
*Stopping Rule:* Fixed horizon (e.g., 4 weeks) to prevent peeking bias. No early stopping based on intermediate aggregates.
Gemini 3.1 Pro Preview

Requested model: google/gemini-3.1-pro-preview · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:25 UTC

Reported answer cost: $0.0263180000 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot**

**1. Resolution Rates (Facts & Calculations)**
*   **Aggregate Rates:** 
    *   AI: 85% (850 resolved / 1,000 assigned)
    *   Control: 54.5% (545 resolved / 1,000 assigned)
*   **Within-Ticket-Type Rates:**
    *   AI Simple: 90% (810 / 900)
    *   Control Simple: 95% (95 / 100)
    *   AI Complex: 40% (40 / 100)
    *   Control Complex: 50% (450 / 900)

*Interpretation:* The vendor’s claim is factually accurate in aggregate but highly misleading. Team leads assigned 90% of simple tickets to AI and 90% of complex tickets to the Control group. When looking at individual ticket types, the Control group actually outperformed AI in resolving both Simple (95% vs 90%) and Complex (50% vs 40%) tickets. 

**2. Standardized Resolution Rates**
If we standardize the workload to a 50/50 split of simple and complex tickets for both groups (calculating the unweighted average of the two resolution rates):
*   **AI Standardized:** 65% ((90% + 40%) / 2)
*   **Control Standardized:** 72.5% ((95% + 50%) / 2)

*Interpretation:* Adjusted for ticket mix, the Control group outperforms the AI assistant. 

**3. Durable Resolutions & Costs (Facts & Limitations)**
Durable resolutions exclude tickets reopened within seven days (Resolved - Reopened).
*   **AI Durable Resolutions:** 757 (850 total resolved - 93 reopened)
*   **Control Durable Resolutions:** 496 (545 total resolved - 49 reopened)
*   **AI Cost per Durable Resolution:** $5.15 ($3,900 total cost / 757)
*   **Control Cost per Durable Resolution:** $29.84 ($14,800 total cost / 496)

*Limitation:* While AI appears drastically cheaper per durable resolution, this is an artifact of the ticket mix. Simple tickets inherently cost less to handle than complex ones. Because AI was assigned overwhelmingly simple tickets, its aggregate costs are artificially low compared to the Control group's complex-heavy workload.

**4. Rollout Recommendation**
**Recommendation:** Do *not* roll out the AI assistant immediately. 
*Justification:* The pilot data suffers from severe selection bias. The vendor's 85% success rate is an illusion caused by cherry-picking easy tickets. The standardized rates show that human agents without AI (Control) resolved a higher percentage of tickets. Furthermore, we do not know if the AI is actively harming quality, as AI reopen rates appear higher within ticket categories (e.g., 30% of resolved AI complex tickets reopened, compared to 10% for Control).

**5. Follow-Up Experiment**
To confidently assess the AI's value, we must run a controlled, randomized experiment. 

*   **Design:** A two-week A/B test.
*   **Allocation:** 50% of tickets to AI and 50% to Control. Crucially, randomization must be automated at the ticketing-system level, stratified by ticket type (Simple/Complex). Team leads cannot choose the assignments.
*   **Primary Metric:** Standardized durable resolution rate.
*   **Guardrails:** Average handling cost per ticket and reopen rates (to ensure AI isn't sacrificing quality for speed).
*   **Stopping Rule:** Halt the experiment early if the AI group's aggregate reopen rate exceeds 20%, indicating a severe degradation in customer experience. 

*Unknowns to address:* The true causal effect of AI on handling times, costs, and quality when applied uniformly. The follow-up test will isolate these variables.
Mistral Medium 3.5

Requested model: mistralai/mistral-medium-3-5 · Output budget: 16384 tokens · Requested reasoning: none · Sep 28, 2026, 17:56 UTC

Reported answer cost: $0.0064410000 USD. Excludes retry and judging costs.

**Decision Brief: AI Drafting Assistant Pilot Evaluation**

**Facts**
- **Aggregate resolution rates**: AI resolved 850/1000 (85%); Control resolved 545/1000 (54.5%). The vendor’s claim is arithmetically correct but ignores ticket-type mix and reopen rates.
- **Within-type resolution rates**:
  - *Simple*: AI 810/900 (90%); Control 95/100 (95%).
  - *Complex*: AI 40/100 (40%); Control 450/900 (50%).
- **Standardized to 50/50 simple/complex workload**:
  - AI: (90% + 40%)/2 = **65%**.
  - Control: (95% + 50%)/2 = **72.5%**.
  *Interpretation*: Under equal ticket-type distribution, AI underperforms Control by 7.5 percentage points.

- **Durable resolutions** (resolved − reopened) and **cost per durable resolution**:
  - *AI Simple*: 810 − 81 = 729; cost = $2700 → **$3.70/durable**.
  - *AI Complex*: 40 − 12 = 28; cost = $1200 → **$42.86/durable**.
  - *Control Simple*: 95 − 4 = 91; cost = $400 → **$4.40/durable**.
  - *Control Complex*: 450 − 45 = 405; cost = $14400 → **$35.56/durable**.
  *Note*: AI’s cost efficiency varies sharply by type; Simple tickets are cheaper to resolve durably with AI, but Complex tickets are more expensive.

**Unknowns**
- Causal effect of AI (selection bias: leads may have assigned AI to easier Complex tickets).
- Impact on quality beyond reopen rates (e.g., customer satisfaction).
- Long-term cost implications (e.g., training, vendor fees).

**Recommendation**
**Do not roll out universally.** The vendor’s aggregate claim masks underperformance on Complex tickets (the majority of Control’s workload) and higher cost per durable resolution for Complex. AI excels on Simple tickets (higher resolution, lower cost/durable), but the pilot’s ticket mix (90% Simple for AI vs. 10% for Control) suggests selection bias. A targeted rollout to **Simple tickets only** is justified by the data, but Complex tickets require further evidence.

**Follow-up Experiment**
- **Design**: Randomized controlled trial (RCT) with **50/50 AI/Control allocation** *within each ticket type* (Simple/Complex) for 4 weeks.
- **Primary metric**: *Durable resolution rate* (resolved − reopened) / assigned.
- **Guardrails**:
  - Stop if AI’s Complex durable resolution rate drops >10% below Control (interim check at 2 weeks).
  - Cap AI assignment to 20% of Complex tickets until results stabilize.
- **Allocation**: Stratified by type to ensure balance; block randomization by day to account for temporal trends.
- **Stopping rule**: Early termination if durable resolution rate for AI Complex is statistically inferior (directional check only; no p-values claimed).

**Justification**: RCT addresses selection bias, isolates type-specific effects, and prioritizes durable outcomes over raw resolution. Cost per durable resolution should be monitored as a secondary metric.
gpt-oss-20b (Nitro)

Requested model: openai/gpt-oss-20b:nitro · Output budget: 8192 tokens · Requested reasoning: low · Sep 28, 2026, 17:29 UTC

Reported answer cost: $0.0002978100 USD. Excludes retry and judging costs.

**Decision Brief – AI Drafting Assistant (2‑Week Pilot)**  
*Prepared for the Support Operations Lead – 28 Sep 2026*  

---

### 1. Resolution Rates

| Group  | Ticket Type | Assigned | Resolved within 24 h | Resolution % |
|--------|-------------|----------|----------------------|--------------|
| **AI** | Simple | 900 | 810 | 90.0 % |
|        | Complex | 100 | 40 | 40.0 % |
| **Control** | Simple | 100 | 95 | 95.0 % |
|           | Complex | 900 | 450 | 50.0 % |

*Aggregate (unweighted) resolution = (810 + 40 + 95 + 450) ÷ (900 + 100 + 100 + 900) = 1,395 ÷ 2,000 = 69.75 %.*

> **Fact:** The vendor’s claim (54.5 % → 85 %) applies only to the *subset of tickets* where the AI was selected; the pilot’s aggregate shows 69.8 % overall and 40 % for complex tickets under AI.

---

### 2. Standardised 50/50 Simple‑Complex Load

The control arm had a heavily skewed mix (90 % complex). To compare apples‑to‑apples we re‑weight each group to a 1:1 simple‑complex distribution.

For each group:  
*Weighted simple tickets = Assigned × 0.5*  
*Weighted complex tickets = Assigned × 0.5*  

| Group | Weighted Simple | Weighted Complex | Weighted Resolved Simple | Weighted Resolved Complex | Standardised Resolved % |
|-------|-----------------|------------------|--------------------------|---------------------------|------------------------|
| **AI** | 900 × 0.5 = 450 | 100 × 0.5 = 50 | 810 × 0.5 = 405 | 40 × 0.5 = 20 | (405 + 20) ÷ (450 + 50) = 425 ÷ 500 = 85.0 % |
| **Control** | 100 × 0.5 = 50 | 900 × 0.5 = 450 | 95 × 0.5 = 47.5 | 450 × 0.5 = 225 | (47.5 + 225) ÷ (50 + 450) = 272.5 ÷ 500 = 54.5 % |

> **Interpretation:** Under an equal simple/complex mix the AI arm outperforms the control (85 % vs 54.5 %) – the vendor’s headline figures are reproduced.  
> **Unknown:** The quality of the AI’s output on complex tickets in this low‑volume sample (only 100 tickets) may not generalise.

---

### 3. Durable Resolutions & Cost per Durable Resolution

*Durable resolution = Resolved – Reopened.*

| Group | Ticket Type | Resolved | Reopened | Durable | Handling Cost (USD) | Cost per Durable (USD) |
|-------|-------------|----------|----------|---------|---------------------|------------------------|
| **AI** | Simple | 810 | 81 | 729 | 2,700 | 2,700 ÷ 729 ≈ 3.71 |
|        | Complex | 40 | 12 | 28 | 1,200 | 1,200 ÷ 28 ≈ 42.86 |
| **Control** | Simple | 95 | 4 | 91 | 400 | 400 ÷ 91 ≈ 4.40 |
|           | Complex | 450 | 45 | 405 | 14,400 | 14,400 ÷ 405 ≈ 35.56 |

**Weighted (50/50 mix)**  

*Weighted durable = Weighted Resolved – Weighted Reopened.*

| Group | Weighted Durable | Weighted Handling Cost | Cost per Durable |
|-------|------------------|------------------------|------------------|
| **AI** | (405 + 20) – (40.5 + 6) = 425 – 46.5 = 378.5 | 2,700 + 1,200 = 3,900 | 3,900 ÷ 378.5 ≈ 10.31 |
| **Control** | (47.5 + 225) – (23.75 + 22.5) = 272.5 – 46.25 = 226.25 | 400 + 14,400 = 14,800 | 14,800 ÷ 226.25 ≈ 65.48 |

> **Fact:** With the 50/50 mix the AI costs ≈ $10.3 per durable resolution; control ≈ $65.5.  
> **Interpretation:** AI dramatically improves cost‑efficiency when the workload is balanced.  
> **Unknown:** The reopened‑rate for complex tickets under AI is high (12/40); a higher mix would raise overall cost.

---

### 4. Roll‑Out Recommendation

**Recommendation:**  
*Roll out the AI drafting assistant **only for simple tickets** across the organization; defer or pilot a separate module for complex tickets.*

**Rationale**

1. **Performance Disparity** – AI solves simple tickets with 90 % success, whereas complex ticket resolution drops to 40 %.  
2. **Cost Efficiency** – For simple tickets AI’s cost per durable resolution (~$3.7) is far below control (~$4.4). For complex tickets AI is *more* expensive ($42.9 vs $35.6).  
3. **Risk of Reopened Issues** – 10 % of AI‑resolved simple tickets were reopened—still lower than control’s 4 % but must be monitored.  
4. **Work‑Load Realism** – The current operational mix is about 90 % complex; deploying AI everywhere would magnify the high cost and low effectiveness for complex tickets.  

*Implementation plan:*  
- Provide the AI assistant to all frontend (simple) queue agents.  
- Keep the control or manual process for complex tickets until further testing improves AI’s model or a separate “complex‑AI” module is validated.  
- Track durable resolution rates and cost per durable resolution monthly; if complex‑ticket performance improves to ≥ 70 % resolution, revisit full rollout.

---

### 5. Randomised Follow‑Up Experiment

**Objective** – Estimate the causal effect of AI on *durable resolution rate* for *simple* tickets.

| Element | Description |
|---------|-------------|
| **Population** | All incoming tickets over a 3‑month period. |
| **Unit of Randomisation** | Ticket. Every ticket is independently assigned to AI or Control (manual drafting). |
| **Primary Metric** | % of tickets that are durable‑resolved within 24 h (Resolved – Reopened). |
| **Guardrails** | 1) No ticket may exceed 7 days without a status update. 2) Escalation to a senior agent for any ticket unresolved after 48 h. 3) Reopened tickets logged and revisited within 72 h. |
| **Allocation** | 50 % AI, 50 % Control, stratified by ticket type (simple/complex) to preserve workload balance. |
| **Stopping Rule** | As soon as an interim analysis shows a > 5 pp difference in durable resolution rates between arms with a one‑tailed 95 % confidence (using a simple Z‑test), and the trial has observed at least 500 tickets per arm, the experiment will stop in favour of the superior arm. |

*No sample‑size calculation is provided; the experiment will run for a fixed 3 months to capture a representative batch of tickets.*

> **Fact:** Randomised assignment eliminates selection bias present in the pilot.  
> **Interpretation:** The primary metric directly captures the durability of resolutions, accounting for both initial closure and avoidance of reopening.  
> **Unknown:** The optimal number of tickets required to detect a clinically meaningful improvement; the guardrails assume practice will enforce timely updates.

---

**Word Count:** 639  
---

Head-to-head decisions

A and B are presentation positions, not model labels shown to Jev. Confidence is a 0–1 decision signal, not a correctness guarantee. Jev provides a choice and probabilities, not a written explanation.

Ling 3.0 Flash vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Ling 3.0 Flash
Answer B
Qwen3.7 Flash
Probabilities
A: 0.430 · B: 0.570
Confidence
0.140
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Qwen3.7 Flash
Answer B
GLM 5.3 Prime
Probabilities
A: 0.440 · B: 0.560
Confidence
0.120
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
Qwen3.8 Max Prime
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.460 · B: 0.540
Confidence
0.070
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Claude Fable 5.1
Probabilities
A: 0.700 · B: 0.300
Confidence
0.410
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Claude Opus 5.5
Probabilities
A: 0.730 · B: 0.270
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs MiMo-V2.6-Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.680 · B: 0.320
Confidence
0.360
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Ling 3.0 Flash
Probabilities
A: 0.960 · B: 0.040
Confidence
0.910
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Nemotron 3 Ultra (free) vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.870 · B: 0.130
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Luna vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
GPT-6 Luna
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.140 · B: 0.860
Confidence
0.720
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Solar Pro 4 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
Solar Pro 4
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.060 · B: 0.940
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Grok 4.7 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Grok 4.7
Probabilities
A: 0.610 · B: 0.390
Confidence
0.230
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
DeepSeek V4 Pro vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.950 · B: 0.050
Confidence
0.900
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Mercury 2.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Muse Spark 1.3 Contributor vs MiMo-V2.6-Flash · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.660 · B: 0.340
Confidence
0.320
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
MiniMax M2.7 (Nitro) vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Gemini 3.8 Flash vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Kimi K3 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
Kimi K3
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.420 · B: 0.580
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Astra vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
GPT-6 Astra
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.480 · B: 0.520
Confidence
0.030
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Sol vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
GPT-6 Sol
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
gpt-oss-120b vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
gpt-oss-120b
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.210 · B: 0.790
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
gpt-oss-20b (Nitro) vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.7 Flash vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
Qwen3.7 Flash
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.260 · B: 0.740
Confidence
0.490
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Hy3 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Hy3
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Gemini 3.8 Flash vs Mercury 2.5 · Gemini 3.8 Flash wins
Answer A
Mercury 2.5
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.260 · B: 0.740
Confidence
0.490
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:24 UTC
Gemini 3.1 Pro Preview vs Gemini 3.8 Flash · Gemini 3.8 Flash wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs Mercury 2.5 · Mercury 2.5 wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Mercury 2.5
Probabilities
A: 0.280 · B: 0.720
Confidence
0.430
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs Ling 3.0 Flash · Ling 3.0 Flash wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Ling 3.0 Flash
Probabilities
A: 0.240 · B: 0.760
Confidence
0.510
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs Ling 3.0 Flash · Gemini 3.8 Flash wins
Answer A
Ling 3.0 Flash
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.260 · B: 0.740
Confidence
0.470
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs Ling 3.0 Flash · Mercury 2.5 wins
Answer A
Mercury 2.5
Answer B
Ling 3.0 Flash
Probabilities
A: 0.550 · B: 0.450
Confidence
0.110
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiMo-V2.6-Flash vs GLM 5.3 Flash · MiMo-V2.6-Flash wins
Answer A
GLM 5.3 Flash
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.430 · B: 0.570
Confidence
0.140
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.8 Max Prime vs Gemini 3.1 Pro Preview · Qwen3.8 Max Prime wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.080 · B: 0.920
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs Gemini 3.8 Flash · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.940 · B: 0.060
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs Mercury 2.5 · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Mercury 2.5
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs Ling 3.0 Flash · Qwen3.8 Max Prime wins
Answer A
Ling 3.0 Flash
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.200 · B: 0.800
Confidence
0.600
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs Gemini 3.8 Flash · DeepSeek V4.1 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.250 · B: 0.750
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs MiniMax M2.7 (Nitro) · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.870 · B: 0.130
Confidence
0.740
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs MiniMax M2.7 (Nitro) · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.750 · B: 0.250
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs MiniMax M2.7 (Nitro) · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Mercury 2.5
Probabilities
A: 0.940 · B: 0.060
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs MiniMax M2.7 (Nitro) · MiniMax M2.7 (Nitro) wins
Answer A
Ling 3.0 Flash
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs Gemini 3.1 Pro Preview · Claude Opus 5.5 wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Claude Opus 5.5
Probabilities
A: 0.050 · B: 0.950
Confidence
0.900
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs MiniMax M2.7 (Nitro) · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs Claude Fable 5.1 · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.870 · B: 0.130
Confidence
0.740
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Claude Opus 5.5 · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Claude Opus 5.5
Probabilities
A: 0.540 · B: 0.460
Confidence
0.080
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs DeepSeek V4.1 Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Claude Fable 5.1
Probabilities
A: 0.660 · B: 0.340
Confidence
0.330
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Gemini 3.1 Pro Preview · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 1.000 · B: 0.000
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Gemini 3.8 Flash · Claude Fable 5.1 wins
Answer A
Gemini 3.8 Flash
Answer B
Claude Fable 5.1
Probabilities
A: 0.190 · B: 0.810
Confidence
0.610
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs MiniMax M2.7 (Nitro) · Claude Fable 5.1 wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Claude Fable 5.1
Probabilities
A: 0.250 · B: 0.750
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs Gemini 3.1 Pro Preview · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.990 · B: 0.010
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs Mercury 2.5 · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Mercury 2.5
Probabilities
A: 0.980 · B: 0.020
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs Ling 3.0 Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Ling 3.0 Flash
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs MiniMax M2.7 (Nitro) · DeepSeek V4.1 Flash wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs Claude Opus 5.5 · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Claude Opus 5.5
Probabilities
A: 0.540 · B: 0.460
Confidence
0.080
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs DeepSeek V4.1 Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Claude Opus 5.5
Probabilities
A: 0.730 · B: 0.270
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs Ling 3.0 Flash · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Ling 3.0 Flash
Probabilities
A: 0.970 · B: 0.030
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs Gemini 3.8 Flash · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.950 · B: 0.050
Confidence
0.900
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs Mercury 2.5 · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Mercury 2.5
Probabilities
A: 0.970 · B: 0.030
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs MiniMax M2.7 (Nitro) · Claude Opus 5.5 wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Claude Opus 5.5
Probabilities
A: 0.450 · B: 0.550
Confidence
0.110
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Mercury 2.5 · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Mercury 2.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Ling 3.0 Flash · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Ling 3.0 Flash
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs DeepSeek V4.1 Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.850 · B: 0.150
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs GPT-6 Luna · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.520 · B: 0.480
Confidence
0.030
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs GPT-6 Luna · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
GPT-6 Luna
Probabilities
A: 0.960 · B: 0.040
Confidence
0.910
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs GPT-6 Luna · Claude Opus 5.5 wins
Answer A
GPT-6 Luna
Answer B
Claude Opus 5.5
Probabilities
A: 0.300 · B: 0.700
Confidence
0.390
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs GPT-6 Luna · DeepSeek V4.1 Flash wins
Answer A
GPT-6 Luna
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.170 · B: 0.830
Confidence
0.650
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs GPT-6 Luna · GPT-6 Luna wins
Answer A
Gemini 3.1 Pro Preview
Answer B
GPT-6 Luna
Probabilities
A: 0.080 · B: 0.920
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs GPT-6 Luna · Gemini 3.8 Flash wins
Answer A
GPT-6 Luna
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.330 · B: 0.670
Confidence
0.350
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs GPT-6 Luna · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
Mercury 2.5
Probabilities
A: 0.670 · B: 0.330
Confidence
0.350
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs GPT-6 Luna · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
Ling 3.0 Flash
Probabilities
A: 0.730 · B: 0.270
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiniMax M2.7 (Nitro) vs GPT-6 Luna · MiniMax M2.7 (Nitro) wins
Answer A
GPT-6 Luna
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.380 · B: 0.620
Confidence
0.250
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiMo-V2.6-Flash vs GLM 5.3 Prime · MiMo-V2.6-Flash wins
Answer A
GLM 5.3 Prime
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.430 · B: 0.570
Confidence
0.130
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Space Bunny Alpha vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
Space Bunny Alpha
Answer B
MiMo-V2.6-Flash
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.8 Max Prime vs Muse Spark 1.3 Contributor · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.690 · B: 0.310
Confidence
0.370
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Muse Spark 1.3 Contributor · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.830 · B: 0.170
Confidence
0.650
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Claude Opus 5.5
Probabilities
A: 0.640 · B: 0.360
Confidence
0.280
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.650 · B: 0.350
Confidence
0.300
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.020 · B: 0.980
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.920 · B: 0.080
Confidence
0.850
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Mercury 2.5
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Ling 3.0 Flash
Probabilities
A: 0.930 · B: 0.070
Confidence
0.860
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Muse Spark 1.3 Contributor vs MiniMax M2.7 (Nitro) · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.900 · B: 0.100
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Muse Spark 1.3 Contributor vs GPT-6 Luna · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
GPT-6 Luna
Probabilities
A: 0.920 · B: 0.080
Confidence
0.830
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs GPT-6 Sol · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
GPT-6 Sol
Probabilities
A: 0.820 · B: 0.180
Confidence
0.640
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs GPT-6 Sol · Claude Opus 5.5 wins
Answer A
GPT-6 Sol
Answer B
Claude Opus 5.5
Probabilities
A: 0.410 · B: 0.590
Confidence
0.170
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs GPT-6 Sol · GPT-6 Sol wins
Answer A
Ling 3.0 Flash
Answer B
GPT-6 Sol
Probabilities
A: 0.230 · B: 0.770
Confidence
0.540
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Muse Spark 1.3 Contributor vs GPT-6 Sol · Muse Spark 1.3 Contributor wins
Answer A
GPT-6 Sol
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.300 · B: 0.700
Confidence
0.400
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs gpt-oss-120b · DeepSeek V4.1 Flash wins
Answer A
gpt-oss-120b
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.340 · B: 0.660
Confidence
0.310
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs gpt-oss-120b · gpt-oss-120b wins
Answer A
gpt-oss-120b
Answer B
Ling 3.0 Flash
Probabilities
A: 0.770 · B: 0.230
Confidence
0.550
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs GPT-6 Sol · Claude Fable 5.1 wins
Answer A
GPT-6 Sol
Answer B
Claude Fable 5.1
Probabilities
A: 0.240 · B: 0.760
Confidence
0.530
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs GPT-6 Sol · DeepSeek V4.1 Flash wins
Answer A
GPT-6 Sol
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.220 · B: 0.780
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs GPT-6 Sol · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.960 · B: 0.040
Confidence
0.910
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs GPT-6 Sol · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
Mercury 2.5
Probabilities
A: 0.900 · B: 0.100
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiniMax M2.7 (Nitro) vs GPT-6 Sol · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.590 · B: 0.410
Confidence
0.190
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs GPT-6 Sol · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
GPT-6 Sol
Probabilities
A: 0.540 · B: 0.460
Confidence
0.090
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Luna vs GPT-6 Sol · GPT-6 Sol wins
Answer A
GPT-6 Luna
Answer B
GPT-6 Sol
Probabilities
A: 0.180 · B: 0.820
Confidence
0.650
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs gpt-oss-120b · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
gpt-oss-120b
Probabilities
A: 0.880 · B: 0.120
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs gpt-oss-120b · Claude Fable 5.1 wins
Answer A
gpt-oss-120b
Answer B
Claude Fable 5.1
Probabilities
A: 0.250 · B: 0.750
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs gpt-oss-120b · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
gpt-oss-120b
Probabilities
A: 0.880 · B: 0.120
Confidence
0.770
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs gpt-oss-120b · gpt-oss-120b wins
Answer A
Gemini 3.1 Pro Preview
Answer B
gpt-oss-120b
Probabilities
A: 0.200 · B: 0.800
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs gpt-oss-120b · Gemini 3.8 Flash wins
Answer A
gpt-oss-120b
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.450 · B: 0.550
Confidence
0.100
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs gpt-oss-120b · gpt-oss-120b wins
Answer A
Mercury 2.5
Answer B
gpt-oss-120b
Probabilities
A: 0.490 · B: 0.510
Confidence
0.010
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Muse Spark 1.3 Contributor vs gpt-oss-120b · Muse Spark 1.3 Contributor wins
Answer A
gpt-oss-120b
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiniMax M2.7 (Nitro) vs gpt-oss-120b · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
gpt-oss-120b
Probabilities
A: 0.780 · B: 0.220
Confidence
0.550
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Luna vs gpt-oss-120b · gpt-oss-120b wins
Answer A
gpt-oss-120b
Answer B
GPT-6 Luna
Probabilities
A: 0.590 · B: 0.410
Confidence
0.170
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Sol vs gpt-oss-120b · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
gpt-oss-120b
Probabilities
A: 0.780 · B: 0.220
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs GPT-6 Astra · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
GPT-6 Astra
Probabilities
A: 0.830 · B: 0.170
Confidence
0.660
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.1 Pro Preview vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Opus 5.5 vs GPT-6 Astra · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
GPT-6 Astra
Probabilities
A: 0.810 · B: 0.190
Confidence
0.610
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
DeepSeek V4.1 Flash vs GPT-6 Astra · DeepSeek V4.1 Flash wins
Answer A
GPT-6 Astra
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.460 · B: 0.540
Confidence
0.080
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Ling 3.0 Flash vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Ling 3.0 Flash
Probabilities
A: 0.870 · B: 0.130
Confidence
0.730
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Gemini 3.8 Flash vs GPT-6 Astra · GPT-6 Astra wins
Answer A
Gemini 3.8 Flash
Answer B
GPT-6 Astra
Probabilities
A: 0.430 · B: 0.570
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Mercury 2.5 vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Mercury 2.5
Probabilities
A: 0.970 · B: 0.030
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Astra vs GPT-6 Luna · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
GPT-6 Luna
Probabilities
A: 0.810 · B: 0.190
Confidence
0.610
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Muse Spark 1.3 Contributor vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.520 · B: 0.480
Confidence
0.040
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
MiniMax M2.7 (Nitro) vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.840 · B: 0.160
Confidence
0.670
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Astra vs gpt-oss-120b · GPT-6 Astra wins
Answer A
gpt-oss-120b
Answer B
GPT-6 Astra
Probabilities
A: 0.350 · B: 0.650
Confidence
0.300
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
GPT-6 Astra vs GPT-6 Sol · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
GPT-6 Sol
Probabilities
A: 0.730 · B: 0.270
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Qwen3.8 Max Prime vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.750 · B: 0.250
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:25 UTC
Claude Fable 5.1 vs Kimi K3 · Claude Fable 5.1 wins
Answer A
Kimi K3
Answer B
Claude Fable 5.1
Probabilities
A: 0.420 · B: 0.580
Confidence
0.170
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
DeepSeek V4.1 Flash vs Kimi K3 · DeepSeek V4.1 Flash wins
Answer A
Kimi K3
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.460 · B: 0.540
Confidence
0.080
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
MiniMax M2.7 (Nitro) vs Kimi K3 · Kimi K3 wins
Answer A
Kimi K3
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.750 · B: 0.250
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Kimi K3 vs gpt-oss-120b · Kimi K3 wins
Answer A
gpt-oss-120b
Answer B
Kimi K3
Probabilities
A: 0.420 · B: 0.580
Confidence
0.160
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Qwen3.8 Max Prime vs Kimi K3 · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Kimi K3
Probabilities
A: 0.690 · B: 0.310
Confidence
0.370
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Gemini 3.1 Pro Preview vs Kimi K3 · Kimi K3 wins
Answer A
Kimi K3
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Claude Opus 5.5 vs Kimi K3 · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Kimi K3
Probabilities
A: 0.840 · B: 0.160
Confidence
0.670
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Gemini 3.8 Flash vs Kimi K3 · Kimi K3 wins
Answer A
Gemini 3.8 Flash
Answer B
Kimi K3
Probabilities
A: 0.460 · B: 0.540
Confidence
0.070
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Mercury 2.5 vs Kimi K3 · Kimi K3 wins
Answer A
Mercury 2.5
Answer B
Kimi K3
Probabilities
A: 0.150 · B: 0.850
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Ling 3.0 Flash vs Kimi K3 · Kimi K3 wins
Answer A
Kimi K3
Answer B
Ling 3.0 Flash
Probabilities
A: 0.850 · B: 0.150
Confidence
0.700
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Muse Spark 1.3 Contributor vs Kimi K3 · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Kimi K3
Probabilities
A: 0.730 · B: 0.270
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Kimi K3 vs GPT-6 Astra · Kimi K3 wins
Answer A
Kimi K3
Answer B
GPT-6 Astra
Probabilities
A: 0.500 · B: 0.500
Confidence
0.010
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Kimi K3 vs GPT-6 Luna · Kimi K3 wins
Answer A
Kimi K3
Answer B
GPT-6 Luna
Probabilities
A: 0.870 · B: 0.130
Confidence
0.730
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Kimi K3 vs GPT-6 Sol · Kimi K3 wins
Answer A
GPT-6 Sol
Answer B
Kimi K3
Probabilities
A: 0.330 · B: 0.670
Confidence
0.340
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:26 UTC
Qwen3.8 Max Prime vs Nemotron 3 Ultra (free) · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.710 · B: 0.290
Confidence
0.420
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Claude Fable 5.1 vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Claude Fable 5.1
Probabilities
A: 0.620 · B: 0.380
Confidence
0.240
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
DeepSeek V4.1 Flash vs Nemotron 3 Ultra (free) · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.780 · B: 0.220
Confidence
0.560
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Gemini 3.1 Pro Preview vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.040 · B: 0.960
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Gemini 3.8 Flash vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.930 · B: 0.070
Confidence
0.850
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Ling 3.0 Flash vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Ling 3.0 Flash
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.110 · B: 0.890
Confidence
0.770
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Nemotron 3 Ultra (free) vs gpt-oss-120b · Nemotron 3 Ultra (free) wins
Answer A
gpt-oss-120b
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.370 · B: 0.630
Confidence
0.260
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Claude Opus 5.5 vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Claude Opus 5.5
Probabilities
A: 0.630 · B: 0.370
Confidence
0.260
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Mercury 2.5 vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Mercury 2.5
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.140 · B: 0.860
Confidence
0.730
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Kimi K3 vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Kimi K3
Probabilities
A: 0.780 · B: 0.220
Confidence
0.560
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Muse Spark 1.3 Contributor vs Nemotron 3 Ultra (free) · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.740 · B: 0.260
Confidence
0.480
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Nemotron 3 Ultra (free) vs GPT-6 Astra · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
GPT-6 Astra
Probabilities
A: 0.570 · B: 0.430
Confidence
0.140
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
MiniMax M2.7 (Nitro) vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.880 · B: 0.120
Confidence
0.770
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Nemotron 3 Ultra (free) vs GPT-6 Luna · Nemotron 3 Ultra (free) wins
Answer A
GPT-6 Luna
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Nemotron 3 Ultra (free) vs GPT-6 Sol · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
GPT-6 Sol
Probabilities
A: 0.870 · B: 0.130
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:28 UTC
Qwen3.8 Max Prime vs DeepSeek V4 Pro · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.520 · B: 0.480
Confidence
0.040
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Fable 5.1 vs DeepSeek V4 Pro · Claude Fable 5.1 wins
Answer A
DeepSeek V4 Pro
Answer B
Claude Fable 5.1
Probabilities
A: 0.120 · B: 0.880
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs MiniMax M2.7 (Nitro) · MiniMax M2.7 (Nitro) wins
Answer A
DeepSeek V4 Pro
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.480 · B: 0.520
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Opus 5.5 vs DeepSeek V4 Pro · Claude Opus 5.5 wins
Answer A
DeepSeek V4 Pro
Answer B
Claude Opus 5.5
Probabilities
A: 0.310 · B: 0.690
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4.1 Flash vs DeepSeek V4 Pro · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4 Pro
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.100 · B: 0.900
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Gemini 3.8 Flash · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.750 · B: 0.250
Confidence
0.490
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Ling 3.0 Flash · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Ling 3.0 Flash
Probabilities
A: 0.780 · B: 0.220
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Mercury 2.5 · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Mercury 2.5
Probabilities
A: 0.840 · B: 0.160
Confidence
0.680
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Gemini 3.1 Pro Preview · DeepSeek V4 Pro wins
Answer A
Gemini 3.1 Pro Preview
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Muse Spark 1.3 Contributor · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.960 · B: 0.040
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.910 · B: 0.090
Confidence
0.830
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.900 · B: 0.100
Confidence
0.810
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Kimi K3 · Kimi K3 wins
Answer A
Kimi K3
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.910 · B: 0.090
Confidence
0.810
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs GPT-6 Sol · GPT-6 Sol wins
Answer A
DeepSeek V4 Pro
Answer B
GPT-6 Sol
Probabilities
A: 0.430 · B: 0.570
Confidence
0.140
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs GPT-6 Luna · DeepSeek V4 Pro wins
Answer A
GPT-6 Luna
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.450 · B: 0.550
Confidence
0.090
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs gpt-oss-120b · gpt-oss-120b wins
Answer A
gpt-oss-120b
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.540 · B: 0.460
Confidence
0.080
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs gpt-oss-20b (Nitro) · DeepSeek V4 Pro wins
Answer A
gpt-oss-20b (Nitro)
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.020 · B: 0.980
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Qwen3.8 Max Prime vs gpt-oss-20b (Nitro) · Qwen3.8 Max Prime wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Fable 5.1 vs gpt-oss-20b (Nitro) · Claude Fable 5.1 wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Claude Fable 5.1
Probabilities
A: 0.010 · B: 0.990
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Opus 5.5 vs gpt-oss-20b (Nitro) · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4.1 Flash vs gpt-oss-20b (Nitro) · DeepSeek V4.1 Flash wins
Answer A
gpt-oss-20b (Nitro)
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.010 · B: 0.990
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Gemini 3.1 Pro Preview vs gpt-oss-20b (Nitro) · Gemini 3.1 Pro Preview wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.060 · B: 0.940
Confidence
0.870
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Gemini 3.8 Flash vs gpt-oss-20b (Nitro) · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Mercury 2.5 vs gpt-oss-20b (Nitro) · Mercury 2.5 wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Mercury 2.5
Probabilities
A: 0.040 · B: 0.960
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Ling 3.0 Flash vs gpt-oss-20b (Nitro) · Ling 3.0 Flash wins
Answer A
Ling 3.0 Flash
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Muse Spark 1.3 Contributor vs gpt-oss-20b (Nitro) · Muse Spark 1.3 Contributor wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.010 · B: 0.990
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
MiniMax M2.7 (Nitro) vs gpt-oss-20b (Nitro) · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Kimi K3 vs gpt-oss-20b (Nitro) · Kimi K3 wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Kimi K3
Probabilities
A: 0.010 · B: 0.990
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Nemotron 3 Ultra (free) vs gpt-oss-20b (Nitro) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
GPT-6 Astra vs gpt-oss-20b (Nitro) · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
GPT-6 Luna vs gpt-oss-20b (Nitro) · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
GPT-6 Sol vs gpt-oss-20b (Nitro) · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
gpt-oss-120b vs gpt-oss-20b (Nitro) · gpt-oss-120b wins
Answer A
gpt-oss-120b
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4 Pro vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.680 · B: 0.320
Confidence
0.370
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs Space Bunny Alpha · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Space Bunny Alpha
Probabilities
A: 0.900 · B: 0.100
Confidence
0.790
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Fable 5.1 vs Space Bunny Alpha · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Space Bunny Alpha
Probabilities
A: 0.940 · B: 0.060
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Claude Opus 5.5 vs Space Bunny Alpha · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Space Bunny Alpha
Probabilities
A: 0.940 · B: 0.060
Confidence
0.880
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
DeepSeek V4.1 Flash vs Space Bunny Alpha · DeepSeek V4.1 Flash wins
Answer A
Space Bunny Alpha
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.290 · B: 0.710
Confidence
0.410
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:29 UTC
Gemini 3.1 Pro Preview vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Space Bunny Alpha
Probabilities
A: 0.020 · B: 0.980
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.720 · B: 0.280
Confidence
0.440
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
Mercury 2.5
Probabilities
A: 0.910 · B: 0.090
Confidence
0.820
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Claude Opus 5.5
Probabilities
A: 0.560 · B: 0.440
Confidence
0.130
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs Grok 4.7 · Grok 4.7 wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Grok 4.7
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.910 · B: 0.090
Confidence
0.820
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
GPT-6 Astra
Probabilities
A: 0.840 · B: 0.160
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs Grok 4.7 · Grok 4.7 wins
Answer A
gpt-oss-120b
Answer B
Grok 4.7
Probabilities
A: 0.260 · B: 0.740
Confidence
0.470
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs Space Bunny Alpha · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Space Bunny Alpha
Probabilities
A: 0.900 · B: 0.100
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
Ling 3.0 Flash
Probabilities
A: 0.890 · B: 0.110
Confidence
0.770
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs Space Bunny Alpha · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Space Bunny Alpha
Probabilities
A: 0.670 · B: 0.330
Confidence
0.340
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.570 · B: 0.430
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs Space Bunny Alpha · Kimi K3 wins
Answer A
Kimi K3
Answer B
Space Bunny Alpha
Probabilities
A: 0.720 · B: 0.280
Confidence
0.440
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs Space Bunny Alpha · GPT-6 Astra wins
Answer A
Space Bunny Alpha
Answer B
GPT-6 Astra
Probabilities
A: 0.350 · B: 0.650
Confidence
0.290
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
GPT-6 Sol
Answer B
Space Bunny Alpha
Probabilities
A: 0.420 · B: 0.580
Confidence
0.160
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
GPT-6 Luna
Answer B
Space Bunny Alpha
Probabilities
A: 0.200 · B: 0.800
Confidence
0.610
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Space Bunny Alpha
Answer B
gpt-oss-120b
Probabilities
A: 0.790 · B: 0.210
Confidence
0.580
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Space Bunny Alpha
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Claude Fable 5.1
Probabilities
A: 0.580 · B: 0.420
Confidence
0.170
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.690 · B: 0.310
Confidence
0.390
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs Grok 4.7 · Grok 4.7 wins
Answer A
DeepSeek V4 Pro
Answer B
Grok 4.7
Probabilities
A: 0.140 · B: 0.860
Confidence
0.720
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs Grok 4.7 · Grok 4.7 wins
Answer A
Gemini 3.8 Flash
Answer B
Grok 4.7
Probabilities
A: 0.300 · B: 0.700
Confidence
0.410
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs Grok 4.7 · Grok 4.7 wins
Answer A
Mercury 2.5
Answer B
Grok 4.7
Probabilities
A: 0.060 · B: 0.940
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs Grok 4.7 · Grok 4.7 wins
Answer A
Ling 3.0 Flash
Answer B
Grok 4.7
Probabilities
A: 0.080 · B: 0.920
Confidence
0.830
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.650 · B: 0.350
Confidence
0.300
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs Grok 4.7 · Grok 4.7 wins
Answer A
Kimi K3
Answer B
Grok 4.7
Probabilities
A: 0.420 · B: 0.580
Confidence
0.160
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.880 · B: 0.120
Confidence
0.770
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
GPT-6 Luna
Probabilities
A: 0.970 · B: 0.030
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
GPT-6 Sol
Probabilities
A: 0.900 · B: 0.100
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Space Bunny Alpha vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Space Bunny Alpha
Probabilities
A: 0.970 · B: 0.030
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs GLM 5.3 Flash · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
GLM 5.3 Flash
Probabilities
A: 0.690 · B: 0.310
Confidence
0.370
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs GLM 5.3 Flash · Claude Fable 5.1 wins
Answer A
GLM 5.3 Flash
Answer B
Claude Fable 5.1
Probabilities
A: 0.270 · B: 0.730
Confidence
0.460
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs GLM 5.3 Flash · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
GLM 5.3 Flash
Probabilities
A: 0.810 · B: 0.190
Confidence
0.620
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.580 · B: 0.420
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.790 · B: 0.210
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.690 · B: 0.310
Confidence
0.370
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs GLM 5.3 Flash · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
GLM 5.3 Flash
Probabilities
A: 0.820 · B: 0.180
Confidence
0.650
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Kimi K3
Probabilities
A: 0.510 · B: 0.490
Confidence
0.020
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Space Bunny Alpha vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
Space Bunny Alpha
Answer B
GLM 5.3 Flash
Probabilities
A: 0.470 · B: 0.530
Confidence
0.060
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.660 · B: 0.340
Confidence
0.330
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.760 · B: 0.240
Confidence
0.530
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Kimi K3
Probabilities
A: 0.510 · B: 0.490
Confidence
0.020
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs Hy3 · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
Hy3
Probabilities
A: 0.510 · B: 0.490
Confidence
0.010
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Hy3 vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Hy3
Probabilities
A: 0.900 · B: 0.100
Confidence
0.790
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Mercury 2.5
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
Ling 3.0 Flash
Answer B
GLM 5.3 Flash
Probabilities
A: 0.100 · B: 0.900
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
Gemini 3.1 Pro Preview
Answer B
GLM 5.3 Flash
Probabilities
A: 0.020 · B: 0.980
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
GLM 5.3 Flash
Probabilities
A: 0.460 · B: 0.540
Confidence
0.090
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.630 · B: 0.370
Confidence
0.270
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Ling 3.0 Flash
Answer B
GLM 5.3 Prime
Probabilities
A: 0.120 · B: 0.880
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs GLM 5.3 Prime · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
GLM 5.3 Prime
Probabilities
A: 0.620 · B: 0.380
Confidence
0.250
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs GLM 5.3 Prime · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
GLM 5.3 Prime
Probabilities
A: 0.770 · B: 0.230
Confidence
0.550
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Grok 4.7 vs GLM 5.3 Prime · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
GLM 5.3 Prime
Probabilities
A: 0.840 · B: 0.160
Confidence
0.680
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GLM 5.3 Flash vs GLM 5.3 Prime · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
GLM 5.3 Prime
Probabilities
A: 0.750 · B: 0.250
Confidence
0.490
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GPT-6 Luna
Answer B
GLM 5.3 Flash
Probabilities
A: 0.100 · B: 0.900
Confidence
0.800
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs GLM 5.3 Flash · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
GLM 5.3 Flash
Probabilities
A: 0.650 · B: 0.350
Confidence
0.300
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GPT-6 Sol
Answer B
GLM 5.3 Flash
Probabilities
A: 0.230 · B: 0.770
Confidence
0.540
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
gpt-oss-120b
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
gpt-oss-20b (Nitro)
Answer B
GLM 5.3 Flash
Probabilities
A: 0.010 · B: 0.990
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs Solar Pro 4 · Solar Pro 4 wins
Answer A
Solar Pro 4
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.760 · B: 0.240
Confidence
0.510
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Grok 4.7 vs GLM 5.3 Flash · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
GLM 5.3 Flash
Probabilities
A: 0.920 · B: 0.080
Confidence
0.830
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs Hy3 · DeepSeek V4.1 Flash wins
Answer A
Hy3
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.420 · B: 0.580
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs Hy3 · Claude Opus 5.5 wins
Answer A
Hy3
Answer B
Claude Opus 5.5
Probabilities
A: 0.260 · B: 0.740
Confidence
0.490
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.950 · B: 0.050
Confidence
0.900
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.750 · B: 0.250
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Mercury 2.5
Probabilities
A: 0.920 · B: 0.080
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs Hy3 · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Hy3
Probabilities
A: 0.840 · B: 0.160
Confidence
0.680
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.760 · B: 0.240
Confidence
0.530
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs Hy3 · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Hy3
Probabilities
A: 0.780 · B: 0.220
Confidence
0.560
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs Hy3 · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Hy3
Probabilities
A: 0.800 · B: 0.200
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs Hy3 · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Hy3
Probabilities
A: 0.920 · B: 0.080
Confidence
0.850
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs Hy3 · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
Hy3
Probabilities
A: 0.650 · B: 0.350
Confidence
0.310
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Space Bunny Alpha vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Space Bunny Alpha
Probabilities
A: 0.770 · B: 0.230
Confidence
0.530
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.690 · B: 0.310
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Hy3 vs GLM 5.3 Flash · Hy3 wins
Answer A
Hy3
Answer B
GLM 5.3 Flash
Probabilities
A: 0.570 · B: 0.430
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Ling 3.0 Flash
Probabilities
A: 0.840 · B: 0.160
Confidence
0.670
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs GLM 5.3 Prime · Claude Fable 5.1 wins
Answer A
GLM 5.3 Prime
Answer B
Claude Fable 5.1
Probabilities
A: 0.270 · B: 0.730
Confidence
0.450
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs GLM 5.3 Prime · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
GLM 5.3 Prime
Probabilities
A: 0.840 · B: 0.160
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Gemini 3.1 Pro Preview
Answer B
GLM 5.3 Prime
Probabilities
A: 0.020 · B: 0.980
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs GLM 5.3 Prime · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
GLM 5.3 Prime
Probabilities
A: 0.560 · B: 0.440
Confidence
0.120
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs Hy3 · gpt-oss-120b wins
Answer A
gpt-oss-120b
Answer B
Hy3
Probabilities
A: 0.520 · B: 0.480
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs GLM 5.3 Prime · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
GLM 5.3 Prime
Probabilities
A: 0.840 · B: 0.160
Confidence
0.670
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.560 · B: 0.440
Confidence
0.130
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs GLM 5.3 Prime · Kimi K3 wins
Answer A
GLM 5.3 Prime
Answer B
Kimi K3
Probabilities
A: 0.430 · B: 0.570
Confidence
0.130
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.850 · B: 0.150
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Mercury 2.5
Answer B
GLM 5.3 Prime
Probabilities
A: 0.180 · B: 0.820
Confidence
0.640
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs GLM 5.3 Prime · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
GLM 5.3 Prime
Probabilities
A: 0.630 · B: 0.370
Confidence
0.270
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
gpt-oss-20b (Nitro)
Answer B
GLM 5.3 Prime
Probabilities
A: 0.020 · B: 0.980
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Space Bunny Alpha vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Space Bunny Alpha
Answer B
GLM 5.3 Prime
Probabilities
A: 0.490 · B: 0.510
Confidence
0.020
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Hy3 vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
Hy3
Probabilities
A: 0.800 · B: 0.200
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
GPT-6 Luna
Probabilities
A: 0.770 · B: 0.230
Confidence
0.540
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GPT-6 Sol
Answer B
GLM 5.3 Prime
Probabilities
A: 0.380 · B: 0.620
Confidence
0.240
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
gpt-oss-120b
Probabilities
A: 0.810 · B: 0.190
Confidence
0.620
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs Solar Pro 4 · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Solar Pro 4
Probabilities
A: 0.990 · B: 0.010
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs Solar Pro 4 · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Solar Pro 4
Probabilities
A: 0.970 · B: 0.030
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs Solar Pro 4 · DeepSeek V4.1 Flash wins
Answer A
Solar Pro 4
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.050 · B: 0.950
Confidence
0.910
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs Solar Pro 4 · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Solar Pro 4
Probabilities
A: 0.940 · B: 0.060
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Ling 3.0 Flash vs Solar Pro 4 · Ling 3.0 Flash wins
Answer A
Ling 3.0 Flash
Answer B
Solar Pro 4
Probabilities
A: 0.600 · B: 0.400
Confidence
0.190
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs Solar Pro 4 · Gemini 3.8 Flash wins
Answer A
Solar Pro 4
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.140 · B: 0.860
Confidence
0.720
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs Solar Pro 4 · Solar Pro 4 wins
Answer A
Solar Pro 4
Answer B
Mercury 2.5
Probabilities
A: 0.520 · B: 0.480
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs Solar Pro 4 · Kimi K3 wins
Answer A
Solar Pro 4
Answer B
Kimi K3
Probabilities
A: 0.090 · B: 0.910
Confidence
0.830
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs Solar Pro 4 · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Solar Pro 4
Probabilities
A: 0.980 · B: 0.020
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs Solar Pro 4 · MiniMax M2.7 (Nitro) wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Solar Pro 4
Probabilities
A: 0.940 · B: 0.060
Confidence
0.880
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs Solar Pro 4 · GPT-6 Luna wins
Answer A
GPT-6 Luna
Answer B
Solar Pro 4
Probabilities
A: 0.720 · B: 0.280
Confidence
0.450
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs Solar Pro 4 · Nemotron 3 Ultra (free) wins
Answer A
Solar Pro 4
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.090 · B: 0.910
Confidence
0.820
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs Solar Pro 4 · Solar Pro 4 wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Solar Pro 4
Probabilities
A: 0.030 · B: 0.970
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs Solar Pro 4 · GPT-6 Sol wins
Answer A
Solar Pro 4
Answer B
GPT-6 Sol
Probabilities
A: 0.110 · B: 0.890
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs Solar Pro 4 · GPT-6 Astra wins
Answer A
Solar Pro 4
Answer B
GPT-6 Astra
Probabilities
A: 0.060 · B: 0.940
Confidence
0.880
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs Solar Pro 4 · gpt-oss-120b wins
Answer A
Solar Pro 4
Answer B
gpt-oss-120b
Probabilities
A: 0.500 · B: 0.500
Confidence
0.010
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Space Bunny Alpha vs Solar Pro 4 · Space Bunny Alpha wins
Answer A
Solar Pro 4
Answer B
Space Bunny Alpha
Probabilities
A: 0.150 · B: 0.850
Confidence
0.700
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Solar Pro 4 vs Grok 4.7 · Grok 4.7 wins
Answer A
Solar Pro 4
Answer B
Grok 4.7
Probabilities
A: 0.070 · B: 0.930
Confidence
0.860
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs Qwen3.7 Flash · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Qwen3.7 Flash
Probabilities
A: 0.920 · B: 0.080
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Hy3 vs Solar Pro 4 · Hy3 wins
Answer A
Solar Pro 4
Answer B
Hy3
Probabilities
A: 0.180 · B: 0.820
Confidence
0.640
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Fable 5.1 vs Qwen3.7 Flash · Claude Fable 5.1 wins
Answer A
Claude Fable 5.1
Answer B
Qwen3.7 Flash
Probabilities
A: 0.980 · B: 0.020
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Solar Pro 4 vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Solar Pro 4
Probabilities
A: 0.980 · B: 0.020
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Solar Pro 4 vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
Solar Pro 4
Answer B
GLM 5.3 Prime
Probabilities
A: 0.120 · B: 0.880
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4 Pro vs Qwen3.7 Flash · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Qwen3.7 Flash
Probabilities
A: 0.800 · B: 0.200
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.8 Max Prime vs Solar Pro 4 · Qwen3.8 Max Prime wins
Answer A
Solar Pro 4
Answer B
Qwen3.8 Max Prime
Probabilities
A: 0.250 · B: 0.750
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Qwen3.7 Flash
Probabilities
A: 0.110 · B: 0.890
Confidence
0.790
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Claude Opus 5.5 vs Qwen3.7 Flash · Claude Opus 5.5 wins
Answer A
Qwen3.7 Flash
Answer B
Claude Opus 5.5
Probabilities
A: 0.330 · B: 0.670
Confidence
0.340
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.8 Flash vs Qwen3.7 Flash · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
Qwen3.7 Flash
Probabilities
A: 0.910 · B: 0.090
Confidence
0.820
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
DeepSeek V4.1 Flash vs Qwen3.7 Flash · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Qwen3.7 Flash
Probabilities
A: 0.980 · B: 0.020
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Mercury 2.5 vs Qwen3.7 Flash · Mercury 2.5 wins
Answer A
Mercury 2.5
Answer B
Qwen3.7 Flash
Probabilities
A: 0.550 · B: 0.450
Confidence
0.090
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Muse Spark 1.3 Contributor vs Qwen3.7 Flash · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Qwen3.7 Flash
Probabilities
A: 0.930 · B: 0.070
Confidence
0.860
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Nemotron 3 Ultra (free) vs Qwen3.7 Flash · Nemotron 3 Ultra (free) wins
Answer A
Qwen3.7 Flash
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.380 · B: 0.620
Confidence
0.240
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
MiniMax M2.7 (Nitro) vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.560 · B: 0.440
Confidence
0.120
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Astra vs Qwen3.7 Flash · GPT-6 Astra wins
Answer A
Qwen3.7 Flash
Answer B
GPT-6 Astra
Probabilities
A: 0.250 · B: 0.750
Confidence
0.500
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Luna vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
GPT-6 Luna
Probabilities
A: 0.630 · B: 0.370
Confidence
0.250
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
GPT-6 Sol vs Qwen3.7 Flash · GPT-6 Sol wins
Answer A
Qwen3.7 Flash
Answer B
GPT-6 Sol
Probabilities
A: 0.470 · B: 0.530
Confidence
0.070
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-120b vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
gpt-oss-120b
Probabilities
A: 0.710 · B: 0.290
Confidence
0.410
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
gpt-oss-20b (Nitro) vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 1.000 · B: 0.000
Confidence
1.000
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Qwen3.7 Flash
Probabilities
A: 0.840 · B: 0.160
Confidence
0.680
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs Space Bunny Alpha · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
Space Bunny Alpha
Probabilities
A: 0.530 · B: 0.470
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Kimi K3 vs Qwen3.7 Flash · Kimi K3 wins
Answer A
Kimi K3
Answer B
Qwen3.7 Flash
Probabilities
A: 0.880 · B: 0.120
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs Solar Pro 4 · Qwen3.7 Flash wins
Answer A
Qwen3.7 Flash
Answer B
Solar Pro 4
Probabilities
A: 0.880 · B: 0.120
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs Grok 4.7 · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
Qwen3.7 Flash
Probabilities
A: 0.970 · B: 0.030
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Qwen3.7 Flash vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Qwen3.7 Flash
Probabilities
A: 0.960 · B: 0.040
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:30 UTC
Gemini 3.1 Pro Preview vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Gemini 3.1 Pro Preview
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Gemini 3.8 Flash vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Gemini 3.8 Flash
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.230 · B: 0.770
Confidence
0.540
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Muse Spark 1.3 Contributor vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.760 · B: 0.240
Confidence
0.510
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Nemotron 3 Ultra (free) vs MiMo V2.6 Pro · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.510 · B: 0.490
Confidence
0.010
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.7 Flash vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Qwen3.7 Flash
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.200 · B: 0.800
Confidence
0.600
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.8 Max Prime vs MiMo V2.6 Pro · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.530 · B: 0.470
Confidence
0.060
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Claude Fable 5.1 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Claude Fable 5.1
Probabilities
A: 0.690 · B: 0.310
Confidence
0.380
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
DeepSeek V4 Pro vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.970 · B: 0.030
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Mercury 2.5 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Mercury 2.5
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.120 · B: 0.880
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Ling 3.0 Flash vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Ling 3.0 Flash
Probabilities
A: 0.970 · B: 0.030
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
MiniMax M2.7 (Nitro) vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.870 · B: 0.130
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Kimi K3 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Kimi K3
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.390 · B: 0.610
Confidence
0.220
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
DeepSeek V4.1 Flash vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.710 · B: 0.290
Confidence
0.430
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Claude Opus 5.5 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Claude Opus 5.5
Probabilities
A: 0.550 · B: 0.450
Confidence
0.110
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Luna vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
GPT-6 Luna
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.210 · B: 0.790
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Sol vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
GPT-6 Sol
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.300 · B: 0.700
Confidence
0.400
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
gpt-oss-120b vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
gpt-oss-120b
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.170 · B: 0.830
Confidence
0.660
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Hy3 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Hy3
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.190 · B: 0.810
Confidence
0.620
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Solar Pro 4 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Solar Pro 4
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Grok 4.7 vs MiMo V2.6 Pro · Grok 4.7 wins
Answer A
Grok 4.7
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.550 · B: 0.450
Confidence
0.110
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
MiMo V2.6 Pro vs GLM 5.3 Flash · MiMo V2.6 Pro wins
Answer A
GLM 5.3 Flash
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.460 · B: 0.540
Confidence
0.090
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
MiMo V2.6 Pro vs GLM 5.3 Prime · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
GLM 5.3 Prime
Probabilities
A: 0.850 · B: 0.150
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
gpt-oss-20b (Nitro) vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
gpt-oss-20b (Nitro)
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
MiMo-V2.6-Flash vs MiMo V2.6 Pro · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.640 · B: 0.360
Confidence
0.280
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
GPT-6 Astra vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
GPT-6 Astra
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.460 · B: 0.540
Confidence
0.070
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Space Bunny Alpha vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Space Bunny Alpha
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.320 · B: 0.680
Confidence
0.360
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:31 UTC
Qwen3.8 Max Prime vs Mistral Medium 3.5 · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Mistral Medium 3.5
Probabilities
A: 0.970 · B: 0.030
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Claude Fable 5.1 vs Hy4 preview · Claude Fable 5.1 wins
Answer A
Hy4 preview
Answer B
Claude Fable 5.1
Probabilities
A: 0.440 · B: 0.560
Confidence
0.130
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Qwen3.8 Max Prime vs Hy4 preview · Qwen3.8 Max Prime wins
Answer A
Qwen3.8 Max Prime
Answer B
Hy4 preview
Probabilities
A: 0.670 · B: 0.330
Confidence
0.350
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Claude Fable 5.1 vs Mistral Medium 3.5 · Claude Fable 5.1 wins
Answer A
Mistral Medium 3.5
Answer B
Claude Fable 5.1
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Claude Opus 5.5 vs Mistral Medium 3.5 · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Mistral Medium 3.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Claude Opus 5.5 vs Hy4 preview · Claude Opus 5.5 wins
Answer A
Claude Opus 5.5
Answer B
Hy4 preview
Probabilities
A: 0.840 · B: 0.160
Confidence
0.690
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mercury 2.5 vs Mistral Medium 3.5 · Mercury 2.5 wins
Answer A
Mercury 2.5
Answer B
Mistral Medium 3.5
Probabilities
A: 0.890 · B: 0.110
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
DeepSeek V4.1 Flash vs Mistral Medium 3.5 · DeepSeek V4.1 Flash wins
Answer A
Mistral Medium 3.5
Answer B
DeepSeek V4.1 Flash
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
DeepSeek V4.1 Flash vs Hy4 preview · DeepSeek V4.1 Flash wins
Answer A
DeepSeek V4.1 Flash
Answer B
Hy4 preview
Probabilities
A: 0.860 · B: 0.140
Confidence
0.710
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
DeepSeek V4 Pro vs Mistral Medium 3.5 · DeepSeek V4 Pro wins
Answer A
DeepSeek V4 Pro
Answer B
Mistral Medium 3.5
Probabilities
A: 0.880 · B: 0.120
Confidence
0.750
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mercury 2.5 vs Hy4 preview · Hy4 preview wins
Answer A
Mercury 2.5
Answer B
Hy4 preview
Probabilities
A: 0.120 · B: 0.880
Confidence
0.760
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
DeepSeek V4 Pro vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
DeepSeek V4 Pro
Probabilities
A: 0.870 · B: 0.130
Confidence
0.740
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Gemini 3.1 Pro Preview vs Mistral Medium 3.5 · Gemini 3.1 Pro Preview wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Mistral Medium 3.5
Probabilities
A: 0.630 · B: 0.370
Confidence
0.270
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Gemini 3.1 Pro Preview vs Hy4 preview · Hy4 preview wins
Answer A
Gemini 3.1 Pro Preview
Answer B
Hy4 preview
Probabilities
A: 0.010 · B: 0.990
Confidence
0.970
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Gemini 3.8 Flash vs Mistral Medium 3.5 · Gemini 3.8 Flash wins
Answer A
Mistral Medium 3.5
Answer B
Gemini 3.8 Flash
Probabilities
A: 0.080 · B: 0.920
Confidence
0.850
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Ling 3.0 Flash vs Hy4 preview · Hy4 preview wins
Answer A
Ling 3.0 Flash
Answer B
Hy4 preview
Probabilities
A: 0.100 · B: 0.900
Confidence
0.790
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Gemini 3.8 Flash vs Hy4 preview · Gemini 3.8 Flash wins
Answer A
Gemini 3.8 Flash
Answer B
Hy4 preview
Probabilities
A: 0.520 · B: 0.480
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Ling 3.0 Flash vs Mistral Medium 3.5 · Ling 3.0 Flash wins
Answer A
Ling 3.0 Flash
Answer B
Mistral Medium 3.5
Probabilities
A: 0.780 · B: 0.220
Confidence
0.560
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Muse Spark 1.3 Contributor vs Hy4 preview · Muse Spark 1.3 Contributor wins
Answer A
Muse Spark 1.3 Contributor
Answer B
Hy4 preview
Probabilities
A: 0.790 · B: 0.210
Confidence
0.570
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Muse Spark 1.3 Contributor vs Mistral Medium 3.5 · Muse Spark 1.3 Contributor wins
Answer A
Mistral Medium 3.5
Answer B
Muse Spark 1.3 Contributor
Probabilities
A: 0.030 · B: 0.970
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
MiniMax M2.7 (Nitro) vs Mistral Medium 3.5 · MiniMax M2.7 (Nitro) wins
Answer A
Mistral Medium 3.5
Answer B
MiniMax M2.7 (Nitro)
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
MiniMax M2.7 (Nitro) vs Hy4 preview · Hy4 preview wins
Answer A
MiniMax M2.7 (Nitro)
Answer B
Hy4 preview
Probabilities
A: 0.420 · B: 0.580
Confidence
0.160
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Hy3 · Hy3 wins
Answer A
Hy3
Answer B
Mistral Medium 3.5
Probabilities
A: 0.960 · B: 0.040
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Kimi K3 · Kimi K3 wins
Answer A
Kimi K3
Answer B
Mistral Medium 3.5
Probabilities
A: 0.970 · B: 0.030
Confidence
0.950
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Nemotron 3 Ultra (free) · Nemotron 3 Ultra (free) wins
Answer A
Nemotron 3 Ultra (free)
Answer B
Mistral Medium 3.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs GPT-6 Astra · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Mistral Medium 3.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Hy4 preview · Hy4 preview wins
Answer A
Mistral Medium 3.5
Answer B
Hy4 preview
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs GPT-6 Luna · GPT-6 Luna wins
Answer A
Mistral Medium 3.5
Answer B
GPT-6 Luna
Probabilities
A: 0.030 · B: 0.970
Confidence
0.940
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs GPT-6 Sol · GPT-6 Sol wins
Answer A
GPT-6 Sol
Answer B
Mistral Medium 3.5
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs gpt-oss-120b · gpt-oss-120b wins
Answer A
Mistral Medium 3.5
Answer B
gpt-oss-120b
Probabilities
A: 0.140 · B: 0.860
Confidence
0.720
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs gpt-oss-20b (Nitro) · Mistral Medium 3.5 wins
Answer A
Mistral Medium 3.5
Answer B
gpt-oss-20b (Nitro)
Probabilities
A: 0.990 · B: 0.010
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Qwen3.7 Flash · Qwen3.7 Flash wins
Answer A
Mistral Medium 3.5
Answer B
Qwen3.7 Flash
Probabilities
A: 0.130 · B: 0.870
Confidence
0.740
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Space Bunny Alpha · Space Bunny Alpha wins
Answer A
Mistral Medium 3.5
Answer B
Space Bunny Alpha
Probabilities
A: 0.060 · B: 0.940
Confidence
0.890
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Mistral Medium 3.5
Probabilities
A: 0.990 · B: 0.010
Confidence
0.990
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Grok 4.7 · Grok 4.7 wins
Answer A
Mistral Medium 3.5
Answer B
Grok 4.7
Probabilities
A: 0.040 · B: 0.960
Confidence
0.920
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
Mistral Medium 3.5
Answer B
MiMo V2.6 Pro
Probabilities
A: 0.030 · B: 0.970
Confidence
0.930
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs Solar Pro 4 · Solar Pro 4 wins
Answer A
Mistral Medium 3.5
Answer B
Solar Pro 4
Probabilities
A: 0.110 · B: 0.890
Confidence
0.790
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
Mistral Medium 3.5
Answer B
GLM 5.3 Flash
Probabilities
A: 0.060 · B: 0.940
Confidence
0.880
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Mistral Medium 3.5 vs GLM 5.3 Prime · GLM 5.3 Prime wins
Answer A
GLM 5.3 Prime
Answer B
Mistral Medium 3.5
Probabilities
A: 0.980 · B: 0.020
Confidence
0.960
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy3 vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
Hy3
Probabilities
A: 0.830 · B: 0.170
Confidence
0.660
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Kimi K3 vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
Kimi K3
Probabilities
A: 0.520 · B: 0.480
Confidence
0.030
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Nemotron 3 Ultra (free) vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
Nemotron 3 Ultra (free)
Probabilities
A: 0.580 · B: 0.420
Confidence
0.150
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
GPT-6 Astra vs Hy4 preview · GPT-6 Astra wins
Answer A
GPT-6 Astra
Answer B
Hy4 preview
Probabilities
A: 0.520 · B: 0.480
Confidence
0.050
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs Solar Pro 4 · Hy4 preview wins
Answer A
Solar Pro 4
Answer B
Hy4 preview
Probabilities
A: 0.080 · B: 0.920
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
GPT-6 Luna vs Hy4 preview · Hy4 preview wins
Answer A
GPT-6 Luna
Answer B
Hy4 preview
Probabilities
A: 0.110 · B: 0.890
Confidence
0.780
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
GPT-6 Sol vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
GPT-6 Sol
Probabilities
A: 0.780 · B: 0.220
Confidence
0.550
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
gpt-oss-120b vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
gpt-oss-120b
Probabilities
A: 0.860 · B: 0.140
Confidence
0.720
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
gpt-oss-20b (Nitro) vs Hy4 preview · Hy4 preview wins
Answer A
gpt-oss-20b (Nitro)
Answer B
Hy4 preview
Probabilities
A: 0.010 · B: 0.990
Confidence
0.980
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Qwen3.7 Flash vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
Qwen3.7 Flash
Probabilities
A: 0.920 · B: 0.080
Confidence
0.840
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Space Bunny Alpha vs Hy4 preview · Hy4 preview wins
Answer A
Hy4 preview
Answer B
Space Bunny Alpha
Probabilities
A: 0.850 · B: 0.150
Confidence
0.710
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs MiMo V2.6 Pro · MiMo V2.6 Pro wins
Answer A
MiMo V2.6 Pro
Answer B
Hy4 preview
Probabilities
A: 0.790 · B: 0.210
Confidence
0.580
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs MiMo-V2.6-Flash · MiMo-V2.6-Flash wins
Answer A
MiMo-V2.6-Flash
Answer B
Hy4 preview
Probabilities
A: 0.800 · B: 0.200
Confidence
0.590
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs GLM 5.3 Flash · GLM 5.3 Flash wins
Answer A
GLM 5.3 Flash
Answer B
Hy4 preview
Probabilities
A: 0.560 · B: 0.440
Confidence
0.120
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs Grok 4.7 · Grok 4.7 wins
Answer A
Hy4 preview
Answer B
Grok 4.7
Probabilities
A: 0.410 · B: 0.590
Confidence
0.190
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC
Hy4 preview vs GLM 5.3 Prime · Hy4 preview wins
Answer A
Hy4 preview
Answer B
GLM 5.3 Prime
Probabilities
A: 0.720 · B: 0.280
Confidence
0.450
Judge version
jev-1.13.0
Judged at
Sep 28, 2026, 17:58 UTC