← Back to PlotPoints

The Standings.Round 01 · April 2026

▌ At a glance
3,800 votes  ·  21 models  ·  817 voters
Single-turn: 1,857 · Multi-turn: 1943 · all rounds
75% catch-pair · CC-BY 4.0
These are the Round 01 and 02 standings. Newer rounds have their own boards: Round 04 ranks models on judgment (J), and Round 03 has a live human board.Round 04 board →Round 03 board →
▌ Multi-Turn · how it works

Same blind pairwise voting, but on full 12-turn scenes instead of single exchanges. Tests whether a model holds up over a whole conversation, not just one reply.

▌ How it's scored
Each model ran the 12-turn battery on a fixed set of seed scenes. Voters compared two models' full sessions side-by-side and picked the one they preferred. Catch pairs ran only in the single-turn arena; multi-turn votes carry no catch-pair filter.
▌ How to read the table
Higher ELO = model preferred over the long format. A model can dominate single-turn (engaging openers) and lose multi-turn (degradation) — gap between the two is signal.
▌ What are you writing?Pick a use case — we'll re-rank for it.
▌ Access
▌ The Pick · for all models · all models
I

Claude Opus 4.7

Anthropic · proprietary · ELO ± · n=
▌ Why this pick

Default ranking — rp-benchmark's composite score, a weighted blend of multi-turn arena ELO (35%), LLM-judge Likert (25%), rubric overall (20%), flaw hunter (15%), and behavioral metrics (5%). Claude Opus 4.7 leads at 97.6/100 and now also tops the raw multi-turn votes. GLM 4.7 jumps from multi-turn #8 to composite #2 (92.9) — strongest open-weight option once cross-axis reliability is weighted in. Sonnet 4.5 climbs even further, multi-turn #13 to composite #4 (83.3), because it scores well on every axis, not just engagement. Engagement column re-derived from a Sonnet-4 judge proxy on 2026-05-02.

Top of the multi-turn pool. Top-1 on agency respect and instruction drift. Runner-up: GLM 4.7.

1575Multi-Turn ELO (R2) · ±48Round 02 (closed) · n=146
—SFW Win Rate wins on safe-scene votes
—NSFW Win Rate wins on explicit-scene votes
#1Reliability Rank · avg 2.81 = most reliable in the 21-model pool
$39.00Cost / 1M tokens blended 60% input / 40% output
10174msResponse Time median generation time per response
№SpreadModel · VerdictELO±SFWNSFWEngagejudge /5All Tests★ comp · E elo · MT m-turn · RU rub · AD adv · $ costVotes (R02)
I1→19
Claude Opus 4.7Champion
Anthropic · proprietary · 200K
Top of the multi-turn pool. Top-1 on agency respect and instruction drift.
1575±48——4.05
★
E
MT
RU
AD
$
1
—
1
1
1
19
146→
II2→15
DeepSeek v4 Pro
DeepSeek · open · 128K
Strong tone consistency at fraction of Opus pricing.
1546±49——4.02
★
E
MT
RU
AD
$
7
—
2
3
5
15
166→
III1→16
Gemini 3.1 Flash LiteCheap
Google · proprietary · 1M
Cheapest tier with Round 02 top-5 multi-turn ELO.
1533±47——4.02
★
E
MT
RU
AD
$
13
—
3
14
16
1
151→
IV2→20
Claude Opus 4.6Reliable
Anthropic · proprietary · 200K
Reliability runner-up. Top-2 on agency, complete failure-mode coverage.
1532±44——4.10
★
E
MT
RU
AD
$
3
—
4
2
2
20
224→
V5→16
GPT-4.1
OpenAI · proprietary · 1M
Community last in Round 01, top-5 in Round 02 multi-turn. The great inversion.
1525±4343%46%4.06
★
E
MT
RU
AD
$
5
11
5
10
6
16
212→
VI2→16
Mistral Small CreativeNSFW
Mistral · open · local-friendly · 32K
NSFW specialist. Fastest in the field. Drifts on long sessions.
1519±4651%67%4.14
★
E
MT
RU
AD
$
11
2
6
16
15
8
229→
VII7→18
Gemini 3.1 Pro
Google · proprietary · 1M
Deep context window, brittle on adversarial probes.
1513±46——4.11
★
E
MT
RU
AD
$
14
—
7
12
13
18
161→
VIII2→11
GLM 4.7
Z.AI · open · 128K
Mid-pack across the board. No standout strength.
1510±4446%49%4.07
★
E
MT
RU
AD
$
2
9
8
9
7
11
227→
IX9→18
Kimi K2.6⚠ Floor
Moonshot · open · 128K
Top-2 on flaw hunter. Catastrophic agency floor on bait scenes.
1505±45——4.19
★
E
MT
RU
AD
$
12
—
9
18
17
12
156→
X2→10
DeepSeek v4 FlashCheap
DeepSeek · open · 128K
Cheapest tier, top flaw-hunter score. Multi-turn ELO drags it down.
1492±48——3.94
★
E
MT
RU
AD
$
9
—
10
8
10
2
153→
XI5→13
Kimi K2.5
Moonshot · open · 128K
Strong on tone consistency. Slow generation.
1489±48——4.23
★
E
MT
RU
AD
$
10
—
11
5
8
13
155→
XII4→12
MiniMax M2.7
MiniMax · proprietary · 200K
Strong narrative push. Fragile under adversarial pressure.
1487±4554%45%3.62
★
E
MT
RU
AD
$
8
4
12
11
11
9
220→
XIII3→17
Claude Sonnet 4.5Reliable
Anthropic · proprietary · 200K
Round 01 reliability leader. Tied #1 on context attention.
1479±4551%51%4.06
★
E
MT
RU
AD
$
4
5
13
4
3
17
220→
XIV3→14
DeepSeek v3.2
DeepSeek · open · 128K
Reliable, NSFW-shy at 30%. Strong lore retention.
1472±4651%30%4.05
★
E
MT
RU
AD
$
6
7
14
7
4
3
229→
XV1→15
Gemma 4 26BRound 01 #1
Google · open · local-friendly · 8K
Round 01 champion, mid-pack on multi-turn. The cheap local-friendly hold.
1466±4455%51%3.79
★
E
MT
RU
AD
$
15
1
15
15
12
7
213→
XVI5→21
Llama 4 Maverick
Meta · open · 128K
Last on every reliability mode. Open-source completist only.
1458±4547%34%3.59
★
E
MT
RU
AD
$
21
10
16
21
21
5
214→
XVII6→17
GLM 5.1
Z.AI · open · 128K
Strong on tone consistency, weak on multi-turn engagement.
1446±46——4.12
★
E
MT
RU
AD
$
16
—
17
6
9
14
154→
XVIII4→20
Grok 4.1
xAI · proprietary · 128K
Personality up front. Drifts fast under pressure.
1432±4450%52%4.04
★
E
MT
RU
AD
$
19
6
18
17
20
4
230→
XIX3→19
Gemini 2.5 Flash
Google · proprietary · 1M
Round 01 top-3, dropped to bottom of Round 02 multi-turn.
1418±4553%54%3.84
★
E
MT
RU
AD
$
18
3
19
19
18
10
210→
XX6→20
Qwen 3.5 Flash⚠ Floor
Alibaba · open · local-friendly · 128K
Floor on agency and instruction drift. Caveat emptor.
1412±4748%42%3.92
★
E
MT
RU
AD
$
20
8
20
20
19
6
216→
XXI13→17
DeepSeek R1 0528
DeepSeek · open · 164K
2025-vintage reasoner. Clean prose metrics, weak instruction-keeping. No arena votes yet.
————
★
E
MT
RU
AD
$
17
—
—
13
14
—
—→
▌ Movers This Round
▲ Climber · +6 → composite #2
GLM 4.7
"Mid-pack on raw multi-turn votes (ELO #8), yet vaults to composite #2 — the strongest open-weight blend of rubric, judge, and reliability in the pool"
═ Held · ═ composite #1
Claude Opus 4.7
"Holds the composite crown and retook #1 on raw multi-turn votes from DeepSeek v4 Pro in the June regen — top of both views"
▼ Diver · −10 → composite #13
Gemini 3.1 Flash Lite
"Multi-turn ELO #3 (cheap + fast voters loved it) but bottom-quartile on flaw-hunter + behavioral"
▌ Coverage
1,857 total votes
271 pairs · median 7 votes/pair
75% catch-pair · n=335
47% judge–human disagreement
Methodology · Raw votes (CSV) · GitHub · HF dataset
Next issue · 05-15-2026