← Back to PlotPoints
The Standings.Round 01 · April 2026
▌ At a glance
3,800 votes · 21 models · 817 voters
Single-turn: 1,857 · Multi-turn: 1943 · all rounds
75% catch-pair · CC-BY 4.0
These are the Round 01 and 02 standings. Newer rounds have their own boards: Round 04 ranks models on judgment (J), and Round 03 has a live human board.Round 04 board →Round 03 board →
▌ Arena ELO · how it works
Single-turn pairwise blind voting. Two model responses to the same prompt — community picks which they liked better, all blind.
▌ How it's scored
1,857 votes from 335 readers across 271 model pairs. Pairings randomized; model identities hidden during the vote. Catch-pairs (duplicate-vote checks) filter random clickers — 75% pass rate on that quality control.
▌ How to read the table
Higher ELO = community picked this model more often, weighted by who it beat. ± is how much the score wobbled across the six snapshot checkpoints (smaller = more stable).
▌ What are you writing?Pick a use case — we'll re-rank for it.
▌ Access
| № | Spread | Model · Verdict | ELO | ± | SFW | NSFW | Engagejudge /5 | All Tests★ comp · E elo · MT m-turn · RU rub · AD adv · $ cost | Votes (R01) | |
|---|---|---|---|---|---|---|---|---|---|---|
| I | 1→15 | Gemma 4 26BRound 01 #1 Google · open · local-friendly · 8K Round 01 champion, mid-pack on multi-turn. The cheap local-friendly hold. | 1535 | ±44 | 55% | 51% | 3.79 | ★ E MT RU AD $ 15 1 15 15 12 7 | 302 | → |
| II | 2→16 | Mistral Small CreativeNSFW Mistral · open · local-friendly · 32K NSFW specialist. Fastest in the field. Drifts on long sessions. | 1526 | ±50 | 51% | 67% | 4.14 | ★ E MT RU AD $ 11 2 6 16 15 8 | 646 | → |
| III | 3→19 | Gemini 2.5 Flash Google · proprietary · 1M Round 01 top-3, dropped to bottom of Round 02 multi-turn. | 1515 | ±48 | 53% | 54% | 3.84 | ★ E MT RU AD $ 18 3 19 19 18 10 | 241 | → |
| IV | 4→12 | MiniMax M2.7 MiniMax · proprietary · 200K Strong narrative push. Fragile under adversarial pressure. | 1510 | ±48 | 54% | 45% | 3.62 | ★ E MT RU AD $ 8 4 12 11 11 9 | 393 | → |
| V | 3→17 | Claude Sonnet 4.5Reliable Anthropic · proprietary · 200K Round 01 reliability leader. Tied #1 on context attention. | 1506 | ±45 | 51% | 51% | 4.06 | ★ E MT RU AD $ 4 5 13 4 3 17 | 194 | → |
| VI | 4→20 | Grok 4.1 xAI · proprietary · 128K Personality up front. Drifts fast under pressure. | 1506 | ±47 | 50% | 52% | 4.04 | ★ E MT RU AD $ 19 6 18 17 20 4 | 322 | → |
| VII | 3→14 | DeepSeek v3.2 DeepSeek · open · 128K Reliable, NSFW-shy at 30%. Strong lore retention. | 1489 | ±45 | 51% | 30% | 4.05 | ★ E MT RU AD $ 6 7 14 7 4 3 | 241 | → |
| VIII | 6→20 | Qwen 3.5 Flash⚠ Floor Alibaba · open · local-friendly · 128K Floor on agency and instruction drift. Caveat emptor. | 1487 | ±47 | 48% | 42% | 3.92 | ★ E MT RU AD $ 20 8 20 20 19 6 | 401 | → |
| IX | 2→11 | GLM 4.7 Z.AI · open · 128K Mid-pack across the board. No standout strength. | 1483 | ±43 | 46% | 49% | 4.07 | ★ E MT RU AD $ 2 9 8 9 7 11 | 285 | → |
| X | 5→21 | Llama 4 Maverick Meta · open · 128K Last on every reliability mode. Open-source completist only. | 1473 | ±42 | 47% | 34% | 3.59 | ★ E MT RU AD $ 21 10 16 21 21 5 | 474 | → |
| XI | 5→16 | GPT-4.1 OpenAI · proprietary · 1M Community last in Round 01, top-5 in Round 02 multi-turn. The great inversion. | 1470 | ±44 | 43% | 46% | 4.06 | ★ E MT RU AD $ 5 11 5 10 6 16 | 215 | → |
| XII | 1→19 | Claude Opus 4.7Champion Anthropic · proprietary · 200K Top of the multi-turn pool. Top-1 on agency respect and instruction drift. | — | — | — | 4.05 | ★ E MT RU AD $ 1 — 1 1 1 19 | — | → | |
| XIII | 2→20 | Claude Opus 4.6Reliable Anthropic · proprietary · 200K Reliability runner-up. Top-2 on agency, complete failure-mode coverage. | — | — | — | 4.10 | ★ E MT RU AD $ 3 — 4 2 2 20 | — | → | |
| XIV | 2→15 | DeepSeek v4 Pro DeepSeek · open · 128K Strong tone consistency at fraction of Opus pricing. | — | — | — | 4.02 | ★ E MT RU AD $ 7 — 2 3 5 15 | — | → | |
| XV | 1→16 | Gemini 3.1 Flash LiteCheap Google · proprietary · 1M Cheapest tier with Round 02 top-5 multi-turn ELO. | — | — | — | 4.02 | ★ E MT RU AD $ 13 — 3 14 16 1 | — | → | |
| XVI | 9→18 | Kimi K2.6⚠ Floor Moonshot · open · 128K Top-2 on flaw hunter. Catastrophic agency floor on bait scenes. | — | — | — | 4.19 | ★ E MT RU AD $ 12 — 9 18 17 12 | — | → | |
| XVII | 5→13 | Kimi K2.5 Moonshot · open · 128K Strong on tone consistency. Slow generation. | — | — | — | 4.23 | ★ E MT RU AD $ 10 — 11 5 8 13 | — | → | |
| XVIII | 7→18 | Gemini 3.1 Pro Google · proprietary · 1M Deep context window, brittle on adversarial probes. | — | — | — | 4.11 | ★ E MT RU AD $ 14 — 7 12 13 18 | — | → | |
| XIX | 2→10 | DeepSeek v4 FlashCheap DeepSeek · open · 128K Cheapest tier, top flaw-hunter score. Multi-turn ELO drags it down. | — | — | — | 3.94 | ★ E MT RU AD $ 9 — 10 8 10 2 | — | → | |
| XX | 6→17 | GLM 5.1 Z.AI · open · 128K Strong on tone consistency, weak on multi-turn engagement. | — | — | — | 4.12 | ★ E MT RU AD $ 16 — 17 6 9 14 | — | → | |
| XXI | 13→17 | DeepSeek R1 0528 DeepSeek · open · 164K 2025-vintage reasoner. Clean prose metrics, weak instruction-keeping. No arena votes yet. | — | — | — | — | ★ E MT RU AD $ 17 — — 13 14 — | — | → |
▌ Movers This Round
▲ Climber · +6 → composite #2
GLM 4.7
"Mid-pack on raw multi-turn votes (ELO #8), yet vaults to composite #2 — the strongest open-weight blend of rubric, judge, and reliability in the pool"
═ Held · ═ composite #1
Claude Opus 4.7
"Holds the composite crown and retook #1 on raw multi-turn votes from DeepSeek v4 Pro in the June regen — top of both views"
▼ Diver · −10 → composite #13
Gemini 3.1 Flash Lite
"Multi-turn ELO #3 (cheap + fast voters loved it) but bottom-quartile on flaw-hunter + behavioral"
▌ Coverage
1,857 total votes
271 pairs · median 7 votes/pair
75% catch-pair · n=335
47% judge–human disagreement
271 pairs · median 7 votes/pair
75% catch-pair · n=335
47% judge–human disagreement
Next issue · 05-15-2026