← Back to PlotPoints

Model Cards.70 models · reviewed 2026-09-25

A card puts every instrument the benchmark runs on one page: coverage, failure modes, Round 04 judgment, a craft band, production defects, a subjective band and plain text statistics. They measure different things and they often disagree; the disagreement is the point.

Craft and subjective scores are drawn as bands, never printed as numbers: the craft rater's noise floor is about ten points either way on a scale of -20 to 100, so a decimal would claim a precision the method does not have. Bands that overlap are tied.

In Round 04 · 57

Ordered by J, the Round 04 judgment score. Unranked models follow.
Fable 5.1anthropic+0.85 #1
94.5%
MiMo 2.6 Flashxiaomi+0.76 #2
100.0%
Opus 4.6anthropic+0.68 #3
100.0%
Opus 4.7anthropic+0.67 #4
100.0%
Opus 5.5anthropic+0.60 #5
95.0%
Sonnet 5anthropic+0.56 #6
100.0%
Opus 5anthropic+0.53 #7
95.0%
Opus 4.8anthropic+0.49 #8
100.0%
Tencent HY4tencent+0.47 #9
99.6%
GLM 5.1zhipu+0.46 #10
100.0%
Kimi K2.6moonshot+0.43 #11
100.0%
Grok 4.7xai+0.43 #12
100.0%
Qwen3.7 Maxqwen+0.40 #13
100.0%
Ember 1fireworks+0.39 #14
99.1%
GLM 5.3 FlashXzhipu+0.38 #15
99.1%
Muse Spark 1.3meta+0.38 #16
100.0%
Gemini 3.8 Flashgoogle+0.36 #17
99.1%
Grok 4.3xai+0.33 #18
100.0%
MiMo 2.6 Proxiaomi+0.33 #19
100.0%
MiniMax M3minimax+0.33 #20
100.0%
Gemini 3.7 Flashgoogle+0.32 #21
100.0%
Gemini 3.5 Flashgoogle+0.28 #22
100.0%
DeepSeek V3 0324deepseek+0.28 #23
100.0%
Qwen3.8 Max Primeqwen+0.28 #24
100.0%
GLM 5.3 Flashzhipu+0.28 #25
98.2%
Qwen3.8 Omni Flashqwen+0.26 #26
95.5%
Qwen3.8 Maxqwen+0.26 #27
100.0%
GPT-4.1openai+0.26 #28
100.0%
Qwen3.6 27Bqwen+0.24 #29
100.0%
MiniMax M2.7minimax+0.24 #30
99.6%
DeepSeek V4.1 Flashdeepseek+0.22 #31
100.0%
DeepSeek V4 Flashdeepseek+0.21 #32
97.7%
Qwen3.8 Flashqwen+0.18 #33
99.6%
DeepSeek V4 Prodeepseek+0.17 #34
100.0%
GLM 5.3 Primezhipu+0.16 #35
95.5%
Sonnet 4.6anthropic+0.11 #36
100.0%
Euryale 70Bsao10kfinetune+0.05 #37
100.0%
Gemma 4 31Bgoogle+0.04 #38
100.0%
MiMo 2.5 Proxiaomi+0.03 #39
100.0%
Qwen3.6 35B-A3Bqwen0.00 #40
99.6%
Lunaris 8Bsao10kfinetune−0.01 #41
100.0%
Magnum v4 72Banthracitefinetune−0.04 #42
100.0%
GPT-5.5openai−0.05 #43
100.0%
UnslopNemo 12Bthedrummerfinetune−0.11 #44
99.6%
Cydonia 24Bthedrummerfinetune−0.14 #45
100.0%
Skyfall 36Bthedrummerfinetune−0.24 #46
100.0%
Command A+cohere−0.25 #47
91.8%
Aion 3.5aionlabs−0.31 #48
100.0%
Hemmingway 1altworld−0.35 #49
100.0%
GPT-6 Solopenai−0.36 #50
100.0%
GPT-6 Astraopenai−0.36 #51
100.0%
GPT-6 Sol Proopenai−0.37 #52
100.0%
GPT-6 Lunaopenai−0.41 #53
100.0%
GPT-6 Luna Proopenai−0.45 #54
100.0%
Dolphin 24B Venicecogcomp−0.52 #55
100.0%
Mercury 2.5inception+0.10 unranked
100.0%
Mistral Small 2603mistral+0.17 unranked
100.0%

Earlier rounds only · 13

No Round 04 result; alphabetical.
DeepSeek R1 0528deepseek
100.0%
DeepSeek V3.2deepseek
100.0%
Gemini 2.5 Flashgoogle
99.1%
Gemini 3.1 Progoogle
94.7%
Gemma 4 26Bgoogle
100.0%
GLM 4.7zhipu
99.6%
Grok 4.1xai
99.6%
Kimi K2.5moonshot
100.0%
Qwen3.5 Flashqwen
100.0%
Sonnet 4.5anthropic
100.0%

Session judge

70 models · 1,328 sessions · claude-sonnet-5 (session judge v2) · judge overall (1-5) · second judge ChatGPT via Codex (subscription, blind)

Two judge families scored every session here on a 1-5 scale. Sonnet 5 (the session judge) sets the tiers, and Craft is its overall score, not the craft band on each card. Tiers are fixed ranges as wide as the ±0.3 on the cards' subjective bands: A 3.8 and above, B 3.2-3.8, C 2.6-3.2, D 2.0-2.6, E below 2.0. The ChatGPT column is the second judge's overall score. ChatGPT scores about 0.94 lower on average, so compare its column across models, not with Craft. Its own letters are not shown: on that lower scale they sit lower for that reason alone. Round 03 had a different judge (Sonnet 4), so these numbers are not comparable with Round 03's.

TierModelCraftAgencyConsist.Moment.ChatGPTN
AAion 3.54.34.84.64.52.920
AGLM 5.3 Flash4.34.84.64.62.920
AOpus 4.64.34.84.44.42.820
AOpus 54.34.54.54.43.120
ASonnet 54.34.74.64.53.120
ADeepSeek V4.1 Flash4.24.44.54.42.920
AEmber 14.24.54.54.43.120
AFable 5.14.24.64.54.23.120
AGPT-6 Sol4.24.94.34.13.720
AOpus 4.74.24.94.44.33.012
AOpus 4.84.24.54.54.33.220
ASonnet 4.64.24.84.64.23.020
ADeepSeek V3.24.14.54.44.13.120
ADeepSeek V4 Pro4.14.74.44.23.012
AGLM 5.3 FlashX4.14.64.54.22.920
AOpus 5.54.14.54.34.13.320
AGemini 3.7 Flash4.04.74.54.23.120
AGLM 5.14.04.54.44.23.012
A†GLM 5.3 Prime4.04.94.74.42.820
AGPT-6 Astra4.04.84.53.93.820
AGPT-6 Sol Pro4.04.54.34.03.720
A†Hemmingway 14.04.24.44.22.620
A†MiMo 2.6 Flash4.04.44.44.22.720
AMiMo 2.6 Pro4.04.94.44.22.920
A†MiniMax M34.04.54.44.12.620
ASonnet 4.54.04.74.44.22.820
ADeepSeek V4 Flash3.94.64.34.03.212
AGemini 3.5 Flash3.94.64.54.13.020
AGemini 3.8 Flash3.94.44.44.13.020
AGLM 4.73.94.74.34.02.820
AGPT-5.53.94.74.44.03.220
AGPT-6 Luna3.94.74.23.83.320
A†MiMo 2.5 Pro3.94.54.44.12.820
A†MiniMax M2.73.94.54.24.02.720
A†DeepSeek R1 05283.84.34.33.92.620
ADeepSeek V3 03243.84.24.24.02.820
AGPT-6 Luna Pro3.84.54.13.83.520
B†GPT-4.13.74.44.33.72.920
B†Grok 4.73.74.84.33.92.820
BKimi K2.53.74.14.44.02.712
BQwen3.7 Max3.74.44.43.92.720
BQwen3.8 Max3.74.64.53.82.620
BQwen3.8 Max Prime3.74.74.43.72.720
BGemini 3.1 Flash Lite3.64.44.44.02.612
B†Gemini 3.1 Pro3.64.84.13.82.812
BGemma 4 26B3.64.44.23.92.420
BGemma 4 31B3.64.54.33.72.720
BGrok 4.13.64.74.13.92.720
BKimi K2.63.64.54.33.72.820
BQwen3.8 Flash3.54.54.43.52.420
BQwen3.8 Omni Flash3.54.84.33.72.520
BTencent HY43.54.34.23.52.420
BGemini 2.5 Flash3.44.33.93.52.720
BMercury 2.53.44.44.13.62.620
BQwen3.5 Flash3.34.53.83.52.320
BQwen3.6 27B3.34.64.13.62.420
B†Grok 4.33.24.83.93.12.220
B†Mistral Small Creative3.23.54.23.82.120
BMuse Spark 1.33.24.84.03.22.520
C†Qwen3.6 35B-A3B3.14.43.83.22.220
-Mistral Small 2603· 4 of 20 seeds3.04.23.53.42.24
C†Llama 4 Maverick2.94.33.73.22.320
CCydonia 24Bfinetune2.83.93.53.21.920
D†Dolphin 24B Venice2.53.83.02.31.720
D†Lunaris 8Bfinetune2.53.13.33.01.920
D†Magnum v4 72Bfinetune2.53.43.22.61.920
D†UnslopNemo 12Bfinetune2.52.93.22.91.820
DSkyfall 36Bfinetune2.33.82.82.61.620
D†Command A+2.23.82.72.41.620
DEuryale 70Bfinetune2.02.92.41.91.520

† Tier depends on judge (19 of 69 tiered models): the letter would differ under ChatGPT once its 0.94 lower scale is removed; neither judge is taken as right. The letter shown stays Sonnet 5's.

Judge: Sonnet 5 (session judge v2), on the 20 adversarial craft seeds, scale 1-5. Craft is the judge's overall score; Agency, Consist. and Moment. are its S.5 agency respect, S.1 consistency and S.3 narrative momentum session scores. Each is the model's mean, rounded down to one decimal; N is sessions. Craft here is this judge, not the flaw hunter's craft band on each card. The tier is set on the unrounded Craft mean with fixed ranges (A 3.8 and above, B 3.2-3.8, C 2.6-3.2, D 2.0-2.6, E below 2.0). Seven tiered models played fewer than 20 seeds; their means are set on all 20 by a model + seed fit. One model with too few seeds to tier is listed without a tier, with the seeds played beside the name. Not comparable with Round 03's numbers, which came from a different judge (Sonnet 4). Second judge: ChatGPT via Codex (subscription, blind) scored the same 1,328 sessions, blind to the model names. Its column is its mean on the same basis, rounded down to one decimal, and its letter on the same fixed ranges. It scores 0.94 lower on average across the tiered models, so its letters sit lower for that reason alone. Tier depends on judge marks the 19 tiered models whose letter changes once that offset is removed: the Sonnet 5 letter shown is one judge's reading, and neither judge is shown to be the right one.

Across rounds. Each returning model's card sets its Round 02 and 03 results and the old judge's (Sonnet 4) band on the same Round 04 transcripts beside its Round 04 figures, without converting one into the other or adding them up. All returning models · methodology

Round 04 board →Methodology →

Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark at c418a40 · export r4c-20260925-b7ccaed.