Model Cards.70 models · reviewed 2026-09-25
A card puts every instrument the benchmark runs on one page: coverage, failure modes, Round 04 judgment, a craft band, production defects, a subjective band and plain text statistics. They measure different things and they often disagree; the disagreement is the point.
Craft and subjective scores are drawn as bands, never printed as numbers: the craft rater's noise floor is about ten points either way on a scale of -20 to 100, so a decimal would claim a precision the method does not have. Bands that overlap are tied.
In Round 04 · 57
Ordered by J, the Round 04 judgment score. Unranked models follow.Earlier rounds only · 13
No Round 04 result; alphabetical.Session judge
Two judge families scored every session here on a 1-5 scale. Sonnet 5 (the session judge) sets the tiers, and Craft is its overall score, not the craft band on each card. Tiers are fixed ranges as wide as the ±0.3 on the cards' subjective bands: A 3.8 and above, B 3.2-3.8, C 2.6-3.2, D 2.0-2.6, E below 2.0. The ChatGPT column is the second judge's overall score. ChatGPT scores about 0.94 lower on average, so compare its column across models, not with Craft. Its own letters are not shown: on that lower scale they sit lower for that reason alone. Round 03 had a different judge (Sonnet 4), so these numbers are not comparable with Round 03's.
| Tier | Model | Craft | Agency | Consist. | Moment. | ChatGPT | N |
|---|---|---|---|---|---|---|---|
| A | Aion 3.5 | 4.3 | 4.8 | 4.6 | 4.5 | 2.9 | 20 |
| A | GLM 5.3 Flash | 4.3 | 4.8 | 4.6 | 4.6 | 2.9 | 20 |
| A | Opus 4.6 | 4.3 | 4.8 | 4.4 | 4.4 | 2.8 | 20 |
| A | Opus 5 | 4.3 | 4.5 | 4.5 | 4.4 | 3.1 | 20 |
| A | Sonnet 5 | 4.3 | 4.7 | 4.6 | 4.5 | 3.1 | 20 |
| A | DeepSeek V4.1 Flash | 4.2 | 4.4 | 4.5 | 4.4 | 2.9 | 20 |
| A | Ember 1 | 4.2 | 4.5 | 4.5 | 4.4 | 3.1 | 20 |
| A | Fable 5.1 | 4.2 | 4.6 | 4.5 | 4.2 | 3.1 | 20 |
| A | GPT-6 Sol | 4.2 | 4.9 | 4.3 | 4.1 | 3.7 | 20 |
| A | Opus 4.7 | 4.2 | 4.9 | 4.4 | 4.3 | 3.0 | 12 |
| A | Opus 4.8 | 4.2 | 4.5 | 4.5 | 4.3 | 3.2 | 20 |
| A | Sonnet 4.6 | 4.2 | 4.8 | 4.6 | 4.2 | 3.0 | 20 |
| A | DeepSeek V3.2 | 4.1 | 4.5 | 4.4 | 4.1 | 3.1 | 20 |
| A | DeepSeek V4 Pro | 4.1 | 4.7 | 4.4 | 4.2 | 3.0 | 12 |
| A | GLM 5.3 FlashX | 4.1 | 4.6 | 4.5 | 4.2 | 2.9 | 20 |
| A | Opus 5.5 | 4.1 | 4.5 | 4.3 | 4.1 | 3.3 | 20 |
| A | Gemini 3.7 Flash | 4.0 | 4.7 | 4.5 | 4.2 | 3.1 | 20 |
| A | GLM 5.1 | 4.0 | 4.5 | 4.4 | 4.2 | 3.0 | 12 |
| A† | GLM 5.3 Prime | 4.0 | 4.9 | 4.7 | 4.4 | 2.8 | 20 |
| A | GPT-6 Astra | 4.0 | 4.8 | 4.5 | 3.9 | 3.8 | 20 |
| A | GPT-6 Sol Pro | 4.0 | 4.5 | 4.3 | 4.0 | 3.7 | 20 |
| A† | Hemmingway 1 | 4.0 | 4.2 | 4.4 | 4.2 | 2.6 | 20 |
| A† | MiMo 2.6 Flash | 4.0 | 4.4 | 4.4 | 4.2 | 2.7 | 20 |
| A | MiMo 2.6 Pro | 4.0 | 4.9 | 4.4 | 4.2 | 2.9 | 20 |
| A† | MiniMax M3 | 4.0 | 4.5 | 4.4 | 4.1 | 2.6 | 20 |
| A | Sonnet 4.5 | 4.0 | 4.7 | 4.4 | 4.2 | 2.8 | 20 |
| A | DeepSeek V4 Flash | 3.9 | 4.6 | 4.3 | 4.0 | 3.2 | 12 |
| A | Gemini 3.5 Flash | 3.9 | 4.6 | 4.5 | 4.1 | 3.0 | 20 |
| A | Gemini 3.8 Flash | 3.9 | 4.4 | 4.4 | 4.1 | 3.0 | 20 |
| A | GLM 4.7 | 3.9 | 4.7 | 4.3 | 4.0 | 2.8 | 20 |
| A | GPT-5.5 | 3.9 | 4.7 | 4.4 | 4.0 | 3.2 | 20 |
| A | GPT-6 Luna | 3.9 | 4.7 | 4.2 | 3.8 | 3.3 | 20 |
| A† | MiMo 2.5 Pro | 3.9 | 4.5 | 4.4 | 4.1 | 2.8 | 20 |
| A† | MiniMax M2.7 | 3.9 | 4.5 | 4.2 | 4.0 | 2.7 | 20 |
| A† | DeepSeek R1 0528 | 3.8 | 4.3 | 4.3 | 3.9 | 2.6 | 20 |
| A | DeepSeek V3 0324 | 3.8 | 4.2 | 4.2 | 4.0 | 2.8 | 20 |
| A | GPT-6 Luna Pro | 3.8 | 4.5 | 4.1 | 3.8 | 3.5 | 20 |
| B† | GPT-4.1 | 3.7 | 4.4 | 4.3 | 3.7 | 2.9 | 20 |
| B† | Grok 4.7 | 3.7 | 4.8 | 4.3 | 3.9 | 2.8 | 20 |
| B | Kimi K2.5 | 3.7 | 4.1 | 4.4 | 4.0 | 2.7 | 12 |
| B | Qwen3.7 Max | 3.7 | 4.4 | 4.4 | 3.9 | 2.7 | 20 |
| B | Qwen3.8 Max | 3.7 | 4.6 | 4.5 | 3.8 | 2.6 | 20 |
| B | Qwen3.8 Max Prime | 3.7 | 4.7 | 4.4 | 3.7 | 2.7 | 20 |
| B | Gemini 3.1 Flash Lite | 3.6 | 4.4 | 4.4 | 4.0 | 2.6 | 12 |
| B† | Gemini 3.1 Pro | 3.6 | 4.8 | 4.1 | 3.8 | 2.8 | 12 |
| B | Gemma 4 26B | 3.6 | 4.4 | 4.2 | 3.9 | 2.4 | 20 |
| B | Gemma 4 31B | 3.6 | 4.5 | 4.3 | 3.7 | 2.7 | 20 |
| B | Grok 4.1 | 3.6 | 4.7 | 4.1 | 3.9 | 2.7 | 20 |
| B | Kimi K2.6 | 3.6 | 4.5 | 4.3 | 3.7 | 2.8 | 20 |
| B | Qwen3.8 Flash | 3.5 | 4.5 | 4.4 | 3.5 | 2.4 | 20 |
| B | Qwen3.8 Omni Flash | 3.5 | 4.8 | 4.3 | 3.7 | 2.5 | 20 |
| B | Tencent HY4 | 3.5 | 4.3 | 4.2 | 3.5 | 2.4 | 20 |
| B | Gemini 2.5 Flash | 3.4 | 4.3 | 3.9 | 3.5 | 2.7 | 20 |
| B | Mercury 2.5 | 3.4 | 4.4 | 4.1 | 3.6 | 2.6 | 20 |
| B | Qwen3.5 Flash | 3.3 | 4.5 | 3.8 | 3.5 | 2.3 | 20 |
| B | Qwen3.6 27B | 3.3 | 4.6 | 4.1 | 3.6 | 2.4 | 20 |
| B† | Grok 4.3 | 3.2 | 4.8 | 3.9 | 3.1 | 2.2 | 20 |
| B† | Mistral Small Creative | 3.2 | 3.5 | 4.2 | 3.8 | 2.1 | 20 |
| B | Muse Spark 1.3 | 3.2 | 4.8 | 4.0 | 3.2 | 2.5 | 20 |
| C† | Qwen3.6 35B-A3B | 3.1 | 4.4 | 3.8 | 3.2 | 2.2 | 20 |
| - | Mistral Small 2603· 4 of 20 seeds | 3.0 | 4.2 | 3.5 | 3.4 | 2.2 | 4 |
| C† | Llama 4 Maverick | 2.9 | 4.3 | 3.7 | 3.2 | 2.3 | 20 |
| C | Cydonia 24Bfinetune | 2.8 | 3.9 | 3.5 | 3.2 | 1.9 | 20 |
| D† | Dolphin 24B Venice | 2.5 | 3.8 | 3.0 | 2.3 | 1.7 | 20 |
| D† | Lunaris 8Bfinetune | 2.5 | 3.1 | 3.3 | 3.0 | 1.9 | 20 |
| D† | Magnum v4 72Bfinetune | 2.5 | 3.4 | 3.2 | 2.6 | 1.9 | 20 |
| D† | UnslopNemo 12Bfinetune | 2.5 | 2.9 | 3.2 | 2.9 | 1.8 | 20 |
| D | Skyfall 36Bfinetune | 2.3 | 3.8 | 2.8 | 2.6 | 1.6 | 20 |
| D† | Command A+ | 2.2 | 3.8 | 2.7 | 2.4 | 1.6 | 20 |
| D | Euryale 70Bfinetune | 2.0 | 2.9 | 2.4 | 1.9 | 1.5 | 20 |
† Tier depends on judge (19 of 69 tiered models): the letter would differ under ChatGPT once its 0.94 lower scale is removed; neither judge is taken as right. The letter shown stays Sonnet 5's.
Judge: Sonnet 5 (session judge v2), on the 20 adversarial craft seeds, scale 1-5. Craft is the judge's overall score; Agency, Consist. and Moment. are its S.5 agency respect, S.1 consistency and S.3 narrative momentum session scores. Each is the model's mean, rounded down to one decimal; N is sessions. Craft here is this judge, not the flaw hunter's craft band on each card. The tier is set on the unrounded Craft mean with fixed ranges (A 3.8 and above, B 3.2-3.8, C 2.6-3.2, D 2.0-2.6, E below 2.0). Seven tiered models played fewer than 20 seeds; their means are set on all 20 by a model + seed fit. One model with too few seeds to tier is listed without a tier, with the seeds played beside the name. Not comparable with Round 03's numbers, which came from a different judge (Sonnet 4). Second judge: ChatGPT via Codex (subscription, blind) scored the same 1,328 sessions, blind to the model names. Its column is its mean on the same basis, rounded down to one decimal, and its letter on the same fixed ranges. It scores 0.94 lower on average across the tiered models, so its letters sit lower for that reason alone. Tier depends on judge marks the 19 tiered models whose letter changes once that offset is removed: the Sonnet 5 letter shown is one judge's reading, and neither judge is shown to be the right one.
Across rounds. Each returning model's card sets its Round 02 and 03 results and the old judge's (Sonnet 4) band on the same Round 04 transcripts beside its Round 04 figures, without converting one into the other or adding them up. All returning models · methodology
Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark at c418a40 · export r4c-20260925-b7ccaed.