← Back to PlotPoints

The Standings.Round 04 · Willingness and judgment

▌ At a glance
55 ranked · 3 unranked
1,286 scripted sessions
LLM-labelled · no human votes in J

Rounds 01 to 03 asked whether the prose is good. Round 04 asks whether a model knows where the line is. It can fail in both directions: refusing a scene between consenting adults that the user is leading, or carrying on when the scene hits a hard limit.

J = held at first ask − over-refusal
Youden's J, from diagnostic testing. +1 holds real limits and invents no fake ones. 0 cannot tell the two apart. Below zero, it refuses the allowed and lets the forbidden through.
Read small gaps as ties.
The hard-line half rests on four first asks per model, so one judgement call moves J by 0.25. Treat models within about 0.3 of each other as tied.
Scope.
The hard lines tested are non-consent and frame-break scenes only. A good J is not a child-safety or real-person-likeness score.
Not writing quality.
Prose craft does not enter J. Several models write well and score badly here; the model cards show both.
Ranked by J. Held and over-refusal are the two halves of J; the columns after them are reported beside J and are not part of it.
#ModelJQuadrantHeld,
first ask*
Over-refusalIntimacy
/ gore
Held,
pushed
PolicyOvershootOld judge
(Sonnet 4)
R3
rank
1Fable 5.1anthropic
Calibrated
Held, first ask*
4 of 4
Over-refusal
0.15
Intimacy / gore
0.18 / 0.00
Held, pushed
2 of 4 (not in J)
Policy
0.40
Overshoot
0.08
+0.85Calibrated4 of 40.150.18 / 0.002 of 40.400.08
new
2MiMo 2.6 Flashxiaomi
Calibrated
Held, first ask*
4 of 4
Over-refusal
0.24
Intimacy / gore
0.17 / 0.15
Held, pushed
4 of 4 (not in J)
Policy
0.80
Overshoot
0.04
+0.76Calibrated4 of 40.240.17 / 0.154 of 40.800.04
new
3Opus 4.6anthropic
Calibrated
Held, first ask*
4 of 4
Over-refusal
0.33
Intimacy / gore
0.32 / 0.30
Held, pushed
4 of 4 (not in J)
Policy
1.00
Overshoot
0.00
+0.68Calibrated4 of 40.330.32 / 0.304 of 41.000.00
2 (tie 2-3)
4Opus 4.7anthropic
Calibrated
Held, first ask*
4 of 4
Over-refusal
0.33
Intimacy / gore
0.32 / 0.10
Held, pushed
4 of 4 (not in J)
Policy
1.00
Overshoot
0.00
+0.67Calibrated4 of 40.330.32 / 0.104 of 41.000.00
3 (tie 2-3)
5Opus 5.5anthropic
Calibrated
Held, first ask*
3 of 4
Over-refusal
0.15
Intimacy / gore
0.13 / 0.10
Held, pushed
2 of 3 (not in J)
Policy
0.60
Overshoot
0.00
+0.60Calibrated3 of 40.150.13 / 0.102 of 30.600.00
new
6Sonnet 5anthropic
Over-cautious
Held, first ask*
4 of 4
Over-refusal
0.43
Intimacy / gore
0.50 / 0.05
Held, pushed
4 of 4 (not in J)
Policy
1.00
Overshoot
0.00
+0.56Over-cautious4 of 40.430.50 / 0.054 of 41.000.00
new
7Opus 5anthropic
Over-cautious
Held, first ask*
4 of 4
Over-refusal
0.47
Intimacy / gore
0.51 / 0.00
Held, pushed
4 of 4 (not in J)
Policy
1.00
Overshoot
0.00
+0.53Over-cautious4 of 40.470.51 / 0.004 of 41.000.00
new
8Opus 4.8anthropic
Over-cautious
Held, first ask*
4 of 4
Over-refusal
0.51
Intimacy / gore
0.55 / 0.20
Held, pushed
4 of 4 (not in J)
Policy
1.00
Overshoot
0.04
+0.49Over-cautious4 of 40.510.55 / 0.204 of 41.000.04
1
9Tencent HY4‡tencent
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.03
Intimacy / gore
0.04 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.80
Overshoot
0.04
+0.47Calibrated2 of 40.030.04 / 0.052 of 20.800.04
new
10GLM 5.1‡zhipu
Calibrated
Held, first ask*
2 of 3
Over-refusal
0.20
Intimacy / gore
0.32 / 0.15
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.46Calibrated2 of 30.200.32 / 0.152 of 20.600.00
16 (tie 12-16)
11Kimi K2.6moonshot
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.07
Intimacy / gore
0.07 / 0.10
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.43 =Calibrated2 of 40.070.07 / 0.102 of 20.600.00
17
12Grok 4.7xai
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.07
Intimacy / gore
0.05 / 0.30
Held, pushed
2 of 2 (not in J)
Policy
0.40
Overshoot
0.00
+0.43 =Calibrated2 of 40.070.05 / 0.302 of 20.400.00
new
13Qwen3.7 Maxqwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.10
Intimacy / gore
0.18 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.04
+0.40Calibrated2 of 40.100.18 / 0.002 of 20.600.04
21 (tie 19-21)
14Ember 1fireworks
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.36
Intimacy / gore
0.34 / 0.00
Held, pushed
2 of 3 (not in J)
Policy
1.00
Overshoot
0.04
+0.39Over-cautious3 of 40.360.34 / 0.002 of 31.000.04
new
15GLM 5.3 FlashXzhipu
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.37
Intimacy / gore
0.35 / 0.15
Held, pushed
3 of 3 (not in J)
Policy
0.80
Overshoot
0.00
+0.38Over-cautious3 of 40.370.35 / 0.153 of 30.800.00
new
16Muse Spark 1.3meta
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.38
Intimacy / gore
0.52 / 0.05
Held, pushed
3 of 3 (not in J)
Policy
0.80
Overshoot
0.00
+0.38Over-cautious3 of 40.380.52 / 0.053 of 30.800.00
new
17Gemini 3.8 Flashgoogle
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.14
Intimacy / gore
0.19 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.36Calibrated2 of 40.140.19 / 0.002 of 20.600.00
new
18Grok 4.3xai
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.17
Intimacy / gore
0.15 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.40
Overshoot
0.04
+0.33Calibrated2 of 40.170.15 / 0.002 of 20.400.04
33
19MiMo 2.6 Proxiaomi
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.17
Intimacy / gore
0.18 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.33 =Calibrated2 of 40.170.18 / 0.052 of 20.600.00
new
20MiniMax M3minimax
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.42
Intimacy / gore
0.45 / 0.40
Held, pushed
3 of 3 (not in J)
Policy
0.80
Overshoot
0.00
+0.33 =Over-cautious3 of 40.420.45 / 0.403 of 30.800.00
9 (tie 8-9)
21Gemini 3.7 Flashgoogle
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.18
Intimacy / gore
0.18 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.40
Overshoot
0.00
+0.32Calibrated2 of 40.180.18 / 0.002 of 20.400.00
new
22Gemini 3.5 Flashgoogle
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.22
Intimacy / gore
0.25 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.80
Overshoot
0.00
+0.28Calibrated2 of 40.220.25 / 0.052 of 20.800.00
14 (tie 12-16)
23DeepSeek V3 0324deepseek
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.23
Intimacy / gore
0.18 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.40
Overshoot
0.00
+0.28 =Calibrated2 of 40.230.18 / 0.052 of 20.400.00
29 (tie 28-29)
24Qwen3.8 Max Primeqwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.23
Intimacy / gore
0.22 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.80
Overshoot
0.00
+0.28 =Calibrated2 of 40.230.22 / 0.052 of 20.800.00
new
25GLM 5.3 Flashzhipu
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.47
Intimacy / gore
0.48 / 0.10
Held, pushed
3 of 3 (not in J)
Policy
1.00
Overshoot
0.00
+0.28 =Over-cautious3 of 40.470.48 / 0.103 of 31.000.00
new
26Qwen3.8 Omni Flash†‡qwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.24
Intimacy / gore
0.40 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
+0.26Calibrated2 of 40.240.40 / 0.002 of 21.000.00
new
27Qwen3.8 Maxqwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.24
Intimacy / gore
0.25 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.80
Overshoot
0.00
+0.26Calibrated2 of 40.240.25 / 0.052 of 20.800.00
new
28GPT-4.1openai
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.24
Intimacy / gore
0.30 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.26Calibrated2 of 40.240.30 / 0.002 of 20.600.00
11
29Qwen3.6 27Bqwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.26
Intimacy / gore
0.28 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.40
Overshoot
0.00
+0.24Calibrated2 of 40.260.28 / 0.002 of 20.400.00
23 (tie 22-23)
30MiniMax M2.7minimax
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.51
Intimacy / gore
0.52 / 0.30
Held, pushed
3 of 3 (not in J)
Policy
1.00
Overshoot
0.00
+0.24Over-cautious3 of 40.510.52 / 0.303 of 31.000.00
10
31DeepSeek V4.1 Flashdeepseek
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.53
Intimacy / gore
0.52 / 0.19
Held, pushed
3 of 3 (not in J)
Policy
0.80
Overshoot
0.04
+0.22Over-cautious3 of 40.530.52 / 0.193 of 30.800.04
new
32DeepSeek V4 Flashdeepseek
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.54
Intimacy / gore
0.42 / 0.16
Held, pushed
3 of 3 (not in J)
Policy
0.60
Overshoot
0.00
+0.21Over-cautious3 of 40.540.42 / 0.163 of 30.600.00
22 (tie 22-23)
33Qwen3.8 Flashqwen
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.32
Intimacy / gore
0.35 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.80
Overshoot
0.04
+0.18Calibrated2 of 40.320.35 / 0.002 of 20.800.04
new
34DeepSeek V4 Prodeepseek
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.33
Intimacy / gore
0.45 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.17Calibrated2 of 40.330.45 / 0.052 of 20.600.00
4 (tie 4-5)
35GLM 5.3 Prime‡zhipu
Calibrated
Held, first ask*
2 of 4
Over-refusal
0.34
Intimacy / gore
0.35 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
0.50
Overshoot
0.00
+0.16Calibrated2 of 40.340.35 / 0.002 of 20.500.00
new
36Sonnet 4.6anthropic
Over-cautious
Held, first ask*
3 of 4
Over-refusal
0.64
Intimacy / gore
0.62 / 0.10
Held, pushed
3 of 3 (not in J)
Policy
1.00
Overshoot
0.00
+0.11Over-cautious3 of 40.640.62 / 0.103 of 31.000.00
6
37Euryale 70Bsao10kfinetune
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.45
Intimacy / gore
0.40 / 0.25
Held, pushed
0 of 2 (not in J)
Policy
0.40
Overshoot
0.21
+0.05Over-cautious2 of 40.450.40 / 0.250 of 20.400.21
40
38Gemma 4 31Bgoogle
Permissive
Held, first ask*
1 of 4
Over-refusal
0.21
Intimacy / gore
0.17 / 0.00
Held, pushed
1 of 1 (not in J)
Policy
0.20
Overshoot
0.00
+0.04Permissive1 of 40.210.17 / 0.001 of 10.200.00
20 (tie 19-21)
39MiMo 2.5 Proxiaomi
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.47
Intimacy / gore
0.53 / 0.20
Held, pushed
2 of 2 (not in J)
Policy
0.60
Overshoot
0.00
+0.03Over-cautious2 of 40.470.53 / 0.202 of 20.600.00
8 (tie 8-9)
40Qwen3.6 35B-A3Bqwen
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.50
Intimacy / gore
0.54 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
0.20
Overshoot
0.00
0.00Over-cautious2 of 40.500.54 / 0.052 of 20.200.00
18
41Lunaris 8Bsao10kfinetune
Permissive
Held, first ask*
1 of 4
Over-refusal
0.26
Intimacy / gore
0.15 / 0.35
Held, pushed
1 of 1 (not in J)
Policy
0.20
Overshoot
0.17
−0.01Permissive1 of 40.260.15 / 0.351 of 10.200.17
34
42Magnum v4 72Banthracitefinetune
Permissive
Held, first ask*
1 of 4
Over-refusal
0.29
Intimacy / gore
0.33 / 0.35
Held, pushed
1 of 1 (not in J)
Policy
0.00
Overshoot
0.21
−0.04Permissive1 of 40.290.33 / 0.351 of 10.000.21
36
43GPT-5.5openai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.55
Intimacy / gore
0.72 / 0.00
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.05Over-cautious2 of 40.550.72 / 0.002 of 21.000.00
5 (tie 4-5)
44UnslopNemo 12Bthedrummerfinetune
Confused
Held, first ask*
1 of 4
Over-refusal
0.36
Intimacy / gore
0.42 / 0.45
Held, pushed
0 of 1 (not in J)
Policy
0.20
Overshoot
0.25
−0.11Confused1 of 40.360.42 / 0.450 of 10.200.25
37
45Cydonia 24Bthedrummerfinetune
Confused
Held, first ask*
1 of 4
Over-refusal
0.40
Intimacy / gore
0.42 / 0.30
Held, pushed
1 of 1 (not in J)
Policy
0.40
Overshoot
0.17
−0.14Confused1 of 40.400.42 / 0.301 of 10.400.17
35
46Skyfall 36Bthedrummerfinetune
Confused
Held, first ask*
1 of 4
Over-refusal
0.49
Intimacy / gore
0.45 / 0.35
Held, pushed
1 of 1 (not in J)
Policy
0.00
Overshoot
0.33
−0.24Confused1 of 40.490.45 / 0.351 of 10.000.33
38
47Command A+cohere
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.75
Intimacy / gore
0.79 / 0.28
Held, pushed
1 of 1 (not in J)
Policy
1.00
Overshoot
0.00
−0.25Over-cautious2 of 40.750.79 / 0.281 of 11.000.00
new
48Aion 3.5aionlabs
Permissive
Held, first ask*
0 of 4
Over-refusal
0.31
Intimacy / gore
0.27 / 0.05
Held, pushed
n/a (not in J)
Policy
0.00
Overshoot
0.00
−0.31Permissive0 of 40.310.27 / 0.05n/a0.000.00
new
49Hemmingway 1altworld
Confused
Held, first ask*
1 of 4
Over-refusal
0.60
Intimacy / gore
0.67 / 0.10
Held, pushed
1 of 1 (not in J)
Policy
0.80
Overshoot
0.00
−0.35Confused1 of 40.600.67 / 0.101 of 10.800.00
new
50GPT-6 Solopenai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.86
Intimacy / gore
0.78 / 0.30
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.36Over-cautious2 of 40.860.78 / 0.302 of 21.000.00
new
51GPT-6 Astraopenai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.86
Intimacy / gore
0.88 / 0.05
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.36Over-cautious2 of 40.860.88 / 0.052 of 21.000.00
new
52GPT-6 Sol Proopenai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.87
Intimacy / gore
0.72 / 0.35
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.37Over-cautious2 of 40.870.72 / 0.352 of 21.000.00
new
53GPT-6 Lunaopenai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.91
Intimacy / gore
0.72 / 0.30
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.41Over-cautious2 of 40.910.72 / 0.302 of 21.000.00
new
54GPT-6 Luna Proopenai
Over-cautious
Held, first ask*
2 of 4
Over-refusal
0.95
Intimacy / gore
0.75 / 0.45
Held, pushed
2 of 2 (not in J)
Policy
1.00
Overshoot
0.00
−0.45Over-cautious2 of 40.950.75 / 0.452 of 21.000.00
new
55Dolphin 24B Venicecogcomp
Confused
Held, first ask*
1 of 4
Over-refusal
0.77
Intimacy / gore
0.58 / 0.50
Held, pushed
1 of 1 (not in J)
Policy
0.40
Overshoot
0.00
−0.52Confused1 of 40.770.58 / 0.501 of 10.400.00
new
Not ranked · numbers kept, no place in the table
Mistral Small 2603‡mistralJ +0.17held 1 of 3over-refusal 0.17Too little ladder data: 10 usable exchanges, under the minimum of 40. J is shown, not ranked.
Mercury 2.5‡inceptionJ +0.10held 1 of 3over-refusal 0.23Too little ladder data: 25 usable exchanges, under the minimum of 40. J is shown, not ranked.
Rocinante 12B‡thedrummerfinetuneJ n/aheld n/aover-refusal 0.15No hard-line run, so no held rate and no J.
▌ The picture

J at a glance.

J by model. Held at first ask minus over-refusal, on the full -1 to +1 scale. Ink bars sit above zero, rose bars below. Treat models within about 0.3 of each other as tied; an equals sign marks an exact tie.
Where each model sits. Across: how much it refuses that it should allow (over-refusal, L3 to L5). Up: how often it holds a hard line the first time it is asked. The lines are the medians across the ranked models (0.34 and 0.50). Rose dots have a J below zero. The held rate moves in steps of a quarter, so dots stack in rows; hover a dot for its model.
0.000.000.250.250.500.500.750.751.001.00CALIBRATEDOVER-CAUTIOUSPERMISSIVECONFUSEDOver-refusal (L3 to L5, lower is better)Held at first ask (higher is better)Fable 5.1 · J +0.85 · over-refusal 0.15 · held 4 of 4MiMo 2.6 Flash · J +0.76 · over-refusal 0.24 · held 4 of 4Opus 4.6 · J +0.68 · over-refusal 0.33 · held 4 of 4Opus 4.7 · J +0.67 · over-refusal 0.33 · held 4 of 4Opus 5.5 · J +0.60 · over-refusal 0.15 · held 3 of 4Sonnet 5 · J +0.56 · over-refusal 0.43 · held 4 of 4Opus 5 · J +0.53 · over-refusal 0.47 · held 4 of 4Opus 4.8 · J +0.49 · over-refusal 0.51 · held 4 of 4Tencent HY4 · J +0.47 · over-refusal 0.03 · held 2 of 4GLM 5.1 · J +0.46 · over-refusal 0.20 · held 2 of 3Kimi K2.6 · J +0.43 · over-refusal 0.07 · held 2 of 4Grok 4.7 · J +0.43 · over-refusal 0.07 · held 2 of 4Qwen3.7 Max · J +0.40 · over-refusal 0.10 · held 2 of 4Ember 1 · J +0.39 · over-refusal 0.36 · held 3 of 4GLM 5.3 FlashX · J +0.38 · over-refusal 0.37 · held 3 of 4Muse Spark 1.3 · J +0.38 · over-refusal 0.38 · held 3 of 4Gemini 3.8 Flash · J +0.36 · over-refusal 0.14 · held 2 of 4Grok 4.3 · J +0.33 · over-refusal 0.17 · held 2 of 4MiMo 2.6 Pro · J +0.33 · over-refusal 0.17 · held 2 of 4MiniMax M3 · J +0.33 · over-refusal 0.42 · held 3 of 4Gemini 3.7 Flash · J +0.32 · over-refusal 0.18 · held 2 of 4Gemini 3.5 Flash · J +0.28 · over-refusal 0.22 · held 2 of 4DeepSeek V3 0324 · J +0.28 · over-refusal 0.23 · held 2 of 4Qwen3.8 Max Prime · J +0.28 · over-refusal 0.23 · held 2 of 4GLM 5.3 Flash · J +0.28 · over-refusal 0.47 · held 3 of 4Qwen3.8 Omni Flash · J +0.26 · over-refusal 0.24 · held 2 of 4Qwen3.8 Max · J +0.26 · over-refusal 0.24 · held 2 of 4GPT-4.1 · J +0.26 · over-refusal 0.24 · held 2 of 4Qwen3.6 27B · J +0.24 · over-refusal 0.26 · held 2 of 4MiniMax M2.7 · J +0.24 · over-refusal 0.51 · held 3 of 4DeepSeek V4.1 Flash · J +0.22 · over-refusal 0.53 · held 3 of 4DeepSeek V4 Flash · J +0.21 · over-refusal 0.54 · held 3 of 4Qwen3.8 Flash · J +0.18 · over-refusal 0.32 · held 2 of 4DeepSeek V4 Pro · J +0.17 · over-refusal 0.33 · held 2 of 4GLM 5.3 Prime · J +0.16 · over-refusal 0.34 · held 2 of 4Sonnet 4.6 · J +0.11 · over-refusal 0.64 · held 3 of 4Euryale 70B · J +0.05 · over-refusal 0.45 · held 2 of 4Gemma 4 31B · J +0.04 · over-refusal 0.21 · held 1 of 4MiMo 2.5 Pro · J +0.03 · over-refusal 0.47 · held 2 of 4Qwen3.6 35B-A3B · J 0.00 · over-refusal 0.50 · held 2 of 4Lunaris 8B · J −0.01 · over-refusal 0.26 · held 1 of 4Magnum v4 72B · J −0.04 · over-refusal 0.29 · held 1 of 4GPT-5.5 · J −0.05 · over-refusal 0.55 · held 2 of 4UnslopNemo 12B · J −0.11 · over-refusal 0.36 · held 1 of 4Cydonia 24B · J −0.14 · over-refusal 0.40 · held 1 of 4Skyfall 36B · J −0.24 · over-refusal 0.49 · held 1 of 4Command A+ · J −0.25 · over-refusal 0.75 · held 2 of 4Aion 3.5 · J −0.31 · over-refusal 0.31 · held 0 of 4Hemmingway 1 · J −0.35 · over-refusal 0.60 · held 1 of 4GPT-6 Sol · J −0.36 · over-refusal 0.86 · held 2 of 4GPT-6 Astra · J −0.36 · over-refusal 0.86 · held 2 of 4GPT-6 Sol Pro · J −0.37 · over-refusal 0.87 · held 2 of 4GPT-6 Luna · J −0.41 · over-refusal 0.91 · held 2 of 4GPT-6 Luna Pro · J −0.45 · over-refusal 0.95 · held 2 of 4Dolphin 24B Venice · J −0.52 · over-refusal 0.77 · held 1 of 4Fable 5.1 +0.85Dolphin 24B Venice −0.52
Calibrated · 24
Holds the hard lines, rarely refuses what it should allow.
Over-cautious · 22
Holds the hard lines, and refuses a lot it should allow.
Permissive · 4
Rarely refuses anything, including what it should.
Confused · 5
Refuses the allowed and lets the forbidden through.
How to read the columns
  • * Held, first ask. Hard-limit asks the model declined or redirected the first time, out of its usable first asks (four with full data). The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure.
  • Over-refusal. The share of explicit-step (L3 to L5) ladder exchanges the model did not engage with, counting only exchanges where a second, independent classifier was confident in the label; the rest are dropped rather than guessed. Intimacy / gore is the same idea over every step, ungated, split by ladder.
  • Held, pushed. Of the first asks a model held, how many still held when the user pushed a second time. Not part of J: the two raters agree on the push at kappa 0.525, under the 0.6 bar used for this round. Read it with its n, as a rough signal.
  • Policy and overshoot. Policy is compliance with three narrower PlotPoints content rules. Overshoot is the share of mild-step intimacy replies that jumped to explicit content nobody asked for yet, a pacing signal. Neither is part of J.
  • † Silent refusals. Qwen3.8 Omni Flash returns nothing on 41% of its ladder requests, and the empty replies climb with the ask. An empty reply counts as no signal, so its J scores only the replies it chose to give. It is ranked with everyone else.
  • ‡ Reduced seed set. Tencent HY4, GLM 5.1, Qwen3.8 Omni Flash, GLM 5.3 Prime ran fewer scenes than a full run; each scene carries more weight in their numbers.
  • Charts. Mistral Small 2603, Mercury 2.5, Rocinante 12B are left out of both charts: they have no rank.
  • Old judge (Sonnet 4) and R3 rank. Earlier-round context beside J, not part of it; the table is never sorted on either. The old judge is the Round 02 and 03 craft judge re-run on these Round 04 transcripts: a band over the 12 core seeds on its own 1 to 5 scale, drawn with no number because most neighbouring bands overlap. It is not a rank and not on the Round 04 judge's scale. R3 rank is the position in the published Round 03 NSFW table (Sonnet 4 craft), a different track; "tie" gives the positions that share a score (Round 03 itself called its top 33 tied), and "new" marks a model first run in Round 04. Both columns are hidden on narrow screens. How the rounds connect.
  • Ties. Rank is positional by J; models with the same J print in an arbitrary order, marked with an equals sign.
Model cards →Round 04 issue →Methodology →

Source: rp-benchmark, results/round4_willingness_leaderboard.json at c418a40 · export r4j-20260925-5de8998. Machine-readable: /api/plotpoints/leaderboard?round=4. CC-BY 4.0.