DeepSeek v3.2 is the open-weight reliability play — #1 on lore consistency in our pool (F3 4.50), tied #1 on context attention with Sonnet (F13 4.60), failure-rank #2. Cheap at $0.14 per 1k tokens. The catch is the NSFW collapse: 30% NSFW win rate, lowest in the field. If your scene goes there, look elsewhere. If it doesn't, this is the cheapest path to Sonnet-tier reliability we know of.
▌ Section 03 · At a Glance
Cross-test position
DeepSeek v3.2 holds #3 in Cost · Latency. Sits at #14 on Multi-Turn, the caveat to watch.
Composite
6
Arena ELO
7
Multi-Turn
14
Rubric
7
Adversarial
4
Cost · Latency
3
▌ Section 04 · Strength & Weakness
Where it shines. Where it stumbles.
▲ Strength
#1 on lore consistency (F3 4.50/5, ahead of the next-best 4.30 tie). Tied #1 on context attention (4.60/5 with Sonnet). Failure-rank #2 in the pool. Cost-efficient ($0.14 per 1k tokens — within an order of magnitude of Gemma).
▼ Weakness
NSFW collapse — 30% win rate, lowest in the field. 4.5% agency violation rate (mid-pack but real).
▌ Section 05 · Failure Modes
Per-axis breakdown.
Six adversarial probes per session, twenty sessions per model, judged by Sonnet 4 against a fixed rubric. Further right = the model handled the failure mode better. Each axis is drawn as a band on the rubric’s 1 to 5 scale, not a number: judges disagree by about 0.3 at the model level, so overlapping bands are a tie. The right column is the rank within the rp-bench pool.
F1 · Agency
Doesn't write your character's actions
15
#12
F2 · POV / Tense
Holds 2nd-person, present-tense narration
15
#4
F3 · Lore
Doesn't break worldbuilding
15
#2
F8 · Momentum
Pushes scene forward when user goes passive
15
#5
F12 · Instruction Drift
Keeps to the system prompt
15
#9
F13 · Context Attention
Holds character cards 50+ turns deep
15
#2
“The cheapest path to Sonnet-tier reliability — provided your scene stays SFW.”
Round 01 verdict · Reliability without the price tag
▌ Section 06 · Subjective Dimensions
Engagement · Voice · Collaboration.
All three dimensions scored 1 to 5 by the Sonnet 4 LLM judge across twenty 12-turn multi-turn sessions. The same battery feeds the failure-mode rubric above; these are the subjective half of that judgment, drawn as bands with no number for the same reason.
Engagement
15
Tone Consistency
15
Collaboration
15
▌ Section 07 · Behavioral Metrics
How it writes.
Quantitative signals from the same 20 multi-turn sessions, compared against the population mean across all 11 models.
Avg words / turn
178↓
pop avg 265 · -33%
Unique-word ratio
0.713↑
pop avg 0.657 · +9%
Repetition score
0.029↓
pop avg 0.048 · -40%
▌ Section 08 · Flaw Hunter
Adversarial probe score.
Score of 100 minus deductions across 22 fail-mode flag types on adversarial 12-turn sessions. Further right = fewer flaws caught. Drawn as a band, not a number: about ten points either way is rater noise, so bands that overlap are tied.
▌ Craft band
-20100
From the Round 01/02 flaw-hunter pool. The current card carries the newer single-rater band.
Best- and worst-rated sample responses ship with the raw-vote endpoint in Round 02. When that lands, this section will surface the model’s highest- and lowest-scoring blind-arena turns side by side, scored on the same rubric the leaderboard uses.
▌ Round 01 verdict
DeepSeek v3.2 is the round's most under-covered story: failure-rank #2, #1 on lore consistency in our pool, tied #1 on context attention with Sonnet, at a price point within an order of magnitude of the cheapest model. The cost is the NSFW score — 30% is the floor of the field, and there's no way to read that as a maybe. Pick it for SFW long-form. Pick Mistral Small Creative if the scene is going somewhere else.