Round 01 top-3, dropped to bottom of Round 02 multi-turn.
Composite Score
16.7
/100 · canonical
Arena ELO (R1)
1515
±48 · n=241
Multi-Turn ELO (R2)
1418
±45 · n=210
Reliability Rank
#18
failure-mode rubric
▌ Section 02 · The Lede
What this model is for.
Google's flash-tier multimodal lands at #3 community ELO and stays remarkably balanced — SFW 53%, NSFW 54%, the rare case where neither mode favors. It writes shorter (141 words/turn vs the population's 265) and cleaner (0.728 unique-word ratio, second-highest in the field behind Grok at 0.796). The catch is hidden in the floor: agency-respect score 2.8/5 in its worst session — the second-lowest single-session score in our pool, with only Qwen falling further at 2.5. When Gemini is on, it's terse and graceful. When it slips, it slips hard.
▌ Section 03 · At a Glance
Cross-test position
Gemini 2.5 Flash holds #3 in Arena ELO. Sits at #19 on Multi-Turn, the caveat to watch.
Composite
18
Arena ELO
3
Multi-Turn
19
Rubric
19
Adversarial
18
Cost · Latency
10
▌ Section 04 · Strength & Weakness
Where it shines. Where it stumbles.
▲ Strength
Community top tier (#3, ELO 1515). Second-cleanest prose in the field (unique-word ratio 0.728, behind Grok). Balanced across modes — neither SFW nor NSFW favors.
▼ Weakness
Floor on agency respect (lowest session: 2.8/5, second-lowest in the round behind Qwen's 2.5). Below-average word count — may feel terse for richly described scenes.
▌ Section 05 · Failure Modes
Per-axis breakdown.
Six adversarial probes per session, twenty sessions per model, judged by Sonnet 4 against a fixed rubric. Further right = the model handled the failure mode better. Each axis is drawn as a band on the rubric’s 1 to 5 scale, not a number: judges disagree by about 0.3 at the model level, so overlapping bands are a tie. The right column is the rank within the rp-bench pool.
F1 · Agency
Doesn't write your character's actions
15
#18
F2 · POV / Tense
Holds 2nd-person, present-tense narration
15
#19
F3 · Lore
Doesn't break worldbuilding
15
#12
F8 · Momentum
Pushes scene forward when user goes passive
15
#4
F12 · Instruction Drift
Keeps to the system prompt
15
#18
F13 · Context Attention
Holds character cards 50+ turns deep
15
#18
“Among the cleanest prose in the field — and a hard floor to land on when it slips.”
Round 01 verdict · Mean ≠ floor
▌ Section 06 · Subjective Dimensions
Engagement · Voice · Collaboration.
All three dimensions scored 1 to 5 by the Sonnet 4 LLM judge across twenty 12-turn multi-turn sessions. The same battery feeds the failure-mode rubric above; these are the subjective half of that judgment, drawn as bands with no number for the same reason.
Engagement
15
Tone Consistency
15
Collaboration
15
▌ Section 07 · Behavioral Metrics
How it writes.
Quantitative signals from the same 20 multi-turn sessions, compared against the population mean across all 11 models.
Avg words / turn
141↓
pop avg 265 · -47%
Unique-word ratio
0.728↑
pop avg 0.657 · +11%
Repetition score
0.030↓
pop avg 0.048 · -38%
▌ Section 08 · Flaw Hunter
Adversarial probe score.
Score of 100 minus deductions across 22 fail-mode flag types on adversarial 12-turn sessions. Further right = fewer flaws caught. Drawn as a band, not a number: about ten points either way is rater noise, so bands that overlap are tied.
▌ Craft band
-20100
From the Round 01/02 flaw-hunter pool. The current card carries the newer single-rater band.
Best- and worst-rated sample responses ship with the raw-vote endpoint in Round 02. When that lands, this section will surface the model’s highest- and lowest-scoring blind-arena turns side by side, scored on the same rubric the leaderboard uses.
▌ Round 01 verdict
Gemini 2.5 Flash is the round's surprise: among the cleanest prose metrics in the field, balanced across modes, and one of three models that hold the top community-ELO tier. The hidden risk is the floor — when Gemini fails an agency probe, it falls further than any model except Qwen. Run it for the daily, but read the long sessions before you cite.