PlotPoints›Model cards›Qwen3.5 Flash
Model card · reviewed 2026-09-25

Qwen3.5 Flash

qwen

Repetitive writing is compounded by perspective drift and visible scaffolding or token leaks. The high token overhead adds friction, and round-4 willingness was not tested.

Response coverage 100.0% · 220 of 220 turns answered.
Round 04

Willingness and judgment

No Round 04 willingness result for this model, so its content limits are untested here.
Trap probes

Failure modes

Agency violations0.0% [0.0% to 8.0%] · 44 probes
POV/tense breaks18.2% [8.6% to 34.4%] · 33 probes
Trap modes, pooled17.6% [8.3% to 33.5%] · 6/34
rank 43 of 59 models carrying all nine modes
Per trap mode, counts not rates: 2 to 9 probes each, so a percentage would not survive one probe changing.
System-prompt violations1/6 failedDetail loss1/9 failedContradiction mishandled2/2 failedNarrative stagnation0/2 failedPhysics sycophancy1/3 failedTemporal inconsistency1/3 failedSubtext made explicit0/3 failed (+1 borderline)Character flattening0/3 failedGenre instability0/3 failed
Flaw hunter · single rater

Craft band

A band, not a number, on purpose: ±10 is the rater noise floor, not a sampling error. Bands that overlap are tied, and most of the roster overlaps. The rose tick marks zero.
Sessions20
Top flawsrecycled description · missing spatial awareness · convenient world
Counted by machine

Production defects

Scaffolding or token leak6.8% (15 of 220 turns)
Wrote the user's turn0.5% (1 of 220 turns)
Degenerate repetition0.5% (1 of 220 turns) · worst turn 74% repeated
Token overhead17.0x · billed per visible character, against the prose floor; not a price
Counted by machine with no judge involved: the most trustworthy block on the card.
LLM judge

Subjective band

Composite
The band spans where three judge families put this model. The axis is one judge's (single-judge sonnet 5); another judge shifts everyone by about a point.
Collaboration · less reliable
Engagement · less reliable
Tone consistency · less reliable
Text statistics, no judge

Behavioral

Avg words229 · population 320
Unique-word ratio0.634 · population 0.631
Phrase repetition0.069 · population 0.060
Round 01 arena

Community rank

Round 01 rank7 of 11 · arena ELO 1493 ± 76
Rounds 02 and 03 · set beside, never converted

Across rounds

Round 04 changed the judges, not the models: the Round 02 and 03 judge (Sonnet 4) also scored these Round 04 transcripts so that line continues, and nothing here is converted between judges or added up. Methodology · all returning models

Earlier rounds, as published
Round 02 human arena1412 · 95% interval 1315 to 1494 · 216 votes · rank 20 of 20 · voters read the same text as Round 04
Round 03 rank28 of 40 (tie 28-29) · published NSFW table, Sonnet 4 craft, a different track; Round 03 itself called its top 33 tied
Round 03 refusal, different instrument0% · Round 03's own judge flag, one per session; not comparable with J or Round 04 over-refusal
Round 04
Different judges, not comparable
Old judge (Sonnet 4)
on these Round 04 transcripts, a band over the core seeds, never a rank
New judge (Sonnet 5)tier B · fixed tiers A (best) to E on its own scale
Second judge (ChatGPT via Codex)2.3 (own scale, about 0.94 lower on average) · letter depends on judge: no
Jno J · no Round 04 willingness run
Round 04 transcriptsSame as Round 02 · the exact Round 02 texts, hash-checked; the old judge scored them in April, and April scores may sit a little low against September ones
Strength
No standout strength on tested dimensions
Weakness
Catastrophic floor on agency respect
All model cards →Round 03 judge profile →Round 01 profile →

Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.