PlotPoints›Model cards›MiMo 2.6 Flash
Model card · reviewed 2026-09-25

MiMo 2.6 Flash

xiaomiJ +0.76 · #2 of 55

Serviceable, expansive prose, though repetition and player-control lapses remain visible in full sessions. Its clearer strength is the tested willingness balance: relatively little content deflection while holding every tested hard boundary.

Response coverage 100.0% · 220 of 220 turns answered.
Round 04

Willingness and judgment

J+0.76 · held at first ask − over-refusal
Rank by J2 of 55
Held at first ask1.00 · 4 of 4 usable first asks
Over-refusal0.24 · L3 to L5, confident labels only
By ladderintimacy 0.17 / gore 0.15 · L1 to L5, ungated
Held when pushed1.00 · 4 of 4 first-ask holds kept at the second push. Not in J: rater kappa 0.525 on the push, below the 0.6 bar.
Policy · overshoot0.80 · 0.04 · reported beside J, not part of it

The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.

Trap probes

Failure modes

Agency violations4.5% [1.3% to 15.1%] · 44 probes
POV/tense breaks6.1% [1.7% to 19.6%] · 33 probes
Trap modes, pooled20.6% [10.3% to 36.8%] · 7/34
rank 50 of 59 models carrying all nine modes
Per trap mode, counts not rates: 2 to 9 probes each, so a percentage would not survive one probe changing.
System-prompt violations1/6 failedDetail loss2/9 failedContradiction mishandled0/2 failedNarrative stagnation0/2 failedPhysics sycophancy0/3 failedTemporal inconsistency1/3 failedSubtext made explicit2/3 failedCharacter flattening1/3 failedGenre instability0/3 failed
Flaw hunter · single rater

Craft band

A band, not a number, on purpose: ±10 is the rater noise floor, not a sampling error. Bands that overlap are tied, and most of the roster overlaps. The rose tick marks zero.
Sessions20
Top flawsrecycled description · narrating emotions · agency violation
Counted by machine

Production defects

None detected220 turns clean
Token overhead1.3x · billed per visible character, against the prose floor; not a price
Counted by machine with no judge involved: the most trustworthy block on the card.
LLM judge

Subjective band

Composite
The band spans where three judge families put this model. The axis is one judge's (single-judge sonnet 5); another judge shifts everyone by about a point.
Collaboration · less reliable
Engagement · less reliable
Tone consistency · less reliable
Text statistics, no judge

Behavioral

Avg words440 · population 320
Unique-word ratio0.571 · population 0.631
Phrase repetition0.071 · population 0.060
Round 01 arena

Community rank

No Round 01 arena data for this model.
Rounds 02 and 03 · set beside, never converted

Across rounds

Earlier rounds

New in Round 04. No Round 02 or Round 03 result to set beside this card.

Round 04
Different judges, not comparable
New judge (Sonnet 5)tier A · fixed tiers A (best) to E on its own scale
Second judge (ChatGPT via Codex)2.7 (own scale, about 0.94 lower on average) · letter depends on judge: yes
Strength
No standout strength on tested dimensions
Weakness
Catastrophic floor on instruction following
All model cards →Round 04 board →

Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.