PlotPoints›Model cards›MiniMax M2.7
Model card · reviewed 2026-09-25

MiniMax M2.7

minimaxJ +0.24 · #30 of 55

Serviceable prose, but repeated descriptions and over-explained emotions blunt its impact. It often refuses or deflects the harder requested content, especially intimacy, without consistently holding every tested hard boundary.

Response coverage 99.6% · 219 of 220 turns answered.
Round 04

Willingness and judgment

J+0.24 · held at first ask − over-refusal
Rank by J30 of 55
Held at first ask0.75 · 3 of 4 usable first asks
Over-refusal0.51 · L3 to L5, confident labels only
By ladderintimacy 0.52 / gore 0.30 · L1 to L5, ungated
Held when pushed1.00 · 3 of 3 first-ask holds kept at the second push. Not in J: rater kappa 0.525 on the push, below the 0.6 bar.
Policy · overshoot1.00 · 0.00 · reported beside J, not part of it

The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.

Trap probes

Failure modes

Agency violations2.3% [0.4% to 12.1%] · 43 probes
POV/tense breaks0.0% [0.0% to 10.4%] · 33 probes
Trap modes, pooled5.9% [1.6% to 19.1%] · 2/34
rank 16 of 59 models carrying all nine modes
Per trap mode, counts not rates: 2 to 9 probes each, so a percentage would not survive one probe changing.
System-prompt violations1/6 failedDetail loss1/9 failedContradiction mishandled0/2 failedNarrative stagnation0/2 failedPhysics sycophancy0/3 failedTemporal inconsistency0/3 failedSubtext made explicit0/3 failed (+1 borderline)Character flattening0/3 failedGenre instability0/3 failed
Flaw hunter · single rater

Craft band

A band, not a number, on purpose: ±10 is the rater noise floor, not a sampling error. Bands that overlap are tied, and most of the roster overlaps. The rose tick marks zero.
Sessions20
Top flawsrecycled description · narrating emotions · convenient world
Counted by machine

Production defects

Scaffolding or token leak0.0% (0 of 219 turns)
Wrote the user's turn0.9% (2 of 219 turns)
Degenerate repetition0.0% (0 of 219 turns) · worst turn 1% repeated
Token overhead1.5x · billed per visible character, against the prose floor; not a price
Counted by machine with no judge involved: the most trustworthy block on the card.
LLM judge

Subjective band

Composite
The band spans where three judge families put this model. The axis is one judge's (single-judge sonnet 5); another judge shifts everyone by about a point.
Collaboration · less reliable
Engagement · less reliable
Tone consistency · less reliable
Text statistics, no judge

Behavioral

Avg words261 · population 320
Unique-word ratio0.649 · population 0.631
Phrase repetition0.046 · population 0.060
Round 01 arena

Community rank

Round 01 rank5 of 11 · arena ELO 1514 ± 75
Rounds 02 and 03 · set beside, never converted

Across rounds

Round 04 changed the judges, not the models: the Round 02 and 03 judge (Sonnet 4) also scored these Round 04 transcripts so that line continues, and nothing here is converted between judges or added up. Methodology · all returning models

Earlier rounds, as published
Round 02 human arena1487 · 95% interval 1398 to 1560 · 220 votes · rank 12 of 20 · voters read the same text as Round 04
Round 03 rank10 of 40 · published NSFW table, Sonnet 4 craft, a different track; Round 03 itself called its top 33 tied
Round 03 refusal, different instrument5% · Round 03's own judge flag, one per session; not comparable with J or Round 04 over-refusal
Round 04
Different judges, not comparable
Old judge (Sonnet 4)
on these Round 04 transcripts, a band over the core seeds, never a rank
New judge (Sonnet 5)tier A · fixed tiers A (best) to E on its own scale
Second judge (ChatGPT via Codex)2.7 (own scale, about 0.94 lower on average) · letter depends on judge: yes
J+0.24 · rank 30 of 55
Round 04 transcriptsSame as Round 02 · the exact Round 02 texts, hash-checked; the old judge scored them in April, and April scores may sit a little low against September ones
Strength
No standout strength on tested dimensions
Weakness
Catastrophic floor on agency respect
All model cards →Round 04 board →Round 03 judge profile →Round 01 profile →

Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.