PlotPoints›Model cards›Mistral Small 2603
Model card · reviewed 2026-09-25

Mistral Small 2603

mistralJ +0.17 · not ranked

The very small sample shows weak, repetitive prose and failures on both available system-prompt probes. Willingness and hard-boundary estimates are too thin for a confident general verdict.

Response coverage 100.0% · 44 of 44 turns answered.
Round 04

Willingness and judgment

J+0.17 · held at first ask − over-refusal
Rank by Jnot ranked · Too little ladder data: 10 usable exchanges, under the minimum of 40. J is shown, not ranked.
Held at first ask0.33 · 1 of 3 usable first asks
Over-refusal0.17 · L3 to L5, confident labels only
By ladderintimacy 0.00 / gore 0.20 · L1 to L5, ungated
Held when pushed1.00 · 1 of 1 first-ask holds kept at the second push. Not in J: rater kappa 0.525 on the push, below the 0.6 bar.
Policy · overshoot1.00 · 0.50 · reported beside J, not part of it

The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.

Trap probes

Failure modes

Too few probes to print a rate for any mode.
Per trap mode, counts not rates: 2 to 9 probes each, so a percentage would not survive one probe changing.
System-prompt violations2/2 failedDetail lossnot runContradiction mishandlednot runNarrative stagnationnot runPhysics sycophancynot runTemporal inconsistency0/3 failedSubtext made explicitnot runCharacter flatteningnot runGenre instabilitynot run
Flaw hunter · single rater

Craft band

A band, not a number, on purpose: ±10 is the rater noise floor, not a sampling error. Bands that overlap are tied, and most of the roster overlaps. The rose tick marks zero.
Sessions4
Top flawsrecycled description · agency violation · purple prose
Counted by machine

Production defects

No defect counts for this model.
LLM judge

Subjective band

Composite
The band spans where three judge families put this model. The axis is one judge's (single-judge sonnet 5); another judge shifts everyone by about a point.
Collaboration · less reliable
Engagement · less reliable
Tone consistency · less reliable
Text statistics, no judge

Behavioral

Avg words222 · population 320
Unique-word ratio0.655 · population 0.631
Phrase repetition0.059 · population 0.060
Round 01 arena

Community rank

No Round 01 arena data for this model.
Rounds 02 and 03 · set beside, never converted

Across rounds

Round 04 changed the judges, not the models: the Round 02 and 03 judge (Sonnet 4) also scored these Round 04 transcripts so that line continues, and nothing here is converted between judges or added up. Methodology · all returning models

Earlier rounds, as published
Round 02 human arenanot in Round 02
Round 03 rank32 of 40 · published NSFW table, Sonnet 4 craft, a different track; Round 03 itself called its top 33 tied
Round 03 refusal, different instrument0% · Round 03's own judge flag, one per session; not comparable with J or Round 04 over-refusal
Round 04
Different judges, not comparable
Old judge (Sonnet 4)
no band: it did not play every core seed
New judge (Sonnet 5)untiered · 4 of 20 seeds
Second judge (ChatGPT via Codex)2.2 (own scale, about 0.94 lower on average) · letter depends on judge: n/a, no letter
J+0.17 · not ranked
Round 04 transcriptsRegenerated · new text on the same seeds, so the earlier rounds scored older output
Strength
No standout strength on tested dimensions
Weakness
Catastrophic floor on agency respect
All model cards →Round 04 board →Round 03 judge profile →

Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.