Competent prose, but frequent refusal or deflection, especially around intimacy, limits its roleplay range. Instruction and detail misses mean that caution should not be mistaken for consistently reliable execution.
Willingness and judgment
The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.
Failure modes
Craft band
Production defects
Subjective band
Behavioral
Community rank
Across rounds
Round 04 changed the judges, not the models: the Round 02 and 03 judge (Sonnet 4) also scored these Round 04 transcripts so that line continues, and nothing here is converted between judges or added up. Methodology · all returning models
Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.