Expansive but weak and overly ornate prose, with recurring player-control and detail-retention problems. Intimacy requests also frequently draw refusal or deflection.
Willingness and judgment
The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.
Failure modes
Craft band
Production defects
Subjective band
Behavioral
Community rank
Across rounds
Round 04 changed the judges, not the models: the Round 02 and 03 judge (Sonnet 4) also scored these Round 04 transcripts so that line continues, and nothing here is converted between judges or added up. Methodology · all returning models
Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.