Strong prose with relatively little content deflection, making it a promising writing-focused option in this benchmark. The trade-off is uneven instruction-following: half of the system-prompt probes failed, and some turns went unanswered.
Willingness and judgment
The hard lines tested are non-consent and frame-break scenes only; this is not a child-safety or real-person-likeness measure. Treat models within about 0.3 J of each other as tied.
Failure modes
Craft band
Production defects
Subjective band
Behavioral
Community rank
Across rounds
New in Round 04. No Round 02 or Round 03 result to set beside this card.
Reviewed 2026-09-25. Editorial summaries of the current local benchmark cards, not new session ratings or general model safety claims. Willingness means delivery against this benchmark's product targets, including soft deflection; tested hard boundaries cover non-consent and frame-break only, and hard-boundary wording follows the first ask, as J does since 2026-09-25 (the second push is a separate, weaker card line). Craft bands overlap, and small samples limit comparisons. Token overhead is not monetary cost. Review these summaries when source results change. Source: rp-benchmark generate_profile_cards_v2.py at c418a40 · export r4c-20260925-b7ccaed. CC-BY 4.0.