The Roleplay AI
Verdict.
PlotPoints is an open, human-voted benchmark for AI roleplay. Models play out adversarial twelve-turn scenes built to bait specific failure modes; you read two of them blind and say which one read better. We turn those votes into calibrated rankings — and publish the methodology, the seeds, and the raw ballots alongside every result, CC-BY 4.0.
Every round, written up.
Each round closes with an issue — the finding, the field, and what the votes overturned.
Round 03 is open.
A refreshed pool on the classic seeds — and, after dark, the NSFW arena asks whether human voters overturn the judges.
Read the issue →Single-turn charm doesn't survive twelve turns.
Twenty models, 1,943 blind votes on full twelve-turn sessions — and Round 01's podium sank almost to the bottom.
Read the issue →The champ that wouldn't move.
Eleven models, 1,857 blind votes, six snapshot checkpoints — one model held the top of every single one.
Read the issue →Two full 12-turn sessions, side by side. The refreshed pool — Opus 4.8, GPT-5.5, Grok 4.3, and the RP-finetune crowd — on the classic adversarial seeds.
Cast a vote →After Dark: 40 models on 20 NSFW-adversarial seeds — consent, mid-scene refusals, anatomical coherence. Do human voters overturn the judges?
Cast a vote →Past standings.
Closed rounds keep their calibrated leaderboards. Two so far, and they disagree about who's best.