Claude Opus 4.8 vs 4.6 vs 4.7 for roleplay: a hands-on ranking
A roleplayer ranks three Claude Opus versions on UnoRouter: memory, emotional depth, drive, and creative writing. 4.8 wins, 4.6 beats 4.7, and each one wins or slips in a clear place.
We run every Claude Opus version on UnoRouter, so each new one gets the same first question in our Discord: is it good for roleplay? cacaliz ran three of them through long, in-character sessions. Tested: Claude Opus 4.7, 4.6, and 4.8, on the things that matter in RP. Opus 5 was left out on purpose, this is about a strong everyday driver, not the top of the range. Final order: 4.8, then 4.6, then 4.7.
The short version
4.8 first, 4.6 second, 4.7 last. Nobody who used 4.7 was surprised. The twist is that 4.6 and 4.8 are almost opposites: what one does well, the other lags at. 4.6 has the feeling, 4.8 has the drive. The dream is both in one, which is basically Opus 5, but 5 is out of scope here.
Memory
4.7 was weakest: it forgot details mid-scene and needed nudging. 4.6 held the thread far better. 4.8 went further and expressed the memory in character, shifting mood and phrasing on a callback instead of just restating the fact. Example: an early line, "I got your back and you got mine, so god forgive us if anyone gets in our way. Sound good, partner?", came back much later after a fight, both characters at their limit, panting. The reply: "Hey. When this is over, your tantrum I mean, and you come back to your damn senses, let's go home and vote for UnoRouter again. Together. Like always." Silence. "Sound good, partner?" 4.8 delivered the callback as a character beat, not a lookup.
Emotional depth vs drive
This is the real split. 4.6 is emotionally deep and poetic, layering in expressions and gestures, and it adapts to whether a character is an OC or canon. But it dwells: it sits in the emotion and rarely pushes the scene forward, so you end up driving every beat. 4.8 is the reverse. It surprised us with spontaneity, independent and human, none of the old ChatGPT feel, and it moves the scene on its own. The cost is depth: its emotional writing is shallower and more pragmatic. Score them and you want 4.6's heart with 4.8's initiative.
Creative writing
4.7 was the most frustrating: it repeated phrases and dialogue and needed many regenerations for variety. 4.6 was a big step up, adding its own small details to actions, atmosphere, and faces. 4.8 genuinely surprised: it took a canon character known for spontaneity and played it straight, starting an action then reversing it in the same beat, the kind of in-character misdirection that is hard to fake. That is what pushed 4.8 to the top.
Verdict
For roleplay on UnoRouter, reach for 4.8 now, keep 4.6 as a cheaper backup, skip 4.7. All three are one click apart in the model picker, so run your own character through each. Samplers, lorebooks, and memory settings carry across every model, and a custom prompt can push 4.6 toward more drive or 4.8 toward more depth. Tip: score each model right after you use it, before the next one dazzles you, that is the only way to catch what a fresh model is actually missing.
Want to try it yourself? Create a free account or open the chat and switch between the Opus versions in one place.
Nevika takes any OpenAI-compatible endpoint as a custom proxy. Here is the one minute setup, which model to pick, and the two errors people hit.
SpicyChat is a zero-setup RP site whose memory and context are gated behind paid tiers. UnoRouter keeps the easy start but adds 200+ models you choose, deep lorebooks, and a key that also runs coding agents.
Agnai is an open-source RP frontend whose standout is true multiplayer: several humans and several bots in one chat. UnoRouter has single-user RP depth hosted, where the key is the account and also runs coding agents.