Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Roleplay

Best AI Models for Roleplay (RP) and Creative Writing

Model rankings updated September 2026 based on real usage data.

Discover the top AI models for roleplay (RP), character chat and creative writing, ranked by real usage data on OpenRouter. These LLMs excel at maintaining consistent personas, rich dialogue and immersive storytelling across long-context sessions.

Whether you're using Janitor AI, SillyTavern or another frontend, or building your own character chatbot or interactive fiction engine, OpenRouter gives you access to the best roleplay models through a single API.

Browse All ModelsCompare Models

LLM Leaderboard for Roleplay Models

1.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
1.3T
27.1%
2.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
412B
8.6%
3.
Favicon for deepseek
Deepseek V3.2
by deepseek
319B
6.6%
4.
Favicon for xiaomi
Mimo V2.5
by xiaomi
295B
6.1%
5.
Favicon for z-ai
GLM 5.3 Flash
by z-ai
270B
5.6%
6.
Favicon for google
Gemini 3 Flash Preview
by google
189B
3.9%
7.
Favicon for deepseek
Deepseek V4 Pro
by deepseek
159B
3.3%
8.
Favicon for google
Gemini 2.5 Flash Lite
by google
146B
3.0%
9.
Favicon for z-ai
GLM 5.2
by z-ai
146B
3.0%
10.
Favicon for unknown
Others
1.57T
32.6%

Top Roleplay Models on OpenRouter

Based on top weekly usage data from millions of users accessing AI models for roleplay through OpenRouter.

Favicon for tencent

Tencent: Hy4 preview

18.2T tokens
Academia (#5)
Finance (#9)
Health (#42)
Legal (#7)
Marketing (#6)

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

by tencent1.05M context$0.834/M input tokens$2.501/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Luna

15.7T tokens
Academia (#4)
Finance (#5)
Health (#5)
Legal (#2)
Marketing (#2)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3 Flash

12.5T tokens
Academia (#2)
Finance (#2)
Health (#4)
Legal (#5)
Marketing (#5)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.31M context$0.075/M input tokens$0.25/M output tokens50% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0731

12.4T tokens
Academia (#1)
Finance (#1)
Health (#1)
Legal (#1)
Marketing (#1)

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

by deepseek1.31M context$0.03/M input tokens$0.07/M output tokens
Favicon for xiaomi

Xiaomi: MiMo-V2.5

6.39T tokens
Academia (#9)
Finance (#14)
Health (#17)
Marketing (#21)
SEO (#21)

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0423

4.78T tokens
Academia (#3)
Finance (#3)
Health (#3)
Legal (#4)
Marketing (#3)

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.06678/M input tokens$0.1336/M output tokens52% off
Favicon for z-ai

Z.ai: GLM 5.3

2.88T tokens
Academia (#13)
Finance (#8)
Health (#20)
Legal (#19)
Marketing (#33)

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

by z-ai1.31M context$0.8727/M input tokens$3.36/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4.1 Flash

2.66T tokens
Academia (#28)
Finance (#4)
Health (#8)
Legal (#50)
Marketing (#37)

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

by deepseek1.05M context$0.15/M input tokens$0.60/M output tokens
Favicon for google

Google: Gemini 3.8 Flash

2.62T tokens
Academia (#11)
Finance (#10)
Health (#11)
Legal (#9)
Marketing (#16)

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

by google1.05M context$0.75/M input tokens$3.75/M output tokens50% off
Favicon for z-ai

Z.ai: GLM 5.2

2.09T tokens
Academia (#6)
Finance (#6)
Health (#18)
Legal (#10)
Marketing (#9)

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

by z-ai1.05M context$0.4802/M input tokens$1.509/M output tokens66% off

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections