Updated
Chinese web fiction (网文) and roleplay apps: context, characters and cost
Serialized Chinese web fiction and in-character chat are two of the most demanding text workloads: thousands of characters per chapter, a cast that must stay consistent for months, and readers who notice every slip. This guide shows how to structure prompts, manage the 100,000-token window, and estimate cost from stated assumptions.
What these workloads look like
Web fiction serials are written in chapters of 2,000 to 4,000 characters, often published daily, with a plot that spans hundreds of chapters. Roleplay apps are the conversational cousin: a persona, a scenario and a long history. Both need the same three things from the model API: enough context to remember, enough output budget to write at length, and a style that does not drift.
The service offers an OpenAI-compatible chat endpoint at https://api.chinesellmapi.com/v1 with one text model, id uncensored. The context window is 100,000 tokens shared between prompt and completion, max_tokens defaults to 2,048 and can go to 16,000, and standard sampling fields such as temperature, top_p and stop are passed through. Streaming is supported, which suits chat. Basic setup and encoding notes live in the app quickstart.
Content rules, stated up front
This is an adults-only service. Lawful adult fiction, mature themes and controversial subjects are not refused as such, and the service is intended for users aged 18 or over. Sexual content involving minors is always blocked and returns a 403 content_blocked error, including when it is framed as fiction or roleplay. There is no setting that changes this.
Build for it from the start. State in your character sheets that every character is an adult and give explicit ages. Add an age gate to your own product. Treat a 403 as a final answer for that request: show a neutral message and do not retry with rephrased prompts. These steps are good product hygiene whatever the platform rules say.
Character sheets and style blocks
Consistency comes from putting stable facts in one place and repeating them on every call. A character sheet is a compact record of name, age, role, speech habits, relationships and hard constraints, including secrets the character must not reveal yet. A separate style block fixes the point of view, pacing, dialogue ratio, chapter length and ending habit. Both go in the system message so they sit at the front of every request.
CHARACTER_SHEET = """\
【人物卡】
姓名:沈清禾(女,28岁)
身份:江城古籍修复师,性格沉静,嘴硬心软
口头禅:“先别急,东西不会跑。”
说话方式:短句,少用感叹号,偶尔引用旧书里的句子
关系:与顾远舟(男,31岁,旧书商)是多年好友,彼此有未说出口的好感
禁忌:不提及她离开出版社的真正原因(第40章才揭晓)
"""
STYLE = """\
【文风】第三人称限知视角,贴近沈清禾;节奏舒缓,多写物件与天气的细节;
对话占比约四成;每章 2500 到 3000 字;结尾留一个小悬念。"""
Write these in Chinese. Speech habits expressed in Chinese examples, such as a catchphrase or sentence length, transfer much more reliably than an English description of them. Keep each sheet under a few hundred tokens, and give every recurring character one; a cast of ten is still only a few thousand tokens. When a character's situation changes, update the sheet itself rather than relying on the history to carry the change.
Long context for serialized chapters
The temptation is to paste every previous chapter into the prompt. Do the arithmetic first. With a 100,000-token window and chapters of about 4,500 tokens, the whole window holds about fourteen chapters with no room left for output, and a serial runs for hundreds. The workable pattern is layered memory:
- The character sheet and style block, always present.
- A rolling story summary, a few hundred characters per chapter, merged into one running recap that you refresh every few chapters.
- The full text of the previous chapter, so the prose connects at the seam.
- The outline for the chapter being written.
After each chapter, ask the model for a short summary that preserves relationship changes and open plot threads, and append it to the recap. The code below writes a chapter and produces the next summary. It also returns usage so you can log real token counts:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.chinesellmapi.com/v1", api_key=os.environ["API_KEY"])
def write_chapter(sheet, style, story_so_far, last_chapter, outline):
messages = [
{"role": "system", "content": "你是一位连载网络小说的作者,所有人物均为成年人。严格遵守人物卡和文风设定。\n" + sheet + "\n" + style},
{"role": "user", "content": (
"【前情提要】\n" + story_so_far +
"\n\n【上一章全文】\n" + last_chapter +
"\n\n【本章大纲】\n" + outline +
"\n\n请直接写出本章正文,不要写标题以外的说明。")},
]
r = client.chat.completions.create(
model="uncensored",
messages=messages,
max_tokens=5000, # ~3,000 characters needs roughly 4,500 tokens under our assumption
temperature=0.9,
top_p=0.95,
)
return r.choices[0].message.content, r.usage
def summarize(chapter_text):
r = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "用 200 字以内概括下面这一章的剧情,保留人物关系的变化和未解的伏笔:\n" + chapter_text}],
max_tokens=400,
temperature=0.3,
)
return r.choices[0].message.content
Roleplay: persona, turn shape and stop sequences
Roleplay needs the same sheet-and-style discipline plus a few chat-specific controls. Put the persona in the system message, state who plays whom, fix the reply length range, and tell the model not to write the user's character's actions. Use stop to cut the output if the model starts writing the user's next line, and stream the reply so the text appears as it is generated.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.chinesellmapi.com/v1", api_key=os.environ["API_KEY"])
PERSONA = (
"你扮演顾远舟,31岁的旧书商,说话随和,爱开玩笑,但在重要的事上很认真。"
"用户扮演沈清禾。所有角色均为成年人。保持第一人称,每次回复 80 到 200 字,"
"用(括号)写简短的动作描写,不要替用户的角色做决定。"
)
history = [{"role": "system", "content": PERSONA}]
def turn(user_text):
history.append({"role": "user", "content": user_text})
stream = client.chat.completions.create(
model="uncensored",
messages=history,
max_tokens=400,
temperature=1.0,
stop=["\n用户:"], # keep the model from writing the user's next line
stream=True,
)
reply = ""
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
piece = chunk.choices[0].delta.content
reply += piece
print(piece, end="", flush=True)
print()
history.append({"role": "assistant", "content": reply})
turn("(推开书店的门)这场雨下得真不是时候。")
History grows by one user message and one assistant message per turn. Once it nears your budget, summarize the oldest part into a single message and keep the most recent turns verbatim. Temperature around 0.9 to 1.0 gives lively dialogue for characters; drop it toward 0.5 when you need the persona to stay strictly on script.
Style control that survives a long serial
Style drift is the failure readers complain about most: the narrator slowly becomes chattier, the dialogue turns formal, the pacing speeds up. Fight it with specifics. Instead of "write in an elegant style", describe measurable habits: sentence length, how often to use idioms, how much weather and object detail to include, whether chapters end on a hook or a quiet beat. Two or three short sample paragraphs in the target voice, placed in the style block, are worth more than a page of adjectives.
Different genres need different dials. A cultivation or fantasy serial benefits from a glossary of invented terms, such as realms, sects and techniques, so the same name is never spelled two ways. A modern romance leans on dialogue rhythm and small physical details. A mystery needs a clue ledger: a list of facts the reader has seen, so the model does not contradict them. Keep each of these lists in the system prompt and update it as the story moves.
Finally, keep temperature and top_p fixed within a serial. If you change sampling between chapters, voice changes follow. Raise temperature only for brainstorming outlines, not for the prose itself.
Cost math with stated assumptions
All figures use the published rates of $0.25 per million input tokens and $1.00 per million output tokens, and an assumed 1.5 tokens per Chinese character. Replace the assumptions with numbers from your own usage logs.
| Scenario | Assumed input tokens | Assumed output tokens | Cost per call |
|---|---|---|---|
| Serial chapter (3,000 chars out) | 8,000 | 4,500 | $0.0020 + $0.0045 = $0.0065 |
| Roleplay turn, 20,000-token history | 20,000 | 300 | $0.0050 + $0.0003 = $0.0053 |
| Roleplay turn, full 60,000-token history | 60,000 | 500 | $0.0150 + $0.0005 = $0.0155 |
In the chapter row the 8,000 input tokens break down as a 1,200-token character sheet and style block, a 2,000-token recap, a 4,500-token previous chapter and a 300-token outline. A hundred such chapters would cost about $0.65. A hundred roleplay turns at the 20,000-token history size cost about $0.53. The takeaway is that history length drives chat cost, so trimming and summarizing pay for themselves quickly.
The $0.50 trial credit, valid for seven days, covers roughly 76 chapters under the first row's assumptions. Check pricing for current rates, and see the trial FAQ for account details.
Operational notes for production
Chapter generation is a long request, so use generous timeouts or stream and assemble the text yourself. Handle 429 with a short backoff, since each key allows 300 requests per minute, and treat 503 with upstream_busy as a retry-in-a-few-seconds event. A 402 means the prepaid balance is used up or the trial has expired, which should alert you rather than retry.
Save every generated chapter and its summary to your own database before generating the next one. If a call fails halfway through a night's batch, you can resume from the last stored chapter, and you never pay twice for the same text. Store the prompt version and the usage figures next to each chapter so cost per chapter is a query rather than a guess.
For roleplay products, keep a per-session budget. Cap the number of turns or the total tokens a single session can consume, and show users a clear message when the cap is reached. Because prompts include history, a very long session costs far more per turn than a fresh one, which is why summarizing old turns matters beyond just fitting the window. Prompts are not used for training, which you can state in your product's privacy notes alongside your own storage policy.
Questions and answers
How long can a chapter be in one request?
Output is capped at 16,000 tokens per request and shares the 100,000-token window with the prompt. A 3,000-character chapter fits comfortably, so set max_tokens a little above your estimate.
How do I keep characters consistent across hundreds of chapters?
Keep a character sheet and style block in the system prompt on every call, and carry the plot forward as a rolling summary plus the previous chapter.
Is adult fiction allowed?
Lawful adult fiction between adult characters is not refused, and the service is for users 18 and over. Sexual content involving minors is always blocked with a 403, including in fiction and roleplay.
Can I stream roleplay replies?
Yes. Set stream to true to receive server-sent events, and use stop sequences to prevent the model from writing the user's lines.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.