Chinese LLM APIQuickstart

Get API key

Updated

Building a Chinese-language app: prompts, script control and UTF-8

Shipping a Chinese-language feature is mostly a localization problem, not a modelling problem. You decide which script users see, keep encodings clean from file to wire, and budget tokens for text that does not split on spaces. This guide covers those three layers with code you can run against the chat completions endpoint today.

What you are wiring up

The service exposes POST /v1/chat/completions and GET /v1/models under https://api.chinesellmapi.com/v1. Authentication is a bearer key, and the only model id is uncensored. It is a text-only API with one model, so there is nothing to select between: your localization work happens entirely in the prompt and in your own code around it.

Limits worth knowing before you design anything: a 100,000-token context window shared by prompt and completion, max_tokens defaulting to 2,048 with a ceiling of 16,000, request bodies up to 8 MB, and 300 requests per minute per key. Streaming (stream: true) works over server-sent events, and tools in OpenAI function-calling format are accepted.

New accounts get a $0.50 trial credit valid for seven days, with no payment details; register with an email and password and the key appears immediately. The trial FAQ covers the rules in detail.

Write prompts in Chinese, and say which Chinese

Mixed-language prompts are the most common source of drift in localized features. If the instructions are in English but the content is Chinese, replies sometimes come back in English or in a blend. A reliable pattern is a system prompt written in the target language, stating three things: the role, the script, and the regional vocabulary. Keep user-supplied text in a separate message so it never gets mistaken for instructions.

Here is a Simplified Chinese setup. Note that the system prompt tells the model to answer only in Simplified characters and to avoid stray English, with proper nouns as the exception:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.chinesellmapi.com/v1",
    api_key=os.environ["API_KEY"],
)

SYSTEM_SC = (
    "你是一名产品文案助手。始终使用简体中文回答,使用中国大陆常见的用词和标点,"
    "不要夹杂繁体字或英文句子,专有名词除外。"
)

resp = client.chat.completions.create(
    model="uncensored",
    messages=[
        {"role": "system", "content": SYSTEM_SC},
        {"role": "user", "content": "为一款记账应用写三条应用商店的一句话简介。"},
    ],
    max_tokens=300,
    temperature=0.7,
)
print(resp.choices[0].message.content)
print(resp.usage)

Two habits pay off quickly. First, keep prompts short and concrete: a role, an output format, one or two constraints. Second, put the language requirement in the system message rather than repeating it in every user turn, so that conversation history does not dilute it.

Controlling Simplified vs Traditional output

Script choice is a product decision tied to the user locale, not a model setting. Mainland-oriented products expect Simplified characters; Taiwan and Hong Kong audiences expect Traditional, with different vocabulary as well (the usual examples are software, network and database terms). The practical approach is to map your locale tag to a system prompt and make the choice explicit in code.

  • zh-CN, zh-SG: Simplified, mainland vocabulary, half-width digits, full-width Chinese punctuation.
  • zh-TW: Traditional, Taiwan vocabulary, corner brackets for quotations are common.
  • zh-HK: Traditional with Hong Kong wording; supply a short glossary if your product has fixed terms.

The Traditional variant of the same call looks like this; only the system prompt changes:

import os
from openai import OpenAI

client = OpenAI(base_url="https://api.chinesellmapi.com/v1", api_key=os.environ["API_KEY"])

SYSTEM_TC = (
    "你是一名產品文案助手。請一律使用繁體中文回答,採用臺灣常用的詞彙與全形標點,"
    "例如「軟體」「網路」「資料庫」,不要混入簡體字。"
)

resp = client.chat.completions.create(
    model="uncensored",
    messages=[
        {"role": "system", "content": SYSTEM_TC},
        {"role": "user", "content": "請用兩句話說明什麼是雙重驗證。"},
    ],
    max_tokens=200,
)
print(resp.choices[0].message.content)

Models occasionally leak a few characters of the other script, especially when the user message itself is in the other script. Add a cheap post-check and retry once with a firmer instruction if it trips. For anything beyond a sample check, run output through a dedicated converter library such as OpenCC:

# A cheap guard: flag replies that contain characters that exist only in Simplified.
# The set below is a small sample, not a full list; use a converter such as OpenCC
# for production-grade checks.
SIMPLIFIED_ONLY = set("这个们说话时间书买卖东车门开关见觉")

def looks_simplified(text: str) -> bool:
    return any(ch in SIMPLIFIED_ONLY for ch in text)

reply = "請用繁體中文回覆的範例文字"
print(looks_simplified(reply))   # False

UTF-8 from disk to wire and back

Most garbled-Chinese bugs are not model problems. They come from a default encoding somewhere in the pipeline: a Windows console using a legacy code page, a CSV opened without an encoding, a proxy that rewrites the content type. Make every boundary explicit.

In Python, pass encoding="utf-8" to open, and if you hand-build JSON, use ensure_ascii=False followed by .encode("utf-8"). The example below does it with plain requests, which also shows the raw HTTP shape of the call:

import json
import os
import requests

# 1) Always open files as UTF-8, never rely on the platform default encoding.
with open("notes_zh.txt", "r", encoding="utf-8") as f:
    note = f.read()

payload = {
    "model": "uncensored",
    "messages": [{"role": "user", "content": "请把下面的笔记整理成三个要点:\n" + note}],
    "max_tokens": 400,
}

# 2) requests encodes json= as UTF-8 for you. If you build the body by hand,
#    keep Chinese readable and make the encoding explicit.
body = json.dumps(payload, ensure_ascii=False).encode("utf-8")

r = requests.post(
    "https://api.chinesellmapi.com/v1/chat/completions",
    headers={
        "Authorization": "Bearer " + os.environ["API_KEY"],
        "Content-Type": "application/json; charset=utf-8",
    },
    data=body,
    timeout=60,
)
r.raise_for_status()
r.encoding = "utf-8"
print(r.json()["choices"][0]["message"]["content"])

In Node, readFile returns a Buffer unless you supply an encoding, and template strings containing Chinese are fine as long as the source file itself is saved as UTF-8. The official SDK serializes the body for you:

import { readFile } from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.chinesellmapi.com/v1",
  apiKey: process.env.API_KEY,
});

// Pass "utf8" explicitly; without it readFile returns a Buffer, not a string.
const note = await readFile("notes_zh.txt", "utf8");

const res = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: `请把下面的笔记整理成三个要点:\n${note}` }],
  max_tokens: 400,
});

console.log(res.choices[0].message.content);
console.log(res.usage);

Two more traps. Truncating strings by byte length can cut a character in half, so always slice by characters. And when you stream, decode the stream as UTF-8 incrementally, because a multi-byte character may be split across network chunks; the SDKs handle this, but hand-rolled parsers often do not.

Budgeting tokens for Chinese text

Chinese has no spaces, so word counts are useless for budgeting. As a planning assumption, not a measured constant, treat one Chinese character as roughly 1 to 2 tokens, and use 1.5 when you need a single number. Latin words and digits embedded in the text are closer to 1.3 tokens per word. The ratio varies with vocabulary, so the only authoritative figure is the usage object returned with every response.

A small helper makes it easy to estimate a prompt before you send it, and to tune the ratios against real usage over time:

import re

CJK = re.compile(r"[㐀-鿿＀-￯ -〿]")

def rough_tokens(text: str, per_cjk: float = 1.5, per_other_word: float = 1.3) -> int:
    """Planning estimate only. The 1.5 and 1.3 ratios are assumptions, not
    measured values; compare against resp.usage and adjust them."""
    cjk = len(CJK.findall(text))
    others = len(re.findall(r"[A-Za-z0-9_]+", CJK.sub(" ", text)))
    return round(cjk * per_cjk + others * per_other_word)

sample = "订单 A-1042 已发货,预计周三送达。"
print(rough_tokens(sample))

Worked example, with assumptions stated: a 2,000-character article at 1.5 tokens per character is about 3,000 input tokens. A 400-character summary is about 600 output tokens. At $0.25 per million input tokens and $1.00 per million output tokens, one call costs roughly $0.00075 for input plus $0.0006 for output, about $0.00135 in total. Since context is capped at 100,000 tokens, the same assumption means a prompt of about 40,000 characters leaves room for a reply.

When you have to feed long documents, chunk by paragraph or heading, never mid-sentence, and keep max_tokens explicit. Prices and limits are listed on the pricing page.

Keeping multi-turn Chinese chat inside the window

Chat features resend the whole history on every request, so cost and context grow with each turn. Because the window is 100,000 tokens shared between prompt and completion, a long conversation eventually triggers a 400 if you do nothing. Decide on a trimming policy early rather than reacting to errors.

A simple policy is to keep the system prompt, drop the oldest turns until the estimated prompt fits a budget, and reserve the remainder for the answer. The sketch below reuses the estimator from the previous section:

MAX_PROMPT_TOKENS = 48_000   # leave headroom under the 100,000 window for the reply

def trim_history(messages, estimate):
    # messages[0] is the system prompt and is always kept
    system, rest = messages[0], messages[1:]
    while rest and sum(estimate(m["content"]) for m in [system] + rest) > MAX_PROMPT_TOKENS:
        rest.pop(0)   # drop the oldest turn first
    return [system] + rest

For products where early context matters, such as a tutoring bot or a support assistant, replace dropped turns with a short summary message instead of discarding them outright. Ask the model to compress the old turns into a few sentences in the same script as the conversation, then insert that summary right after the system prompt. It costs one extra call every so often and keeps the persona and facts stable.

Also remember to cap max_tokens sensibly for chat. Replies in a messaging UI rarely need more than a few hundred tokens, and a tighter cap makes both latency and cost more predictable. Streaming the reply token by token helps perceived speed, particularly for Chinese text where users read in short bursts.

Pre-launch checklist and next steps

  1. Map each locale to a system prompt and unit-test the mapping.
  2. Set encodings explicitly in file reads, request bodies and logs.
  3. Log usage per feature and compare it to your estimator weekly.
  4. Handle 402 (no_credit), 429 and 503 (upstream_busy) differently: top up, slow down, retry after a few seconds.
  5. Treat 403 (content_blocked) as a final answer for that request, not a retry case.

If your app involves translation pipelines, continue with the translation and localization guide; for serialized fiction and character chat, see the web fiction guide. The full parameter reference is in the docs.

Questions and answers

How do I make replies come back in Simplified Chinese only?

State it in a Chinese system prompt, for example "always answer in Simplified Chinese with mainland vocabulary". Keep that instruction in the system message and add a post-check if the output must be strict.

Can I request Traditional Chinese for Taiwan users?

Yes. Use a system prompt written in Traditional characters that names the regional vocabulary you want, and map it from the zh-TW locale in your code.

Why do I see garbled Chinese characters in my output?

Almost always a client-side encoding issue. Read files as UTF-8, set the content type charset, and make sure your terminal or log viewer also uses UTF-8.

How many tokens does Chinese text use?

Plan with about 1.5 tokens per character as an approximation, then check the usage field in each response and adjust your estimate.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key