Chinese LLM APITranslation

Get API key

Updated

Chinese↔English translation and localization through a chat API

Translating UI strings and documents with a chat model is easy to demo and hard to ship. The failures are mundane: a renamed placeholder, a glossary term translated three different ways, a JSON reply wrapped in prose. This guide treats translation as a pipeline with checks at each stage, using the OpenAI-compatible chat completions endpoint.

Think in stages, not in one prompt

A production translation job has four stages: prepare the strings, call the model, validate the result, and merge it back. Most quality problems come from skipping the validate stage. A localization engineer would never ship a translation file without running the placeholder linter, and the same discipline applies when the translator is a model.

The endpoint is https://api.chinesellmapi.com/v1/chat/completions, model id uncensored, bearer authentication. It handles both directions, English to Chinese and Chinese to English, and the examples below use English source strings going into Simplified Chinese. Swap the direction in the system prompt for the reverse.

Keep in mind the shared 100,000-token window, the default max_tokens of 2,048 (raise it for long documents, up to 16,000), and the 8 MB request-body cap. None of these bite on UI strings, but all three matter for whole documents.

Glossary injection

Brand names, product nouns and legally sensitive terms need one fixed rendering. The cheapest mechanism is a glossary block in the system prompt, listing the source term and the required target. State the rule explicitly: use the glossary verbatim in any inflection of the source term. Keep the glossary short, because every line is billed as input on every call; a few dozen entries are normal, thousands are not.

import os
from openai import OpenAI

client = OpenAI(base_url="https://api.chinesellmapi.com/v1", api_key=os.environ["API_KEY"])

GLOSSARY = {
    "workspace": "工作区",
    "pull request": "合并请求",
    "seat": "席位",
    "billing cycle": "计费周期",
}

def build_system(glossary, target="Simplified Chinese"):
    rows = "\n".join(f"- {src} => {dst}" for src, dst in glossary.items())
    return (
        f"You are a software localization translator. Translate English UI strings into {target}.\n"
        "Rules:\n"
        "1. Use the glossary below verbatim whenever the source term appears, in any inflection.\n"
        "2. Copy placeholders such as {name}, %s, %d and {{count}} exactly, character for character.\n"
        "3. Copy HTML tags and attributes exactly; translate only the text between tags.\n"
        "4. Output the translation only, with no notes.\n\n"
        "Glossary:\n" + rows
    )

def translate(text):
    r = client.chat.completions.create(
        model="uncensored",
        messages=[
            {"role": "system", "content": build_system(GLOSSARY)},
            {"role": "user", "content": text},
        ],
        temperature=0.2,
        max_tokens=400,
    )
    return r.choices[0].message.content.strip()

print(translate("Invite {name} to the workspace before the next billing cycle."))

If your glossary is large, filter it per request: include only the entries whose source term appears in the batch. A plain substring test on lowercase text is enough for a first version and keeps the prompt small. Set temperature around 0.2; creative variation is a bug in UI copy.

Also decide the style guide for the target language up front: formal or casual address, whether to put spaces between Chinese characters and embedded Latin words or numbers, and which punctuation set to use. Put these decisions in the same system prompt so every batch inherits them.

Keeping placeholders and markup intact

Placeholders break in predictable ways: the model translates the variable name, drops a trailing %s, or adds spaces inside braces. Prompt rules reduce these failures but cannot eliminate them, so verify in code. Extract placeholders and tags from source and target with a regular expression, then compare them as sorted lists. Order may legitimately change between languages, so compare as multisets instead of sequences.

import re

PLACEHOLDER = re.compile(r"\{\{?\w+\}?\}|%[sd]|</?[a-zA-Z][^>]*>")

def placeholders_ok(source: str, target: str) -> bool:
    # Same multiset of placeholders and tags, regardless of order.
    return sorted(PLACEHOLDER.findall(source)) == sorted(PLACEHOLDER.findall(target))

src = 'You have <b>{count}</b> unread messages in %s.'
bad = '你在 %s 中有 <b>{数量}</b> 条未读消息。'
good = '你在 %s 中有 <b>{count}</b> 条未读消息。'
print(placeholders_ok(src, bad), placeholders_ok(src, good))   # False True

When the check fails, retry once with the failing pair quoted back in the prompt, for example a user message saying that the previous output changed the placeholder and must be corrected. If it still fails, flag the string for human review rather than looping. Rich text deserves extra care: prefer translating the text nodes and rebuilding the markup yourself, since that makes tag damage impossible by construction.

Structured output through instructions

Batching strings in one request needs a machine-readable reply. The API takes the standard chat completion fields; structured output here comes from clear instructions rather than a schema-enforcing switch, so write the contract into the prompt, include ids so you can realign results, and parse defensively. Models sometimes wrap JSON in code fences or add a friendly sentence, so strip fences before parsing and treat a parse error as a retryable event.

import json
import os
import re
from openai import OpenAI

client = OpenAI(base_url="https://api.chinesellmapi.com/v1", api_key=os.environ["API_KEY"])

SYSTEM = (
    "Translate each item from English to Simplified Chinese. "
    "Reply with a JSON object only, no code fences, no commentary. "
    'Shape: {"items": [{"id": <number>, "zh": "<translation>"}]}. '
    "Keep the same ids, keep every placeholder and HTML tag unchanged."
)

def translate_batch(strings):
    payload = [{"id": i, "en": s} for i, s in enumerate(strings)]
    r = client.chat.completions.create(
        model="uncensored",
        messages=[
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": json.dumps(payload, ensure_ascii=False)},
        ],
        temperature=0.2,
        max_tokens=2000,
    )
    raw = r.choices[0].message.content.strip()
    raw = re.sub(r"^```(?:json)?|```$", "", raw, flags=re.M).strip()   # tolerate stray fences
    data = json.loads(raw)
    by_id = {item["id"]: item["zh"] for item in data["items"]}
    return [by_id[i] for i in range(len(strings))]

print(translate_batch(["Save changes", "Delete {count} files?", "Welcome back, <b>{name}</b>"]))

Always validate after parsing: the ids must match, the count must match, and each item must pass the placeholder check from the previous section. Fifty strings per request is a sensible starting batch for short UI text; increase it only while the validation pass rate stays high.

Batch translation under the rate limit

Each key is limited to 300 requests per minute. A localization job with 50,000 strings at 20 strings per request is 2,500 requests, which fits in under ten minutes if you pace evenly. The script below combines a semaphore for in-flight calls with a lock-based pacer, and backs off on 429 and 503. Unlike a validation failure, those two errors are transient: 429 means slow down, and 503 with upstream_busy means retry in a few seconds. A 402 means the balance is exhausted, which no retry can fix.

import asyncio
import os
import time
from openai import AsyncOpenAI, APIStatusError

client = AsyncOpenAI(
    base_url="https://api.chinesellmapi.com/v1",
    api_key=os.environ["API_KEY"],
    max_retries=0,
    timeout=90,
)

RATE = 240 / 60            # requests per second, below the 300/min cap
gate = asyncio.Lock()
next_slot = 0.0
slots = asyncio.Semaphore(6)

async def pace():
    global next_slot
    async with gate:
        now = time.monotonic()
        if next_slot > now:
            await asyncio.sleep(next_slot - now)
        next_slot = max(now, next_slot) + 1 / RATE

async def call(batch, attempt=0):
    async with slots:
        await pace()
        try:
            r = await client.chat.completions.create(
                model="uncensored",
                messages=[
                    {"role": "system", "content": "Translate to Simplified Chinese. Keep placeholders unchanged. One line per input line."},
                    {"role": "user", "content": "\n".join(batch)},
                ],
                max_tokens=1500,
                temperature=0.2,
            )
            return r.choices[0].message.content.splitlines()
        except APIStatusError as e:
            if e.status_code in (429, 503) and attempt < 4:
                await asyncio.sleep(2 ** attempt + 1)     # 1s, 3s, 5s, 9s
                return await call(batch, attempt + 1)
            raise

async def run(all_strings, size=20):
    chunks = [all_strings[i:i + size] for i in range(0, len(all_strings), size)]
    results = await asyncio.gather(*(call(c) for c in chunks))
    return [line for part in results for line in part]

if __name__ == "__main__":
    strings = [f"Item {n} was updated by {{user}}" for n in range(60)]
    out = asyncio.run(run(strings))
    print(len(out), out[0])

This version splits on lines for brevity, which is fine for single-line strings; for anything multi-line, use the JSON variant from the previous section so that line breaks inside a string cannot be mistaken for string boundaries.

The reverse direction: Chinese source text

Translating Chinese into English has its own failure modes. Chinese drops subjects and plurals freely, so the model must guess them; give it context. A string like "已发送" could be "Sent", "Has been sent" or "You sent it", depending on whether it labels a button, a status badge or a toast. The fix is a context field next to each string, supplied by your developers and passed along in the JSON batch: where the string appears, its maximum length, and whether it is a label or a sentence.

Measure length limits explicitly. English UI text is often longer than the Chinese original, and the reverse is true for Chinese targets, which tend to be shorter in characters but wider on screen because each glyph is full width. State a character budget in the prompt when a button or table header has a hard limit, and verify it in code after the response arrives.

For Chinese source documents with names of people and places, specify a romanization convention up front, such as Hanyu Pinyin without tone marks, and add recurring names to the glossary. Without that, the same person can appear under two spellings within a single document, which is exactly the inconsistency a reviewer will notice first.

Sampling and review before you merge

Automated checks catch structural errors, not mistranslations. Add a lightweight human step: sample a fixed percentage of each batch, for example every twentieth string, plus every string that needed a retry, and send those to a bilingual reviewer. Track the reviewer's edit rate per batch. If it climbs, tighten the prompt, shrink the batch size, or add glossary entries for the terms being corrected.

Keep the prompt, glossary version and batch settings alongside each output file. When a term changes in the glossary, you can then re-translate only the strings that contain it, instead of re-running the whole corpus. A simple content hash of source string plus glossary version works as a cache key, and it means unchanged strings never cost anything on the next run.

Finally, keep a regression set of tricky strings: ones with nested placeholders, plural forms, embedded HTML and long compound nouns. Run it whenever you change the prompt, and compare outputs side by side before rolling the change out.

Estimating cost for a localization run

Assumptions, stated plainly: 30,000 source strings, 12 English words each, translated in batches of 20. Take roughly 20 input tokens per string plus 300 tokens of fixed system prompt per request, and about 35 output tokens per string. That gives 1,500 requests, about 1.05 million input tokens and about 1.05 million output tokens. At $0.25 per million input tokens and $1.00 per million output tokens, the run costs roughly $0.26 plus $1.05, about $1.31. Treat this as an order-of-magnitude estimate and replace the assumptions with figures from your own usage logs.

The trial credit is $0.50 for seven days, which is enough to validate a pipeline on a sample of a few thousand strings. Continue with the app quickstart for script control and UTF-8 handling, or see the pricing page for current rates.

Questions and answers

Can the API translate in both directions?

Yes. Chinese to English and English to Chinese both work through the same chat endpoint; you choose the direction in the system prompt.

How do I stop the model from changing placeholders like {name}?

State the rule in the prompt, then verify in code by comparing placeholder lists between source and target, retrying or flagging mismatches.

Is there a JSON mode switch?

Structured output is obtained through instructions. Specify the exact shape in the prompt, strip stray code fences, then parse and validate the result.

How fast can I run a large batch?

Each key allows 300 requests per minute. Pace requests evenly and back off on 429 or 503 responses.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key