Pay per token — no subscription

Chinese LLM API pricing

Pay per token. DeepSeek follows provider peak / off-peak pricing with a transparent 20% XinoAPI markup. No credit purchase fees, no monthly minimum, no annual commitment.

Prices updated: August 2026
Markup
0%
over official provider API price
Credit purchase fee
0%
pay what you see
Free signup credit
$0
≈ 2.5M Flash output tokens at off-peak
Minimum topup
$0
no monthly commitment

One transparent rate card

All prices are USD per 1 million tokens. DeepSeek uses the provider's Beijing-time peak / off-peak schedule; all other model prices below are fixed. Input is split into cache-hit and cache-miss tokens where the provider reports both.

Model Input (per 1M) Output (per 1M) Context
DeepSeek V4-Flash Popular
deepseek-v4-flash · 1M context, 384K output, dual reasoning modes
Off-peak $0.0084 hit · $0.264 miss
Peak $0.0168 hit · $0.528 miss
Off-peak $0.792
Peak $1.584
1M
DeepSeek V4-Pro Flagship
deepseek-v4-pro · 1.6T MoE, SWE-bench 80.6%, approaches Claude Opus
Off-peak $0.0264 hit · $0.792 miss
Peak $0.0528 hit · $1.584 miss
Off-peak $2.376
Peak $4.752
1M
Qwen Turbo
qwen-turbo · fast and cheap, good for high volume
$0.051$0.061$0.204$0.2451M
Qwen Plus New
qwen-plus · Qwen3.7 generation, balanced quality, 1M context
$0.111$0.133$0.667$0.8001M
Qwen Max Flagship
qwen-max · Qwen3.7 Max, top coding benchmarks
$0.333$0.400$1.333$1.600256K
GLM-5.2 New
glm-5.2 · Zhipu new flagship, 1M context, coding near SOTA
$1.40$1.680$4.40$5.2801M
GLM-5.1
glm-5.1 · previous generation flagship, 203K context
$0.97$1.164$3.04$3.648203K
Kimi K3 New
kimi-k3 · 2.8T MoE, frontier coding, 1M context
$3.00$3.600$15.00$18.0001M
Kimi K2.6
kimi-k2.6 · cost-effective tier, 256K context
$0.903$1.083$3.75$4.500256K
MiniMax M3 New
MiniMax-M3 · 428B MoE, MSA architecture, 1M context, multimodal
$0.292$0.350$1.167$1.4001M

On mobile, swipe the table horizontally to view all input, output, and context rates.

DeepSeek peak: 09:00–12:00 and 14:00–18:00 Beijing time; all other times are off-peak. The rate and rate-card version are locked when a request begins. Strikethrough on fixed-price rows shows the direct provider price.

What real workloads cost

Common usage shapes at XinoAPI prices, estimated per month.

Chat Assistant · 500 users/day

$20–$40/mo
ModelDeepSeek V4-Flash
Daily requests500
Avg tokens/req2K in · 1K out
Monthly volume30M in · 15M out

Code Assistant · 10 devs

$33–$67/mo
ModelDeepSeek V4-Flash
Daily requests300
Avg tokens/req8K in · 2K out
Monthly volume72M in · 18M out

Document Summary · bulk

$15/mo
ModelQwen Plus
Documents/month10,000
Avg tokens/doc8K in · 500 out
Monthly volume80M in · 5M out

Reasoning Workflow

$69–$138/mo
ModelDeepSeek V4-Pro
Daily requests100
Avg tokens/req5K in · 8K out
Monthly volume15M in · 24M out

Free to start, no credit card

Every new account gets $2.00 in credits on signup. That covers roughly 2.5M DeepSeek V4-Flash output tokens at off-peak, or any equivalent published usage.

$2.00

Scale without hidden cuts

No token-price discounting. Larger customers get operational guarantees, billing support, and dedicated routing instead of hidden price cuts.

$20 minimum
PAYG
Cards via Stripe
$500+
Invoice
ACH / wire available
$5K+/mo
Enterprise
Priority routing and support path
Enterprise
Custom
Dedicated channels and compliance

For enterprise volumes (>$5K/month), invoicing, or dedicated routing, contact sales.

Pricing is not the whole decision

XinoAPI is designed for users outside mainland China and routes requests to third-party model providers with different data policies.

ControlCurrent policyWhy it matters
Mainland China accessNot permitted for registration, purchase, dashboard access, or API use.Maintains a clear cross-border service boundary for Chinese LLM inference export.
Prompt/response storageNo plaintext content retention by default; billing uses metadata such as model, tokens, status, and timestamps.Reduces data exposure for production agent and application workloads.
Provider termsUsers must comply with each upstream provider's terms, data policy, and regional restrictions.XinoAPI is a gateway, not the developer or operator of upstream models.
Sensitive dataUse the Privacy SDK for local PII and secret redaction before sending prompts.Provider-side policies vary, especially for models operated in mainland China.

See the Compliance Center and Security Whitepaper for the full policy.

Pricing questions, answered

No — direct provider prices are lower before gateway costs. You pay a markup for unified billing, self-service access in supported regions, optimized routing, and privacy tooling. If you are eligible to use a direct provider and only need one model, direct APIs may be cheaper. If you are outside mainland China, need multiple Chinese model families, or want gateway-level controls, XinoAPI is usually the better operational choice despite the markup.
No. You pay per token at the published rate. There's no API request fee separate from tokens, no inactivity fee, and credits never expire. The minimum Stripe top-up is $20 to keep card processing overhead sustainable.
OpenRouter charges 0% token markup + 5.5% fee on credit purchases (minimum $0.80). XinoAPI charges 20% token markup + 0% purchase fee. At typical usage (~$50–100/month), OpenRouter can be cheaper on pure token cost. XinoAPI's value is in Chinese model specialization, gateway controls, and self-service onboarding — not price competition.
Credit and debit cards (Visa, Mastercard, AmEx) via Stripe. Bank transfers (ACH, wire) are available for enterprise top-ups above $500. We do not accept cryptocurrency, Alipay, or WeChat Pay.
No. Credits never expire on any plan. Unused credits remain in your account indefinitely.
Yes, within 30 days of purchase and provided no more than 10% of the credit has been consumed. Contact support@xinoapi.com with your account email and order ID.
DeepSeek is priced at the provider's Beijing-time rate in effect when a request starts. Peak is 09:00–12:00 and 14:00–18:00; all other times are off-peak. We show cache-hit, cache-miss, and output prices separately, and lock the applicable rate before dispatching the request.
It depends on the model, cache status, and time window. On DeepSeek V4-Flash output, $1 buys about 1.26M tokens off-peak or 630K tokens at peak. Cache-hit input is substantially cheaper than cache-miss input when the provider reports a hit.
No. If an upstream provider returns a 5xx error or the request fails before the model generates output, you're not charged. Rate limit errors (429) and your own malformed requests (4xx) also don't consume credits.
Yes. The XinoAPI Privacy SDK is MIT-licensed and free for any use, including with other LLM providers. Install from PyPI with pip install xinoapi-privacy.
Yes. Open-source maintainers and students can apply for $20/month in free credits by emailing community@xinoapi.com with a link to your project or student ID.
DeepSeek V4 pricing is time-of-use. V4-Flash is $0.264 / $0.528 per 1M cache-miss input tokens and $0.792 / $1.584 output tokens off-peak / peak. V4-Pro is $0.792 / $1.584 cache-miss input and $2.376 / $4.752 output. Cache-hit input is listed separately in the rate card. These prices include XinoAPI's 20% markup.
Use deepseek-v4-flash for general tasks (chat, RAG, code completion) and deepseek-v4-pro for flagship reasoning, planning, and code review. Both follow the published DeepSeek time-of-use rate card. Use explicit V4 model IDs in new projects; avoid relying on legacy DeepSeek aliases.
Qwen Turbo at $0.061 input / $0.245 output per 1M tokens is the cheapest fixed-price option. DeepSeek V4-Flash is a strong quality-to-cost choice when its time-of-use rate fits your workload; check the published peak / off-peak rate before estimating spend.

Start building with $2 free credits

No credit card required. 5 Chinese LLM families. Unified OpenAI-compatible API.

Self-service · Credits never expire · Privacy SDK available

Chat with us