API pricing

Pay only for the tokens you use.

Choose the model on each request. Input, cache, and output are calculated separately and deducted from your internal USD balance.

Input tokens

Prompt tokens that were not served from the model's cache.

Cached input

Reused context billed at the model's discounted cache rate.

Output tokens

Tokens generated by the model in the response to your request.

Live catalog

Available models

5 published models

Available modelsprice per 1 million tokens
ModelContextZero Data RetentionAvg. TPSlast 7 daysInputprice per 1 million tokensCacheprice per 1 million tokensOutputprice per 1 million tokens
deepseek-v4-flashdeepseek-v4-flash1,048,576 ctxYes56 tok/s$0.14$0.006$0.22
deepseek-v4-flash-0731deepseek-v4-flash-07311,048,576 ctxYes58 tok/s$0.14$0.006$0.22
Qwen 3.6 27bqwen3.6-27b262,144 ctxYes37 tok/s$0.15$0.03$1.50
glm-5.2glm-5.21,048,576 ctxYes55 tok/s$0.70$0.15$2.20
Kimi K3kimi-k31,048,576 ctxYes72 tok/s$2.50$0.30$9.00

TPS is the weighted average observed over the last 7 days and appears after 10 valid responses. Zero Data Retention means prompts and responses are not persisted; technical telemetry and billing data are retained.

One balance, always in USD.

Add funds in USD by card and use them until they run out. The wallet is managed internally, and usage is deducted only after a successful response.

Add balance