Input tokens
Prompt tokens that were not served from the model's cache.
Choose the model on each request. Input, cache, and output are calculated separately and deducted from your internal USD balance.
Prompt tokens that were not served from the model's cache.
Reused context billed at the model's discounted cache rate.
Tokens generated by the model in the response to your request.
Live catalog
5 published models
| Model | Context | Zero Data Retention | Avg. TPSlast 7 days | Inputprice per 1 million tokens | Cacheprice per 1 million tokens | Outputprice per 1 million tokens |
|---|---|---|---|---|---|---|
deepseek-v4-flashdeepseek-v4-flash | 1,048,576 ctx | Yes | 56 tok/s | $0.14 | $0.006 | $0.22 |
deepseek-v4-flash-0731deepseek-v4-flash-0731 | 1,048,576 ctx | Yes | 58 tok/s | $0.14 | $0.006 | $0.22 |
Qwen 3.6 27bqwen3.6-27b | 262,144 ctx | Yes | 37 tok/s | $0.15 | $0.03 | $1.50 |
glm-5.2glm-5.2 | 1,048,576 ctx | Yes | 55 tok/s | $0.70 | $0.15 | $2.20 |
Kimi K3kimi-k3 | 1,048,576 ctx | Yes | 72 tok/s | $2.50 | $0.30 | $9.00 |
TPS is the weighted average observed over the last 7 days and appears after 10 valid responses. Zero Data Retention means prompts and responses are not persisted; technical telemetry and billing data are retained.
Add funds in USD by card and use them until they run out. The wallet is managed internally, and usage is deducted only after a successful response.