Models and pricing
Per million tokens, USD. OpenAI-compatible — point any client at our base URL and use the model IDs below.
| Model ID | Name | Input | Cached input | Output | Context | Max out |
|---|
| deepseek-v4-flash | DeepSeek V4 Flash | $0.080000 | $0.016000 | $0.180000 | 1.3M | 8K |
| glm-4-7-flash | GLM 4.7 Flash | $0.060000 | $0.010000 | $0.400000 | 203K | 8K |
| glm-5-2 | GLM 5.2 | $0.966000 | $0.193200 | $3.036000 | 1M | 16K |
| kimi-k3 | Kimi K3 | $3.000000 | $0.300000 | $15.000000 | 1M | 16K |
| claude-sonnet-4-5 | Claude Sonnet 4.5 | $3.000000 | $0.300000 | $15.000000 | 1M | 64K |
| claude-opus-4-5 | Claude Opus 4.5 | $5.000000 | $0.500000 | $25.000000 | 200K | 64K |
Cached input is charged at the lower rate when the start of your conversation matches a previous request byte for byte — which is the normal case in a long chat, since only the newest message changes.
Plans
| Plan | Price | Daily allowance | Rate limit | Burst | Concurrent streams |
|---|
| Wisp | $5.00/mo | $8.00 | 20/min | 40 | 2 |
| Shade | $12.00/mo | $25.00 | 40/min | 80 | 3 |
| Umbral | $22.00/mo | $60.00 | 60/min | 120 | 5 |
| Eclipse | $60.00/mo | $120.00 | 100/min | 200 | 8 |
Allowances reset every day and do not roll over. Rate limit is the sustained rate; burst is how many requests can arrive at once before that rate applies, so a run of swipes never blocks. New accounts start with $2 of credit, good for 14 days on any model.