umbra

Models and pricing

Per million tokens, USD. OpenAI-compatible — point any client at our base URL and use the model IDs below.

Model IDNameInputCached inputOutputContextMax out
deepseek-v4-flashDeepSeek V4 Flash$0.080000$0.016000$0.1800001.3M8K
glm-4-7-flashGLM 4.7 Flash$0.060000$0.010000$0.400000203K8K
glm-5-2GLM 5.2$0.966000$0.193200$3.0360001M16K
kimi-k3Kimi K3$3.000000$0.300000$15.0000001M16K
claude-sonnet-4-5Claude Sonnet 4.5$3.000000$0.300000$15.0000001M64K
claude-opus-4-5Claude Opus 4.5$5.000000$0.500000$25.000000200K64K
Cached input is charged at the lower rate when the start of your conversation matches a previous request byte for byte — which is the normal case in a long chat, since only the newest message changes.

Plans

PlanPriceDaily allowanceRate limitBurstConcurrent streams
Wisp$5.00/mo$8.0020/min402
Shade$12.00/mo$25.0040/min803
Umbral$22.00/mo$60.0060/min1205
Eclipse$60.00/mo$120.00100/min2008
Allowances reset every day and do not roll over. Rate limit is the sustained rate; burst is how many requests can arrive at once before that rate applies, so a run of swipes never blocks. New accounts start with $2 of credit, good for 14 days on any model.