An API for roleplay
Six models behind one OpenAI-compatible URL. Paste it into SillyTavern and keep the character card you already have. We publish the same prices you can look up anywhere else — the plan is what makes it cheap, not a markup you can't see.
Allowances are in dollars, which tells you nothing. Here they are in replies, on a 16k-token scene with a 400-token answer and 90% of the context cached — which is what a long chat looks like once it gets going.
messages a day on Wisp — $8.00 of credit, 2 at a time, resets every morning. Unused credit doesn't roll over, which is how the allowance stays this large.
Per million tokens. Cached input is the rate that matters for roleplay, because after the first message most of your context hasn't changed.
| Model | Input | Cached | Output | Context |
|---|---|---|---|---|
| DeepSeek V4 Flashdeepseek-v4-flash | $0.080000 | $0.016000 | $0.180000 | 1.3M |
| GLM 4.7 Flashglm-4-7-flash | $0.060000 | $0.010000 | $0.400000 | 203K |
| GLM 5.2glm-5-2 | $0.966000 | $0.193200 | $3.036000 | 1M |
| Kimi K3kimi-k3 | $3.000000 | $0.300000 | $15.000000 | 1M |
| Claude Sonnet 4.5claude-sonnet-4-5 | $3.000000 | $0.300000 | $15.000000 | 1M |
| Claude Opus 4.5claude-opus-4-5 | $5.000000 | $0.500000 | $25.000000 | 200K |
Allowances reset daily, not monthly. You get the whole thing again tomorrow morning instead of rationing it for four weeks.
Chat Completions, streaming, the usual fields. In SillyTavern pick a custom OpenAI-compatible endpoint, paste the base URL and a key, and carry on. Full walkthrough.
curl https://api.umbraapi.dev/v1/chat/completions \
-H "Authorization: Bearer sk-umbra-..." \
-H "Content-Type: application/json" \
-d '{"model":"claude-opus-4-5",
"messages":[{"role":"user","content":"hello"}],
"stream":true}'The $2 is yours for 14 days and works on every model, including Opus. Spend it before you decide anything.
Start with $2 free