We priced the same coding-agent session on Ollama Cloud, OpenRouter and OpenCode Go. Ollama's new credit plans win on flagships; Go wins on cheap models.
The short answer: for a coding agent that leans on expensive flagship open models like Kimi K3 or GLM-5.3, Ollama Cloud Pro is now the cheapest of the three, because its $20 plan includes $60 of token credits. For cheap, fast models such as MiniMax M3 or GLM-5.3-Flash, OpenCode Go at $10 a month is hard to beat. OpenRouter wins when you need the widest catalogue, occasional use with no subscription, or one key that also reaches Claude and GPT-6.
That ranking changed recently. On August 31, 2026 Ollama dropped its 5-hour and weekly session limits and moved Pro, Max and Team to per-token pricing with a monthly credit pool. Most comparisons written before September still describe the old model.
This guide was last checked against ollama.com/pricing, the Ollama blog and docs, openrouter.ai/pricing and the OpenRouter docs and models API, and opencode.ai/docs/go and /docs/zen on September 24, 2026.
Quick summary
- Ollama Cloud: Free ($0, starter credits, 1 concurrent request), Pro ($20/month or $200/year, $60 of credits, 3 concurrent), Max ($100/month, $300 of credits, 10 concurrent). When the credits run out you keep going at the same per-token rate. No 5-hour or weekly limits on the new plans.
- OpenRouter: no subscription. You buy credits and pay the provider's per-token price plus a 5.5% platform fee on card purchases (5% via crypto). Free
:freemodels are capped at 20 requests per minute and 50 per day, or 1,000 per day once you have bought $10 of credits. - OpenCode Go: $10/month flat. Each model has its own monthly dollar cap ($15, $30 or $60 of usage at list prices), with a 5-hour window worth 20% of that cap and a weekly window worth 50%. The docs list 32 models today.
- Privacy: all three say no training on your prompts by default, with exceptions you should know about (details below).
What each one actually sells
From inside a coding agent the three look the same: a base URL and a model name. Underneath, they sell different things.
Ollama Cloud sells prepaid credit at a 3x discount. Pro costs $20 and loads $60 of usage at Ollama's published token rates each month. Max costs $100 and loads $300. Included credit does not roll over. If you are on an older Pro or Max plan, it keeps its old session and weekly limits until you switch or change billing cycle. Our earlier Ollama Pro vs Max guide covers the local-versus-cloud split, which still holds.
OpenRouter sells routing. It is a marketplace of more than 450 models from dozens of providers. You pay whatever the chosen provider charges, and OpenRouter takes 5.5% when you top up. Bring-your-own-key traffic is free up to $25,000 a month of list-price inference, then 5%.
OpenCode Go sells a subsidised bundle. It costs $10 and gives you a separate dollar budget on every model. On models where OpenCode has negotiated bulk rates, such as GLM-5.3-Flash, MiniMax M3 and Kimi K2.7 Code, the cap is $60 a month, six times what you paid. On new or already-cheap flagships such as Kimi K3, GLM-5.3, Qwen3.8 Max and DeepSeek V4 Pro, the cap is only $15. Our OpenCode Go plan review goes deeper on day-to-day use.
The test session: what one agent run costs
To compare like for like, we priced one realistic coding-agent session:
- 150 requests (an hour or so of Claude Code or OpenCode work)
- 9.0M cached input tokens (about 60K of repeated context per request)
- 0.225M fresh input tokens (1,500 per request)
- 0.06M output tokens (400 per request)
That shape is close to the per-request averages OpenCode publishes for Go, where cached context dominates.
Two more assumptions. Sessions run in US working hours, which falls mostly inside Ollama's DeepSeek peak window (12:00 to 18:00 UTC on weekdays) and outside OpenCode Go's (01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays), so we use Ollama's peak DeepSeek rates and Go's off-peak ones. And OpenRouter prices are the default listed price for each model plus the 5.5% fee.
Cost per session at list price
| Model | Ollama Cloud | OpenRouter (incl. 5.5% fee) | OpenCode Go (counted against cap) | Go cap per month | Sessions per Go 5-hour window |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | $0.19 (off-peak $0.10) | $0.10 | $0.10 | $15 ($60 until Sep 27) | 31 |
| DeepSeek V4 Pro | $0.93 (off-peak $0.47) | $1.04 | $0.47 | $15 | 6 |
| GLM-5.3-Flash | $0.33 | $0.54 | $0.33 | $60 | 36 |
| MiniMax M3 | $1.36 | $0.72 | $0.68 | $60 | 17 |
| Kimi K2.7 Code | $2.16 | $2.07 | $2.16 | $60 | 5 |
| GLM-5.3 | $2.92 | $3.08 | $2.92 | $15 | 1 |
| Qwen3.8 Max | not offered | $3.23 | $3.06 | $15 | 1 |
| Kimi K3 | $4.28 | $4.51 | $4.28 | $15 | 0.7 |
List prices barely differ for the flagships: Ollama, OpenCode and OpenRouter's default listing all charge $1.40 input, $0.26 cached and $4.40 output per million tokens for GLM-5.3, and $3.00, $0.30 and $15.00 for Kimi K3. Ollama does charge double for MiniMax M3 ($0.60 in, $2.40 out versus $0.30 and $1.20). And OpenCode Go's flagship caps are small. A $15 cap on Kimi K3 covers about three and a half of our sessions a month, and the 5-hour window runs out before one session finishes.
The monthly bill: 20 sessions
Per-session prices only tell half the story, because Ollama and Go both bundle usage into a subscription. Here is what 20 of those sessions a month (roughly one per working day) costs on each, picking the cheapest plan that covers it.
Monthly, same workload: Ollama / OpenRouter / OpenCode Go
| Model (20 sessions/month) | Ollama Cloud | OpenRouter | OpenCode Go |
|---|---|---|---|
| DeepSeek V4.1 Flash | $4 (pay-as-you-go credits) | $2 | $10 |
| MiniMax M3 | $20 (Pro) | $14 | $10 |
| GLM-5.3 | $20 (Pro) | $62 | $53 |
| Kimi K3 | $46 (Pro + $26 extra) | $90 | $81 |
How the numbers work. Ollama Pro's $60 of credit covers 20 GLM-5.3 sessions ($58 at list price), so the bill is the $20 subscription. Kimi K3 needs $86, so Pro plus $26 of extra credit. OpenCode Go covers the first few flagship sessions inside the $15 cap, then, if you turn on "Use balance", bills the rest to your OpenCode Zen balance at Zen's rates (which match Go's list prices for these models, plus card fees of 4.4% + $0.30 per top-up). Without that switch, Go simply stops serving the model until the window resets.
At 60 sessions a month the gap widens. Ollama Max ($100, $300 of credit) covers 60 GLM-5.3 sessions (about $175 at list) or 60 Kimi K3 sessions (about $257). The same Kimi K3 workload is about $271 on OpenRouter and about $252 on Go with Zen overflow.
For cheap models the ranking flips. At 60 sessions of MiniMax M3, Go still costs $10 because the whole workload fits inside its $60 cap, while Ollama would need about $82 of usage at its higher MiniMax rate.
Rate limits and what happens when you hit them
This is where the three differ most in daily use.
| Behaviour | Ollama Cloud | OpenRouter | OpenCode Go |
|---|---|---|---|
| Time windows | None on new plans (monthly credit only) | None on paid models | 5-hour (20% of cap), weekly (50%), monthly |
| Concurrency | Free 1, Pro 3, Max and Team 10 | No platform-level cap on paid models | Not published |
| When you run out | Keeps going at per-token rates from purchased credit | Requests fail when credit hits zero | Blocked per model, or falls back to Zen balance if enabled |
| Free usage | Starter credits on starter models | :free models, 20 RPM, 50/day (1,000/day after $10 purchase) | Free models (currently "Space Bunny Free", limited time) |
| Peak pricing | DeepSeek models cost 2x from 12:00 to 18:00 UTC weekdays | Depends on provider | DeepSeek models cost 2x in two UTC morning windows on weekdays |
Concurrency is the Ollama catch. Claude Code and OpenCode both fan out subagents, and Pro allows three requests at once. Extra requests queue, and when the queue fills they are rejected. If you run parallel agents all day, that is the practical reason to pay for Max, not the credit. Speed is a separate question. Users have reported slow token rates on Ollama's hosted models this year, which we covered in Ollama Cloud slowdown.
OpenCode Go's per-model windows push you to rotate models. Spend a 5-hour window of Kimi K3 on planning, then switch to MiniMax M3 or GLM-5.3-Flash for the edit loop. OpenCode has also cut caps before: its DeepSeek allowance was reduced earlier this year, as we reported in OpenCode Go slashes DeepSeek quotas. The docs say plainly that "usage limits may change."
Model catalogue: who has what
| Model | Ollama Cloud | OpenRouter | OpenCode Go |
|---|---|---|---|
| DeepSeek V4.1 Flash | Yes | Yes | Yes |
| DeepSeek V4 Pro | Yes | Yes | Yes |
| GLM-5.3 / GLM-5.3-Flash | Yes | Yes | Yes |
| Kimi K3 / Kimi K2.7 Code | Yes | Yes | Yes |
| MiniMax M3 | Yes | Yes | Yes |
| Qwen3.8 Max / Qwen3.8 Flash | No | Yes | Yes |
| MiMo-V2.6 Pro / Flash | No | Yes | Yes |
| gpt-oss 120B / 20B | Yes | Yes | No |
| Gemma 4, Nemotron 3 | Yes | Yes | No |
| GPT-6 Luna, Grok 4.7 (closed) | No | Yes | Yes |
| Claude, GPT-6 Astra (closed) | No | Yes | No |
| Catalogue size | About 20 cloud models on the pricing table | 450+ models | 32 models listed in the docs |
The overlap is large. All three carry the models most people actually use in coding agents in September 2026: DeepSeek V4.1 Flash, GLM-5.3, Kimi K3 and MiniMax M3. Ollama is missing the Qwen3.8 and MiMo lines. OpenCode Go skips gpt-oss, which is one of the cheapest capable models on Ollama at $0.15 input and $0.60 output. For how the Chinese labs' own token plans compare with these resellers, see AI token plans explained.
Privacy and data terms
| Ollama Cloud | OpenRouter | OpenCode Go | |
|---|---|---|---|
| Training on your prompts | "Never logged or trained on" | Off unless you allow providers that train, set separately for paid and free models | Not used, except Muse Spark "Contributor" models, which trade discounts for training rights |
| Retention | Ollama requires no-logging and zero-retention terms from its hosting partners | OpenRouter keeps nothing unless you opt into logging (1% discount); upstream provider policy varies | 0 days for most open models; 30 days for Grok 4.7/4.6 and GPT Luna models |
| Zero data retention control | Default | Account-wide, per model group or per request with "zdr": true | Per model (DeepSeek ZDR agreement renewed monthly, current one valid to September 30, 2026) |
| Where it runs | Mainly the US, overflow to Europe and Singapore | Depends on the provider OpenRouter routes to | Depends on the provider per model |
OpenRouter gives you the most control, but only if you use it. Its router can send the same model to about 25 different providers (DeepSeek V4.1 Flash has that many today), each with its own terms and its own cache pricing. Turn on ZDR enforcement in your account settings before pointing a work repository at it.
Setup in Claude Code, OpenCode and Cline
All three speak an OpenAI-compatible API. Ollama and OpenCode Go also speak Anthropic's Messages API, which is what Claude Code needs.
| Tool | Ollama Cloud | OpenRouter | OpenCode Go |
|---|---|---|---|
| Claude Code base URL | https://ollama.com | https://openrouter.ai/api | https://opencode.ai/zen/go |
| OpenAI-compatible base URL (Cline, others) | https://ollama.com/v1 | https://openrouter.ai/api/v1 | https://opencode.ai/zen/go/v1 |
| OpenCode | ollama launch opencode | Built-in provider | /connect, then opencode-go/<model-id> |
For Claude Code on Ollama Cloud, the Ollama docs give this pattern (note that the cloud endpoint requires bearer auth):
ANTHROPIC_BASE_URL=https://ollama.com \
ANTHROPIC_AUTH_TOKEN="$OLLAMA_API_KEY" \
ANTHROPIC_API_KEY="" \
claude --model glm-5.3-flash
For OpenRouter, OpenRouter's guide sets ANTHROPIC_BASE_URL="https://openrouter.ai/api", puts your OpenRouter key in ANTHROPIC_AUTH_TOKEN, and blanks ANTHROPIC_API_KEY. OpenRouter warns that Claude Code is tuned for Anthropic models and may misbehave with others.
For OpenCode Go, Claude Code is on OpenCode's list of validated clients. There is one important limit: only the MiniMax and Qwen models are served on Go's Anthropic-style /v1/messages endpoint. DeepSeek, GLM and Kimi are on /chat/completions, so Claude Code cannot reach them directly. Community guides also report that Go wants the key in ANTHROPIC_API_KEY (sent as x-api-key) rather than as a bearer token. If you want GLM or Kimi on Go, use OpenCode itself. The CLI trade-offs are covered in Claude Code vs Codex CLI vs Antigravity CLI vs OpenCode.
Which should you pick?
You live in OpenCode and mostly use fast models. Take OpenCode Go. $10 buys around 88 of our sessions a month on MiniMax M3 or about 180 on GLM-5.3-Flash. Nothing else comes close for that workload.
You want Kimi K3 or GLM-5.3 as your daily driver. Take Ollama Cloud Pro. The $60 of credit covers about 20 GLM-5.3 sessions or 14 Kimi K3 sessions, and extra use is billed at the same rate with no fee. Move to Max if you run several agents in parallel, since Max raises concurrency from 3 to 10.
You code a few hours a month, or want to test many models. Use OpenRouter. There is no subscription to waste, and you can try Qwen3.8 Max, MiMo, gpt-oss and closed models behind one key. Buy $10 of credit once so the free models get 1,000 requests a day instead of 50.
You use Claude Code and want open models for subagents. Ollama Cloud is the smoothest, because every model is reachable through its Anthropic-compatible endpoint. OpenCode Go only exposes MiniMax and Qwen to Claude Code. If you are weighing this against a Claude subscription, see Claude Pro vs Max for Claude Code after Opus 5.5.
You want to spend $0. Run models locally, which stays unlimited on Ollama, and keep OpenRouter's free models as a fallback. Our $0 AI coding stack guide walks through it.
The combination many people land on: OpenCode Go ($10) for the high-volume edit loop plus Ollama Pro ($20) for flagship planning and review. That is $30 a month for $60 of Ollama credit plus Go's per-model caps of $15 to $60 each.
Caveats
- Our session is one shape. An agent that re-reads large files without caching will pay far more for fresh input, and the rankings can shift. Cached input is about 97% of our token count, so a service with poor cache pricing loses badly.
- OpenRouter's price is not fixed per model. The listed price is the default route. DeepSeek V4.1 Flash is offered by providers charging from $0.04 to $0.375 per million input tokens, with cache pricing that differs by more than ten times. Pin a provider if you need predictable bills.
- Promotions expire. OpenCode Go's 4x DeepSeek V4.1 Flash cap ends September 27, 2026, and the DeepSeek ZDR agreement is renewed monthly.
- Ollama's free tier does not state a dollar amount for its starter credits, and you must buy credit to unlock non-starter models. The $4 DeepSeek figure above assumes pay-as-you-go credit with no subscription.
FAQ
Is an Ollama Cloud subscription worth it?
Yes, if you spend more than about $20 a month on open-model tokens. Since August 31, 2026, Pro turns $20 into $60 of usage at Ollama's list prices, with no 5-hour or weekly limits. If you spend less than $20, pay-as-you-go credit on the free plan is cheaper.
What are Ollama Cloud's pricing and usage limits?
Free is $0 with starter credits and 1 concurrent request. Pro is $20 a month ($200 a year) with $60 of credit and 3 concurrent requests. Max is $100 with $300 of credit and 10 concurrent requests. Team is $500 with $1,000 of shared credit. The only hard limits on the new plans are credit and concurrency.
Ollama Cloud vs OpenRouter: which is cheaper?
For steady use, Ollama Pro or Max, because of the credit multiplier. At list price the two are close for most models, but OpenRouter adds a 5.5% fee on card top-ups and Ollama adds nothing. OpenRouter is cheaper for light use and for models where Ollama's rate is higher, such as MiniMax M3.
Ollama Cloud vs OpenCode Go: which should I buy?
OpenCode Go for cheap, high-volume models in OpenCode. Ollama Cloud Pro for Kimi K3, GLM-5.3 and DeepSeek V4 Pro, and for Claude Code. Go's $15 monthly cap on flagship models runs out in a few sessions.
OpenRouter vs OpenCode Go: what's the difference?
OpenRouter is pay-as-you-go access to 450+ models with no bundled discount. OpenCode Go is a $10 subscription to 32 curated coding models with per-model dollar caps worth up to $60 a month each. Go is far cheaper for anyone who uses it daily. OpenRouter has the breadth.
Do these services train on my code?
Ollama says prompts are never logged or trained on. OpenRouter keeps no prompts by default and lets you block providers that train, with separate settings for paid and free models. OpenCode Go says most models are not trained on and kept 0 days, except the Muse Spark Contributor models, which are used for training in exchange for lower prices.