← Blog

How much the Claude API really costs per user: a worked example

A worked example: price one real chat session on Opus, Sonnet, and Haiku from the model registry, then per active user per month, and why it sets tiers.

"How much does the Claude API cost per user" has no single answer, because it depends on which model you route to and how long your conversations run. But you can compute it exactly for your app, and the number decides your pricing tiers, your free-tier limits, and which model each plan gets. Here is the actual math, worked from the model registry the kit prices requests against.

Price one real session first

Pricing is per million tokens: Opus 4.8 is $5 in / $25 out, Sonnet 5 is $3 / $15, and Haiku 4.5 is $1 / $5. A chat session is not one request — it is a handful of turns, and because you resend the growing conversation each turn, the input tokens accumulate fast. Take a realistic support-style session that totals about 30,000 input tokens across its turns and 4,000 output tokens. The cost is just the registry rates applied:

export function costUsd(model, usage) {
  const m = MODELS[model];
  return (
    usage.inputTokens * m.inputPerMTok +
    usage.outputTokens * m.outputPerMTok +
    usage.cacheReadTokens * m.inputPerMTok * 0.1 +      // cache reads ~0.1x
    usage.cacheCreationTokens * m.inputPerMTok * 1.25   // cache writes 1.25x
  ) / 1_000_000;
}

Run those 30k in / 4k out through it and the same session costs $0.25 on Opus, $0.15 on Sonnet, and $0.05 on Haiku. Same conversation, a 5x spread — entirely from the model choice.

Now scale it to a user, per month

One session is not a user. Say an active user runs about 40 of those sessions a month. That is $10 per active user on Opus, $6 on Sonnet, $2 on Haiku — before any caching. If your system prompt is large and stable, prompt caching bills those repeated prefix tokens at roughly a tenth of the input rate, which pulls the input side down further; I worked that break-even out separately in the prompt-caching cost math. The point stands: your cost per user is a number you can name, and it is the floor under whatever you charge.

Why per-user attribution changes the decision

Aggregate spend — one big API bill at month end — hides the shape that actually matters. Usage is never uniform: a small fraction of users run 10x the sessions of the median, so a flat price silently makes your light users subsidize your heavy ones, and a single whale on Opus can erase the margin from dozens of paying accounts. You only see this if cost is attributed per user, one row per request, which is exactly what per-user metering is for.

Once you can see it, the product decisions make themselves:

  • Set the free tier to the cheap model, hard-capped. At $2 per fully-active user on Haiku — and far less for a capped free user — the free tier is a marketing cost, not a leak. That is why the kit routes free to Haiku and caps it, the free-tier safety pattern.
  • Price paid tiers above their model cost, not below. If Pro users get Opus and cost $10, a $20 plan has real margin; a $9 plan is a slow bleed you will not notice without per-user numbers.
  • Route by plan deliberately. The gap between $2 and $10 per user is the whole argument for gating models by tier instead of giving everyone your best model.

The takeaway

The cost per user is not a mystery you absorb — it is costUsd applied to real usage, summed per user, per month. Shipwright, the Next.js and Claude starter kit this blog documents, computes it on every request and shows it live: open the demo and watch the exact cost of each message accrue as you type.

Shipwright is a Next.js 16 + Claude starter kit that ships these patterns already done.

Try the live demo →