Last week, I wrote a post about how new models get cheaper. While doing research for that, I started digging deeper into how models are actually served.
I also wanted to understand why:
Fable is capped at half of a Max plan's weekly usage
OpenAI paused new $200 subscriptions a week after Astra launched.
So I rented a GPU (Nvidia H100) and ran vLLM, an open-source serving program, and served open Qwen models at three sizes. Didn’t optimize anything - just the default vLLM settings.
These are the current prices for these models per million tokens:
Model | Input | Output | Cached input |
|---|---|---|---|
Sonnet 5 | $2 | $10 | $0.20 |
Opus 5.5 | $4 | $20 | $0.20 |
Fable 5.1 | $10 | $50 | $0.25 |
GPT-5.6 Sol | $2 | $10 | $0.20 |
GPT-6 Astra | $10 | $50 | $1.00 |
A model is a file
A model is essentially a file of numbers called weights. The smallest model I tested (Qwen 3 8B), stored at two bytes per weight, uses 16.4 GB of memory.
If you want to understand how LLMs work under the hood, these two resources are helpful:
3Blue1Brown's "But what is a GPT?"
Why models run on GPUs
GPUs are built for matrix multiplication and addition. Running a model is mostly matrix multiplication and addition, so that's where it runs. You could run one on a CPU, but the performance would be pretty terrible: a CPU does a few multiplications at a time and reads memory about ten times slower, while a GPU does thousands at once.
Why a bigger model costs more per token
The model's weights sit in the GPU's memory, but the arithmetic happens in its cores, which hold only a few megabytes. So for every output token, the entire weights stream through the cores a few megabytes at a time.
Memory bandwidth becomes the bottleneck. The more weights a model has, the longer it takes for the weights to stream into the cores, and the longer each output token takes. That's also why tokens per second fall as the model gets bigger.
An H100 moves memory at 3,350 GB per second. Qwen 8B's weights are about 16.4 GB. So it can stream that file at most about 204 times per second (3,350 ÷ 16.4), which is roughly the most output tokens per second one user can get. In my unoptimized setup, I measured 149.
Double the model and each token takes twice as long to stream, so the same card makes half as many tokens per second, and each token costs twice as much GPU time.
Why output costs more than input
Output tokens take more GPU time to make than input tokens. Input tokens all exist up front, so the GPU multiplies each chunk of weights against all of them at once, and one stream of the file handles the whole prompt. Output tokens come one at a time, and each one needs its own full stream. The two phases are called prefill and decode.
In the experiment, for this 8B model and one user, the GPU processed a 9,000-token prompt in about a quarter of a second. Producing 9,000 output tokens took about 65 seconds. With 32 users sharing the card the gap narrows to somewhere between 8x and 24x. The exact numbers depend on the model, its format, and the serving setup.
Why sharing the card lowers the cost
One user writing tokens leaves most of the card idle. Each pass through the weights produces one token. Servers fix this by batching: many users' next-token steps run together, so one pass through the weights produces a token for every one of them.
From here on the numbers use one-byte (FP8) versions of the models instead of the two-byte versions above. The 32B model at two bytes per weight is about 65 GB, which barely fits on an 80 GB card. So I ran all three sizes at one byte to keep them comparable.
H100 rents for $3.95 an hour. With one user, the 8B model wrote 210 tokens per second, which is 756,000 tokens an hour, so a million output tokens cost $3.95 ÷ 0.756 = $5.22 of GPU time. With 32 users sharing the card it made 3,751 tokens per second in total, and the same million cost $0.29. With 128 users, 8,061 per second and $0.14. That's a 38x drop in the GPU cost of each added token. Each user also got slower as the card filled: 63 tokens per second each at 128 users, against 210 alone. All of these assume the card is busy every second of the hour. Idle time raises the real number.
That's probably why Anthropic has offered more usage off-peak. In March 2026 it doubled limits outside peak hours for two weeks, and it tightens session limits during peak hours. Work that moves to a quiet hour fills a card that would otherwise sit idle. The labs' batch APIs at 50% off do the same for work that can wait.
Why cached input costs 10% or less
During the prefill, the model saves each token's working state, the numbers it will look back at while writing. This is called the KV cache. If the next request starts with the same text, the server can reuse that saved state instead of reading those tokens again. Only the matching part from the start of the request counts.
I timed the same 9,000-token prompt cold and warm, measuring how long the server took to produce one token. Cold, 248 milliseconds. Warm, on the second repeat, 21. On the 32B model, 856 versus 35.
That skipped work is why a cache read costs 10% or less of the input price. The first time the server sees a token, it computes that token's state and stores it. Anthropic charges 1.25x the input price for those tokens. That's the cache write. Every later request that starts with the same text pays the cache read price for them instead.
Same card, three sizes
The earlier sections explained why a bigger model should make fewer tokens per second and cost more per token. Here's what happened when I tested it across three sizes of the same model family, with everything else held the same: same GPU, same serving program, same FP8 format, same test.
Size (file) | One user, tokens/sec | 128 users, tokens/sec total | GPU cost per million output |
|---|---|---|---|
8B (9.5 GB) | 210 | 8,061 | $0.14 |
14B (16.3 GB) | 134 | 4,644 | $0.24 |
32B (34.3 GB) | 67 | 1,930 | $0.57 |
Serving speed drops as the file grows. The 14B to 32B step is almost exact: 2.1x the bytes, 2.0x slower for one user. Cost per token rises even with 128 users sharing the card: 1.7x from 8B to 14B, 2.4x from 14B to 32B.
What a long conversation costs
Every turn re-sends the whole conversation, so the input grows as the session goes on. Almost all of it is cache reads at a tenth of the input price. A turn at 100,000 tokens of context costs about $0.02 of input on Sonnet, about the same as turn 1's 8,200 fresh tokens. A long session is cheap per token. It gets expensive in one case: the cache misses, because you paused more than five minutes or something early in the prompt changed, and the whole conversation bills as fresh input again. My prompt caching post covers what breaks the cache.
On the GPU, the cost is the saved state. For each token, the model keeps the numbers it computed at every layer, so later tokens can look back at it. On these Qwen models that's 144 to 256 KiB per token, against 4 bytes for the text. A 100,000-token conversation takes 15 to 26 GB of memory for one user. The card has 80 GB and the weights take their share first, so a bigger model has less room. About 430K tokens of state fit beside the 8B, about 140K beside the 32B. A bigger model holds fewer long conversations at once.
Long contexts also slow the card down. Each new token has to read back the saved state for every token before it. In the experiment, 32 users at 9,000-token contexts ran 2.87x slower than the same 32 users at short prompts, on the 8B model. Fewer tokens per hour from the same card means more GPU time per token.
Anthropic doesn't have a long-context rate anymore. It did until the 4.6 models, when prompts over 200,000 tokens billed at double. Now a 900,000-token request bills at the same per-token rate as a 9,000-token one. Anthropic covers the memory and the slowdown another way: the cache read price on every re-sent token, the 1.25x write price, and the five-minute expiry that frees the memory when you go idle. OpenAI still charges a higher rate past a context threshold.
Why the top tier gets capped
The best guess is that these models need a lot of resources and both labs have given the same answer: not enough capacity, too much strain on the system.
A bigger model makes fewer tokens per second and holds fewer conversations per GPU, so it uses up a fleet faster. When everyone tries the newest model in the same week, the machines don't exist yet. That's the cap.