OpenAI recently announced their fastest inference - Ultrafast during DevDay. It is supposed to be up to 8 times faster (>300 tok/s) in Codex.
Unfortunately, ultrafast is only available on the new Pro $500 plan so I ran some quick tests via the API. My numbers lined up with the numbers publicly mentioned (6-8x faster for 6x the price).
This got me thinking: what would it take to make an open weight model ultrafast (>300 tok/s). OpenAI hasn’t said how it works under the hood. Nvidia wrote a blog post mentioning that the ultrafast mode runs on Blackwell GPUs.
I wanted to see what levers you could tune to get the most speed up and to get a sense of what OpenAI must be doing under the hood to support ultrafast.
Disclaimer: I am not an AI infra expert. OpenAI has dedicated tuning their infra. I am just exploring this with Claude / Codex.
Here's where I ended up:
Qwen3-235B on 4 Nvidia B200s (Blackwell) | Speed per person | GPU cost per million tokens |
|---|---|---|
Where I started: 64 people sharing, no tricks | 71 tok/s | $1.63 |
Where I got to: 1 person, plus a small model guessing ahead | 306 tok/s | $23.52 |
4.3x faster | 14x the cost |
The setup
Qwen3-235B has 235B weights but only uses ~22B per token (it's a mixture-of-experts model). The file is 236 GB.
I rented 4 Nvidia B200s (the Blackwell chips Ultrafast runs on) on Modal for $25/hour, served the model with vLLM, and timed every token on real coding prompts.
Cost = hourly rental ÷ tokens produced.
I started with 64 people sending requests at once, the most I tested. Each got 71 tok/s, at $1.63 per million tokens.
Lever 1: fewer people sharing
People at once | Speed per person | Cost per million tokens |
|---|---|---|
64 | 71 tok/s | $1.63 |
16 | 89 tok/s | $4.91 |
4 | 110 tok/s | $15.74 |
1 | 136 tok/s | $51.19 |
I tested different levels of concurrency (going from 1 to 64) and compared the speed vs cost for each level.
Going from 64 people to 1 got me 1.9x the speed for 31x the cost. Each token still waits for the weights to be read, no matter how many people share the machine.
Lever 2: a small model guessing ahead
This is called speculative decoding. A small helper model guesses the next few tokens and the big model checks them all in one pass. Each correct guess is a token you don't wait for, and the output is identical.
I used an off-the-shelf 2.4 GB helper for Qwen (EAGLE3). The big model accepted ~55% of its guesses on code.
People at once | Speed per person | Cost per million tokens |
|---|---|---|
64 | 138 tok/s | $0.95 |
16 | 177 tok/s | $2.79 |
4 | 236 tok/s | $8.43 |
1 | 306 tok/s | $23.52 |
It roughly doubled the speed at every level and cut the cost. At 64 people, speed went from 71 to 138 tok/s and cost dropped 41%.
Lever 3: more GPUs (didn't make one person faster)
On H100s, I split the model across 8 GPUs instead of 4. Cost doubled for minimal speedup.
One person got 112 tok/s on 4 GPUs and 115 on 8. Here's roughly where the time goes for each token:

Totals are measured. The blue/gray split is my estimate from the chips' memory speed.
Reading the weights from memory (blue) is the part more GPUs speed up. Doubling the GPUs cut it from 1.6 ms to 0.8 ms.
Everything else (gray) is mostly the GPUs syncing. The model has 94 layers, and after each one, every GPU waits for the others before starting the next. That didn't get faster with 8 GPUs. It got a bit slower, since there are more GPUs to wait on.
The math itself takes about 0.04 ms, too small to show.
So more GPUs alone don't make one person faster. They help when lots of people share the machine. With 16 people, 8 GPUs gave each person 81 tok/s vs 59 on 4 GPUs, because there's more math to split up.
What would help one person is less reading and less waiting. That's what Cerebras does. Its chip is one giant wafer with the memory built in next to the math, so weights load fast and there are fewer chips to sync.
What I couldn’t test
Custom GPU code: according to Nvidia, OpenAI used its own models to write GPU code tuned for Blackwell. I used vLLM's off the shelf which probably isn’t fully optimized for the B200 yet.
4-bit weights: half the bytes to read. My 4-bit run didn't fit on one B200.
Bigger racks: Nvidia's GB300 NVL72 connects 72 GPUs so they spend less time syncing. Nvidia didn't say which setup OpenAI uses.
How close did I get?
Normal | Fast | Speed jump | Cost jump | |
|---|---|---|---|---|
OpenAI: Astra Standard → Ultrafast | 74 tok/s | 457 tok/s | 6.2x | 6x the price |
Me: same B200s, 64 people → 1 person | 138 tok/s | 306 tok/s | 2.2x | 25x my GPU cost |
I wasn’t really able to get my ultrafast model to anywhere close to either Astra’s speed or relative price increase.
Now, a 6x price increase seems reasonable for a 6-8x speed increase.
You can use fast open models today
You don't need the $500 plan to try this. Some OpenRouter providers sell fast and normal versions of the same open model. I tested a few:
Same open model | Fast | Normal | Speed | Price |
|---|---|---|---|---|
GLM-5.2 on Decart | 798 tok/s, $8.00 | 115 tok/s, $2.40 | 6.9x | 3.3x |
GLM-5.2 on BaseTen | 381 tok/s, $6.60 | 118 tok/s, $4.40 | 3.2x | 1.5x |
GLM-5.3 on BaseTen | 407 tok/s, $6.60 | 189 tok/s, $4.40 | 2.2x | 1.5x |
MiniMax M2.7, MiniMax's own "highspeed" | 47 tok/s, $2.40 | 54 tok/s, $1.20 | 0.9x | 2x |
Prices are per million output tokens.
Decart's fast GLM-5.2 hit ~800 tok/s for $8 per million, faster than Ultrafast. (OpenRouter numbers include thinking tokens and my Astra numbers don't, so it's a rough comparison.)
MiniMax's "highspeed" tier was slower than its normal one in my test, at 2x the price, so time any speed tier before paying for it.
What if I bought the hardware?
I also priced buying instead of renting. These are reseller/analyst estimates from Sept 2026, spread over 3 years of 24/7 use, plus power or data center space.
Setup | Buy | Speed, 1 person | Cost per million, 1 person | Cost per million, busy |
|---|---|---|---|---|
4 B200s (half of an 8-GPU server) | ~$257,500 | 295 tok/s | $10.99 | $0.45 |
8 used H100s | ~$165,000 | 223 tok/s | $11.25 | $0.38 |
4 RTX PRO 6000 workstation | ~$72,000 | 56 tok/s | $15.66 | $1.71 |
Mac Studio M3 Ultra 512 GB | $9,499 | 24 tok/s | $4.60 | n/a |
Owning is 2-3.5x cheaper than renting, but only if the machine is busy all day. At 10% use, the hardware cost per token is 10x higher.
What I learned
Getting an open model to ~300 tok/s for one person was doable on rented Blackwell GPUs and free software. Doing it cheaply is the hard part.
Sharing is what makes tokens cheap. Giving one person the whole machine only made it ~2x faster, for 25x the cost per token. That's most of why speed tiers cost more.
The guessing helper (speculative decoding) was the best lever: 2x faster and cheaper. Providers likely use it already.
The rest of OpenAI's speed comes from things you can't rent: custom GPU code and bigger, better-connected racks.