OpenAI recently announced their fastest inference - Ultrafast during DevDay. It is supposed to be up to 8 times faster (>300 tok/s) in Codex.

Unfortunately, ultrafast is only available on the new Pro $500 plan so I ran some quick tests via the API. My numbers lined up with the numbers publicly mentioned (6-8x faster for 6x the price).

This got me thinking: what would it take to make an open weight model ultrafast (>300 tok/s). OpenAI hasn’t said how it works under the hood. Nvidia wrote a blog post mentioning that the ultrafast mode runs on Blackwell GPUs.

I wanted to see what levers you could tune to get the most speed up and to get a sense of what OpenAI must be doing under the hood to support ultrafast.

Disclaimer: I am not an AI infra expert. OpenAI has dedicated tuning their infra. I am just exploring this with Claude / Codex.

Here's where I ended up:

Qwen3-235B on 4 Nvidia B200s (Blackwell)

Speed per person

GPU cost per million tokens

Where I started: 64 people sharing, no tricks

71 tok/s

$1.63

Where I got to: 1 person, plus a small model guessing ahead

306 tok/s

$23.52

4.3x faster

14x the cost

The setup

Qwen3-235B has 235B weights but only uses ~22B per token (it's a mixture-of-experts model). The file is 236 GB.

I rented 4 Nvidia B200s (the Blackwell chips Ultrafast runs on) on Modal for $25/hour, served the model with vLLM, and timed every token on real coding prompts.

Cost = hourly rental ÷ tokens produced.

I started with 64 people sending requests at once, the most I tested. Each got 71 tok/s, at $1.63 per million tokens.

Lever 1: fewer people sharing

People at once

Speed per person

Cost per million tokens

64

71 tok/s

$1.63

16

89 tok/s

$4.91

4

110 tok/s

$15.74

1

136 tok/s

$51.19

I tested different levels of concurrency (going from 1 to 64) and compared the speed vs cost for each level.

Going from 64 people to 1 got me 1.9x the speed for 31x the cost. Each token still waits for the weights to be read, no matter how many people share the machine.

Lever 2: a small model guessing ahead

This is called speculative decoding. A small helper model guesses the next few tokens and the big model checks them all in one pass. Each correct guess is a token you don't wait for, and the output is identical.

I used an off-the-shelf 2.4 GB helper for Qwen (EAGLE3). The big model accepted ~55% of its guesses on code.

People at once

Speed per person

Cost per million tokens

64

138 tok/s

$0.95

16

177 tok/s

$2.79

4

236 tok/s

$8.43

1

306 tok/s

$23.52

It roughly doubled the speed at every level and cut the cost. At 64 people, speed went from 71 to 138 tok/s and cost dropped 41%.

Lever 3: more GPUs (didn't make one person faster)

On H100s, I split the model across 8 GPUs instead of 4. Cost doubled for minimal speedup.

One person got 112 tok/s on 4 GPUs and 115 on 8. Here's roughly where the time goes for each token:

Time to make one token on H100s for one person: 4 GPUs take 8.9 ms (1.6 ms reading weights, 7.2 ms everything else); 8 GPUs take 8.7 ms (0.8 ms reading, 7.8 ms everything else)

Totals are measured. The blue/gray split is my estimate from the chips' memory speed.

Reading the weights from memory (blue) is the part more GPUs speed up. Doubling the GPUs cut it from 1.6 ms to 0.8 ms.

Everything else (gray) is mostly the GPUs syncing. The model has 94 layers, and after each one, every GPU waits for the others before starting the next. That didn't get faster with 8 GPUs. It got a bit slower, since there are more GPUs to wait on.

The math itself takes about 0.04 ms, too small to show.

So more GPUs alone don't make one person faster. They help when lots of people share the machine. With 16 people, 8 GPUs gave each person 81 tok/s vs 59 on 4 GPUs, because there's more math to split up.

What would help one person is less reading and less waiting. That's what Cerebras does. Its chip is one giant wafer with the memory built in next to the math, so weights load fast and there are fewer chips to sync.

What I couldn’t test

  • Custom GPU code: according to Nvidia, OpenAI used its own models to write GPU code tuned for Blackwell. I used vLLM's off the shelf which probably isn’t fully optimized for the B200 yet.

  • 4-bit weights: half the bytes to read. My 4-bit run didn't fit on one B200.

  • Bigger racks: Nvidia's GB300 NVL72 connects 72 GPUs so they spend less time syncing. Nvidia didn't say which setup OpenAI uses.

How close did I get?

Normal

Fast

Speed jump

Cost jump

OpenAI: Astra Standard → Ultrafast

74 tok/s

457 tok/s

6.2x

6x the price

Me: same B200s, 64 people → 1 person

138 tok/s

306 tok/s

2.2x

25x my GPU cost

I wasn’t really able to get my ultrafast model to anywhere close to either Astra’s speed or relative price increase.

Now, a 6x price increase seems reasonable for a 6-8x speed increase.

You can use fast open models today

You don't need the $500 plan to try this. Some OpenRouter providers sell fast and normal versions of the same open model. I tested a few:

Same open model

Fast

Normal

Speed

Price

GLM-5.2 on Decart

798 tok/s, $8.00

115 tok/s, $2.40

6.9x

3.3x

GLM-5.2 on BaseTen

381 tok/s, $6.60

118 tok/s, $4.40

3.2x

1.5x

GLM-5.3 on BaseTen

407 tok/s, $6.60

189 tok/s, $4.40

2.2x

1.5x

MiniMax M2.7, MiniMax's own "highspeed"

47 tok/s, $2.40

54 tok/s, $1.20

0.9x

2x

Prices are per million output tokens.

Decart's fast GLM-5.2 hit ~800 tok/s for $8 per million, faster than Ultrafast. (OpenRouter numbers include thinking tokens and my Astra numbers don't, so it's a rough comparison.)

MiniMax's "highspeed" tier was slower than its normal one in my test, at 2x the price, so time any speed tier before paying for it.

What if I bought the hardware?

I also priced buying instead of renting. These are reseller/analyst estimates from Sept 2026, spread over 3 years of 24/7 use, plus power or data center space.

Setup

Buy

Speed, 1 person

Cost per million, 1 person

Cost per million, busy

4 B200s (half of an 8-GPU server)

~$257,500

295 tok/s

$10.99

$0.45

8 used H100s

~$165,000

223 tok/s

$11.25

$0.38

4 RTX PRO 6000 workstation

~$72,000

56 tok/s

$15.66

$1.71

Mac Studio M3 Ultra 512 GB

$9,499

24 tok/s

$4.60

n/a

Owning is 2-3.5x cheaper than renting, but only if the machine is busy all day. At 10% use, the hardware cost per token is 10x higher.

What I learned

Getting an open model to ~300 tok/s for one person was doable on rented Blackwell GPUs and free software. Doing it cheaply is the hard part.

  • Sharing is what makes tokens cheap. Giving one person the whole machine only made it ~2x faster, for 25x the cost per token. That's most of why speed tiers cost more.

  • The guessing helper (speculative decoding) was the best lever: 2x faster and cheaper. Providers likely use it already.

  • The rest of OpenAI's speed comes from things you can't rent: custom GPU code and bigger, better-connected racks.