Tokens Per Second (TPS) Benchmarks · RTX 4090 24GB

Tokens per second —
for your GPU.

We pushed 8 local models — from 9B to 80B — across context windows from 64K to 262K on a single RTX 4090, and logged every result as tokens per second (TPS): the one number that tells you how a GPU actually performs under real inference load.

This is just the first dataset — gputps.com will soon track TPS numbers across many more GPUs, so you know exactly what fits before you buy.
0
Models tested
0
Max context
0
Peak TPS
0
TPS test runs
Live Readout · qwen3.6-35b-a3b
112K CTX
23.5
/ 24.0 GB
Dedicated VRAM
98% utilized under full GPU offload — right at the edge of the swap cliff.
TPS 144.94 tok/s
Tokens 594
Gen Time 1.86 s
Shared Mem 0.7 / 63.9 GB
87.30 → 141.06 → 142.91 → 144.94 → 26.58 → 14.16 → 10.05 tok/s ·  
TPS Ranking

The fastest tokens per second in the test

Top results across every model, quantization, and context size — sorted by raw tokens-per-second (TPS) throughput.

Models

Eight models, one card slot

From the frugal 9B to the 80B heavyweight — here's how every tested model behaves on a single 24GB card.

Deep Dive

The context cliff

Qwen 3.6 35B A3B at full GPU offload, across rising context windows — the point where VRAM pressure makes performance collapse.

Speed vs. Context Window

qwen3.6-35b-a3b · full GPU offload · tok/s
Stable / VRAM < 24GB
Swap zone / performance collapse
Insights

What the numbers actually show

Four takeaways that matter for any 24GB setup — not just these particular models.

01

Full offload backfires on MoE models

For models with few active parameters (A3B, A4B), 100% GPU offload often slows things down through swapping — more than targeted partial offload or CPU-forced MoE layers would.

02

144K is the sweet spot

For Qwen 3.6 35B A3B, a context window around 144K delivers the best balance of memory footprint and tokens/second — beyond that, things get unstable.

03

80B is a bridge too far

Qwen 3 Coder Next 80B A3B needs over 46GB of combined GPU and shared memory even at just 64K context — practically unusable on 24GB.

04

Quantization saves runs

Qwen 3 Coder 30B: 4.86 tok/s at full precision → 96.88 tok/s with Q4_0. Same context, 20x faster.

Test Rig

The bench behind the numbers

GPUNVIDIA GeForce RTX 4090 · 24GB
System RAM128 GB
CPUIntel Xeon E5 2630
Inference EngineLM Studio
Test PeriodMay 29, 2026
QuantizationsFull · Q8_0 · Q4_0

GPU▍TPS tracks real-world tokens-per-second performance for local LLM inference. This RTX 4090 / 24GB dataset is the first entry in a growing library — more GPUs, more models, and more context-window breakdowns are on the way.