real output TPS, TTFT, and throughput-per-dollar — measured, not claimed.
Every number on this page comes from a real streaming request sent through surp to the Surplus Intelligence marketplace. No vendor spec sheets, no extrapolated benchmarks — only observed generation throughput. This is the data the Surplus dashboard doesn't show.
Models ranked by throughput-value score (p50 output TPS × tokens-per-dollar). Green ≥ 80 TPS, amber 40-79, red < 40.
| # | model | p50 TPS | p95 TPS | range | p50 TTFT | price /1M | throughput/$ | runs |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.6-luna | 279.6 | 400000.0 | 49.4-400000.0 | 5038ms | $0.0230 | 12.18B | 196/216 |
| 2 | glm-5.2 | 117.4 | 276.8 | 15.8-1619.4 | 3670ms | $0.0171 | 6.89B | 221/221 |
| 3 | deepseek-v4-flash | 71.7 | 204.4 | 9.7-492.6 | 2280ms | $0.0149 | 4.82B | 225/225 |
| 4 | deepseek-v4-flash-0731 | 84.5 | 134.3 | 8.4-380.9 | 2789ms | $0.0360 | 2.35B | 224/225 |
| 5 | deepseek-v4-pro | 69.8 | 117.2 | 36.6-208.9 | 2428ms | $0.0312 | 2.24B | 220/220 |
| 6 | llama-3.3-70b-instruct | 54.6 | 516.4 | 12.3-551.7 | 1563ms | $0.0611 | 893.8M | 213/215 |
Output TPS = generated completion tokens ÷ generation seconds (time from first token to last token). This is the metric LLM buyers care about: how fast the model produces text.
This is different from request RPS (requests per second), which measures how many separate requests the gateway handles. The health board tracks RPS; this page tracks output TPS.
POST /v1/chat/completions through surp with max_tokens=400.completion_tokens / generation_seconds.p50_output_tps × (1,000,000 / price_per_1m) — higher is better.usage.completion_tokens from the API; falls back to len(text)/4 estimate if usage is absent.Last 10 raw runs — verify the numbers yourself.
| output TPS | TTFT | wall | tokens | price/1M | status |
|---|---|---|---|---|---|
| 1612.9 | 4668ms | 4916ms | 400 | $0.0187 | ok |
| 1632.7 | 6217ms | 6462ms | 400 | $0.0187 | ok |
| 207.4 | 2216ms | 3209ms | 206 | $0.0187 | ok |
| 178.4 | 2160ms | 4077ms | 342 | $0.0187 | ok |
| 400000.0 | 4595ms | 4596ms | 400 | $0.0187 | ok |
| 0.0 | 6322ms | 6322ms | 400 | $0.0169 | ok |
| 264.6 | 4091ms | 5603ms | 400 | $0.0169 | ok |
| 201.2 | 1840ms | 2814ms | 196 | $0.0169 | ok |
| 209.9 | 2713ms | 4595ms | 395 | $0.0169 | ok |
| 1646.1 | 6149ms | 6392ms | 400 | $0.0169 | ok |
The benchmark runner is open source. Clone the repo and run:
python3 benchmark_runner.py --model deepseek-v4-flash-0731 --runs 10 --api-key sk-sur-...
You'll get the same kind of numbers we publish here. Or just make a streaming request and time it:
curl -N -X POST https://surp.ivc.lol/v1/chat/completions \
-H "Authorization: Bearer sk-sur-..." \
-H "Content-Type: application/json" \
-d '{"model":"surp/direct/deepseek-v4-flash-0731",
"messages":[{"role":"user","content":"Write 200 words about HTTP/2"}],
"max_tokens":400,
"stream":true}'
related: provider health board (RPS, latency) · free models · cache-affinity auction · benchmark API