surp/free — real responses, zero payment, treasury-sponsored.
surp/free gives users genuinely free LLM inference through the same OpenAI-compatible endpoint. There is no x402 payment challenge and no API key. The surp treasury pays the upstream model cost, with strict daily budgets, per-IP limits, cheap-model guardrails, and automatic fallback.
curl -X POST https://surp.ivc.lol/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"surp/free",
"messages":[{"role":"user","content":"Explain x402 in one sentence"}],
"max_tokens":64,
"stream":false}'
Limits: 10 requests/IP/day · 128 max output tokens · non-streaming · no tool calls · availability is best-effort.
| requests served today | 1 / 100 |
| tokens served today | 50 / 20,000 |
| failed upstream attempts | 9 |
| maximum sponsored model price | $0.0500 / 1M tokens |
| fallback attempts per request | up to 3 live models |
These are real, liquid Surplus marketplace models currently below the configured treasury ceiling. The router tries the cheapest first and falls back when a seller is unavailable.
| # | model | USD/1M | healthy sellers | requests 24h | volume 24h |
|---|---|---|---|---|---|
| 1 | jev-1.13 | $0.0004 | 22 | 32,889 | 8,632 |
| 2 | openai-gpt-oss-20b | $0.0011 | 355 | 613 | 51 |
| 3 | gemma-3-4b-it | $0.0015 | 34 | 470 | 18 |
| 4 | gemma-4-26b-a4b-it | $0.0020 | 320 | 1,508 | 51,078 |
| 5 | gemma-3-12b-it | $0.0020 | 34 | 445 | 318 |
| 6 | deepseek-v4.1-flash | $0.0021 | 262 | 867,551 | 214,711,846 |
| 7 | ministral-3-3b-instruct | $0.0022 | 62 | 496 | 141 |
| 8 | gemma-3-27b-it | $0.0028 | 94 | 380 | 13 |
| 9 | gemini-2.5-flash-image | $0.0033 | 231 | 10,145 | 7,514,132 |
| 10 | qwen3-coder-30b-a3b-instruct | $0.0035 | 74 | 655 | 588 |
| 11 | qwen3-32b | $0.0036 | 68 | 862 | 1,380 |
| 12 | openai-gpt-oss-120b | $0.0037 | 328 | 113,583 | 286,071 |
| 13 | nvidia-nemotron-3-nano-30b-a3b | $0.0037 | 123 | 398 | 31 |
| 14 | gemma-4-31b | $0.0043 | 47 | 439 | 1,081 |
| 15 | gemma-4-31b-it | $0.0043 | 381 | 16,235 | 5,365,075 |
| 16 | ministral-3-14b-instruct | $0.0045 | 64 | 569 | 895 |
| 17 | glm-4.7-flash | $0.0046 | 79 | 5,925 | 570,267 |
| 18 | ministral-3-8b-instruct | $0.0046 | 64 | 502 | 303 |
| 19 | glm-4.5-air | $0.0052 | 115 | 1,638 | 14,035 |
| 20 | nvidia-nemotron-nano-9b-v2 | $0.0057 | 50 | 195 | 12 |
| 21 | gpt-6-luna | $0.0058 | 54 | 229,423 | 200,522,875 |
| 22 | qwen3.8-flash | $0.0062 | 58 | 47,471 | 1,238,092 |
| 23 | qwen3-235b-a22b-2507 | $0.0064 | 131 | 2,948 | 23,757 |
| 24 | glm-5.3-flash | $0.0065 | 177 | 340,008 | 128,447,777 |
| 25 | deepseek-v3.2 | $0.0066 | 59 | 1,032,729 | 350,553,524 |
| 26 | openai-gpt-oss-safeguard-20b | $0.0074 | 65 | 326 | 107 |
| 27 | nvidia-nemotron-3-super-120b | $0.0080 | 68 | 317 | 40 |
| 28 | nvidia-nemotron-nano-12b-v2 | $0.0090 | 60 | 196 | 17 |
| 29 | qwen3-coder-next | $0.0091 | 54 | 2,289 | 63,824 |
| 30 | deepseek-v4-flash | $0.0091 | 49 | 104,271 | 19,900,543 |
| 31 | qwen3-next-80b-a3b-instruct | $0.0107 | 125 | 1,463 | 4,533 |
| 32 | gemini-3.8-flash | $0.0109 | 282 | 44,260 | 3,419,798 |
| 33 | minimax-m2.5 | $0.0121 | 121 | 2,633 | 304,291 |
| 34 | qwen3-coder | $0.0129 | 80 | 2,066 | 40,287 |
| 35 | openai-gpt-oss-safeguard-120b | $0.0148 | 74 | 193 | 2 |
| 36 | minimax-m2.1 | $0.0150 | 62 | 1,292 | 25,867 |
| 37 | minimax-m2 | $0.0150 | 27 | 514 | 1,077 |
| 38 | minimax-m3 | $0.0151 | 30 | 5,488 | 586,902 |
| 39 | glm-4.7 | $0.0152 | 192 | 46,601 | 13,191,520 |
| 40 | palmyra-vision-7b | $0.0170 | 53 | 183 | 0 |
| 41 | glm-5.2 | $0.0171 | 192 | 122,176 | 6,037,396 |
| 42 | mistral-small-3.2-24b-instruct | $0.0171 | 64 | 396 | 147 |
| 43 | qwen3.8-omni-flash | $0.0173 | 12 | 127 | 41,568 |
| 44 | gemini-3.1-flash-lite | $0.0175 | 220 | 99,700 | 1,438,061 |
| 45 | gpt-5.6-luna | $0.0184 | 65 | 272,391 | 37,110,575 |
| 46 | deepseek-v3.1 | $0.0205 | 42 | 24,032 | 806,207 |
| 47 | deepseek-v4-pro | $0.0211 | 345 | 49,544 | 22,868,768 |
| 48 | qwen3-vl-235b-a22b-instruct | $0.0211 | 80 | 1,090 | 11,513 |
| 49 | glm-4.6 | $0.0216 | 216 | 14,991 | 1,147,529 |
| 50 | gemini-3.7-flash | $0.0238 | 211 | 8,538 | 7,493,192 |
| model | free requests | tokens |
|---|---|---|
| glm-5.3-flash | 1 | 50 |
OmniRoute is an MIT-licensed AI gateway with an unusually careful free-tier catalog. We incorporated the parts that improve honesty and transparency:
Source: OmniRoute release/v3.8.50, curated 2026-07-22, MIT License. Snapshot stored with license attribution in this repository.
| providers represented | 81 |
| permanently free, uncapped providers | 13 (listed, never summed) |
| ToS generally ok | 49 |
| ToS ambiguous | 38 |
| ToS caution | 334 |
| ToS avoid | 93 |
| catalog entries after excluding avoid | 426 |
Because "free" often means free for one developer's account—not free to resell through a public proxy. OmniRoute itself surfaces these ToS risks. surp will only integrate a third-party free provider into public routing when its terms explicitly allow proxying/resale or we have written permission. Until then, the catalog is informational and our public free requests are treasury-sponsored.
related: live system status · cache-aware routing · reward proposal · token-gated access · API docs