▸ free ai models usage
new: genuinely free AI models are live — try surp/free + see live budgets · token-gating prototype · vote on SRP

free AI models

surp/free — real responses, zero payment, treasury-sponsored.

surp/free gives users genuinely free LLM inference through the same OpenAI-compatible endpoint. There is no x402 payment challenge and no API key. The surp treasury pays the upstream model cost, with strict daily budgets, per-IP limits, cheap-model guardrails, and automatic fallback.

important: We do not resell or proxy personal free-tier credentials from OmniRoute's catalog. Many provider free tiers prohibit proxying or third-party access. OmniRoute's MIT-licensed methodology powers our catalog intelligence; actual public inference is purchased through our normal Surplus account and sponsored by surp.

try it now

curl -X POST https://surp.ivc.lol/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"surp/free",
       "messages":[{"role":"user","content":"Explain x402 in one sentence"}],
       "max_tokens":64,
       "stream":false}'

Limits: 10 requests/IP/day · 128 max output tokens · non-streaming · no tool calls · availability is best-effort.

live sponsored budget (UTC)

99
requests remaining today
19.9K
tokens remaining today
63
eligible live models
3911ms
average free latency
requests served today1 / 100
tokens served today50 / 20,000
failed upstream attempts9
maximum sponsored model price$0.0500 / 1M tokens
fallback attempts per requestup to 3 live models

live models we may sponsor

These are real, liquid Surplus marketplace models currently below the configured treasury ceiling. The router tries the cheapest first and falls back when a seller is unavailable.

#modelUSD/1Mhealthy sellersrequests 24hvolume 24h
1jev-1.13$0.00042232,8898,632
2openai-gpt-oss-20b$0.001135561351
3gemma-3-4b-it$0.00153447018
4gemma-4-26b-a4b-it$0.00203201,50851,078
5gemma-3-12b-it$0.002034445318
6deepseek-v4.1-flash$0.0021262867,551214,711,846
7ministral-3-3b-instruct$0.002262496141
8gemma-3-27b-it$0.00289438013
9gemini-2.5-flash-image$0.003323110,1457,514,132
10qwen3-coder-30b-a3b-instruct$0.003574655588
11qwen3-32b$0.0036688621,380
12openai-gpt-oss-120b$0.0037328113,583286,071
13nvidia-nemotron-3-nano-30b-a3b$0.003712339831
14gemma-4-31b$0.0043474391,081
15gemma-4-31b-it$0.004338116,2355,365,075
16ministral-3-14b-instruct$0.004564569895
17glm-4.7-flash$0.0046795,925570,267
18ministral-3-8b-instruct$0.004664502303
19glm-4.5-air$0.00521151,63814,035
20nvidia-nemotron-nano-9b-v2$0.00575019512
21gpt-6-luna$0.005854229,423200,522,875
22qwen3.8-flash$0.00625847,4711,238,092
23qwen3-235b-a22b-2507$0.00641312,94823,757
24glm-5.3-flash$0.0065177340,008128,447,777
25deepseek-v3.2$0.0066591,032,729350,553,524
26openai-gpt-oss-safeguard-20b$0.007465326107
27nvidia-nemotron-3-super-120b$0.00806831740
28nvidia-nemotron-nano-12b-v2$0.00906019617
29qwen3-coder-next$0.0091542,28963,824
30deepseek-v4-flash$0.009149104,27119,900,543
31qwen3-next-80b-a3b-instruct$0.01071251,4634,533
32gemini-3.8-flash$0.010928244,2603,419,798
33minimax-m2.5$0.01211212,633304,291
34qwen3-coder$0.0129802,06640,287
35openai-gpt-oss-safeguard-120b$0.0148741932
36minimax-m2.1$0.0150621,29225,867
37minimax-m2$0.0150275141,077
38minimax-m3$0.0151305,488586,902
39glm-4.7$0.015219246,60113,191,520
40palmyra-vision-7b$0.0170531830
41glm-5.2$0.0171192122,1766,037,396
42mistral-small-3.2-24b-instruct$0.017164396147
43qwen3.8-omni-flash$0.01731212741,568
44gemini-3.1-flash-lite$0.017522099,7001,438,061
45gpt-5.6-luna$0.018465272,39137,110,575
46deepseek-v3.1$0.02054224,032806,207
47deepseek-v4-pro$0.021134549,54422,868,768
48qwen3-vl-235b-a22b-instruct$0.0211801,09011,513
49glm-4.6$0.021621614,9911,147,529
50gemini-3.7-flash$0.02382118,5387,493,192

models actually served today

modelfree requeststokens
glm-5.3-flash150

what we adopted from OmniRoute

OmniRoute is an MIT-licensed AI gateway with an unusually careful free-tier catalog. We incorporated the parts that improve honesty and transparency:

  • Pool deduplication: if ten models share one provider quota, count the pool once using its maximum—not ten times.
  • Recurring vs one-time separation: signup credits never inflate the steady monthly headline.
  • Uncapped-provider honesty: permanently free providers with no published token cap are listed but never multiplied by RPM × 24/7.
  • ToS risk labels: each catalog entry is classified as ok, ambiguous, caution, avoid, or unknown.
  • Quota visibility: used, remaining, reset/budget limits and actual served models are public.
  • Automatic fallback: failed free models are skipped and the next live candidate is tried.

OmniRoute free-tier catalog snapshot

Source: OmniRoute release/v3.8.50, curated 2026-07-22, MIT License. Snapshot stored with license attribution in this repository.

1.53B
documented recurring tokens/month
2.15B
realistic first month + credits
43
deduped recurring pools
519
catalog model entries
providers represented81
permanently free, uncapped providers13 (listed, never summed)
ToS generally ok49
ToS ambiguous38
ToS caution334
ToS avoid93
catalog entries after excluding avoid426

why not directly pool all those free tiers?

Because "free" often means free for one developer's account—not free to resell through a public proxy. OmniRoute itself surfaces these ToS risks. surp will only integrate a third-party free provider into public routing when its terms explicitly allow proxying/resale or we have written permission. Until then, the catalog is informational and our public free requests are treasury-sponsored.

failure behavior

  • If a model has no healthy sellers, it is excluded before routing.
  • If the first model returns an error, the router tries up to two more candidates.
  • If the daily global budget is exhausted, the API returns HTTP 429 with a paid x402 alternative.
  • If one IP reaches its daily allowance, only that client is throttled.
  • If every sponsored model fails, the API returns structured JSON—never a paid settlement followed by an HTML 500.

related: live system status · cache-aware routing · reward proposal · token-gated access · API docs