compute once. reuse safely. pay less.
LLM APIs waste money when they recompute the same context over and over. Agent tool definitions, system prompts, codebases, policies, and long conversation prefixes are often identical between requests. surp.ivc.lol now uses a two-layer cache engine to avoid that waste while preserving live marketplace pricing.
Some upstream providers maintain prompt-prefix or KV caches. If a request reaches the same provider worker with the same stable prefix, the provider may avoid part of the input computation and return faster or charge less.
Surp keeps a recently selected model for up to five minutes when it remains within 30% of the current cheapest eligible model. This improves the odds of upstream cache reuse, but Surplus may still choose a different seller or worker behind that model. Sticky routing is therefore a locality hint, not a guaranteed provider-cache hit.
request 1 → cheapest model A → provider writes prefix cache request 2 → A is still within tolerance → reuse A → cache read request 3 → model B becomes 40% cheaper → switch to B
This balances two things that normally conflict: marketplace arbitrage and cache locality.
Some requests are completely repeatable: temperature zero, no tools, one response, and no streaming. For those requests, surp.ivc.lol stores the completed JSON response under a SHA-256 fingerprint of the full request.
When the exact request arrives again within the cache window:
X-Surp-Cache: HIT and cache metadata.We deliberately do not cache every request.
| request type | exact response cache | sticky/provider cache |
|---|---|---|
| temperature 0, no tools, non-streaming | eligible | eligible |
| streaming response | bypass | eligible |
| tool or agent call | bypass | eligible |
| creative/nonzero temperature | bypass | eligible |
multiple candidates (n > 1) | bypass | eligible |
Privacy boundary: raw prompts are not persisted in the exact-cache database. It stores a SHA-256 request fingerprint and the completed response for up to 15 minutes. Upstream sellers still receive the original request and apply their own retention policies.
temperature: 0 for deterministic tasks that may repeat.Every response reports its cache path:
X-Surp-Cache: HIT | MISS | BYPASS X-Surp-Cache-Type: exact-response X-Surp-Routing: sticky-within-tolerance | live-cheapest
The public status page and GET /api/stats expose cumulative hits, misses, eligible lookups, saved tokens, live entries, entry ages, and the 15-minute TTL. The hit rate is historical: with only 19 eligible lookups so far, it proves the path has been used but is not yet a stable traffic benchmark.
| layer | behavior |
|---|---|
| Surplus marketplace | Routes a requested model to the cheapest healthy seller. Its current public chat docs do not expose a marketplace cache API, cache-hit header, or discounted exact-answer tier. |
| upstream seller | May maintain a prompt/KV cache. It can reuse a shared prefix while still generating a new answer; availability and pricing depend on the seller. |
| Surp exact cache | Stores the completed answer for an identical deterministic request and skips the seller call entirely on a hit. |
| Surp sticky routing | Keeps the same model briefly to improve locality odds. It does not guarantee the same Surplus seller or worker. |
| Surp affinity research | Collects privacy-preserving prefix/model/provider latency samples. The auction remains experimental and is not a mature routing signal. |
Most gateways optimize either price or caching. Surp combines two separate mechanisms: exact deterministic repeats can skip generation entirely, while brief model stickiness tries to improve upstream prompt-cache locality without accepting an unlimited price gap.
Cache hits create measurable economic value: the network avoids an upstream model call and keeps the difference. Instead of capturing all of that value, surp runs an experimental off-chain reward ledger called SRP:
more usage → more shared cache → lower network costs
↑ ↓
SRP claim on revenue ← more rebates ← more margin
SRP's estimated redemption value is rebate pool ÷ outstanding SRP. As protocol revenue enters the pool, existing SRP gains claim value. Heavy users who create useful cache entries can push their effective marginal cost toward zero over time.
Live pool backing, SRP outstanding, holders, and implied value per SRP are visible on the status page. An agent can query its own balance with GET /api/rewards?payer=0x....
You don't have to trust the default routing. Every request can carry its own constraints without changing the combo:
| parameter | type | effect |
|---|---|---|
max_price_per_1m | float | Reject any model above this USD/1M-token ceiling. Returns 404 if nothing qualifies. |
provider | string or array | Pin routing to one or more providers (allow-list). Cheapest within the set wins. |
surp_bypass_cache | bool | Skip the exact-response cache entirely. Always calls the model, never returns a cached answer. |
surp/strict/<combo> | model prefix | Disable sticky routing for this request. Always picks the live cheapest, never reuses a model for cache locality. |
Example: cheapest coder, but never above $0.50/1M and never from the cache:
curl -X POST https://surp.ivc.lol/v1/chat/completions \\
-H "Content-Type: application/json" \\
-d '{"model":"surp/best-coding",
"max_price_per_1m": 0.50,
"surp_bypass_cache": true,
"messages":[{"role":"user","content":"write a binary search"}],
"max_tokens":100}'
Example: cheapest chat, but only from a specific provider:
curl -X POST https://surp.ivc.lol/v1/chat/completions \\
-H "Content-Type: application/json" \\
-d '{"model":"surp/best-chat",
"provider": "deepseek",
"messages":[{"role":"user","content":"hi"}],
"max_tokens":50}'
These guardrails compose. A request can pin a provider, cap the price, and bypass the cache at the same time. The surp/strict/ prefix layers on top of any combo, including custom ones (surp/strict/my/your-slug).
related: x402 LLM API · cheapest LLM API · x402 gateway · API docs · live cache metrics