// metered routing
Route by activity and effort — and see every µUSD.
Every request matches a rule by (activity, effort), picks the best model across providers and is metered in µUSD. The playground shows the route — and the cost — before you spend a token.
Changing an activity's model is editing one cell of the grid: no product changes code, no deploy happens.
A rule is an ordered chain — when the first one fails, the Relay tries the next within the same request.
Eleven filters judge every candidate, and the trace names which one stopped whom.
// how the Relay decides
From the request to the metered µUSD
The preview and the real call return the same trace — there is no brochure version and a production one.
- Requestcapability + effort
- Rulematches the activity
- Capabilitiesfilters eligible models
- Choicebest cost/quality
- Fallbackon failure, next in line
- MeteringµUSD, tokens and latency in the log
// modalities
Not text alone
Every modality is a routable activity, with the same trace, the same budget and the same metering.
Text, code and reasoning
Chat, long-context summaries, code with tool use and reasoning models — as a single answer or streamed through the local bridge.
Embeddings
From 1 to 96 texts per call. The required capability is enforced by the route, not by whoever wrote the rule remembering it.
Vision
Images arrive as signed references, never as bytes in the body: the Relay fetches within an allowlist and forces the vision capability.
Audio: transcription and speech
Speech becomes text and text becomes speech. Metering counts seconds of audio or characters, in the same µUSD arithmetic.
Avatar video
asynchronous · 202 + pollA render takes minutes: the submit answers immediately and the poll is the engine. No webhook required, and a ceiling of concurrent renders per Space.
// providers
A closed catalog, switched on by you
The operator switches a provider on, imports the models it advertises in one click and reviews the inferred capabilities. Curated catalog: nobody plugs in an arbitrary adapter.
Each provider's key goes to the platform vault — never to the database, never to the browser. A product talking to the Relay knows no AI key at all.
Your own machine is a provider too
A light CLI pairs the machine, exposes Ollama through an outbound tunnel and starts taking jobs from the Relay — no open port, no public address.
A local model runs at zero cost, in the same catalog as the paid ones.
The answer can grow on screen as it streams, outside the gateway's time ceiling.
“Sensitive data” is a Space policy: switched on, it drops every candidate that is not provably local.
// the console
Seven screens to run all of it
The console is the operator surface: configure, test and account for it. Nothing here needs a deploy.
Overview
Today's requests, cost, p50 latency and fallbacks, with the traffic split by provider.
Playground
A pretend call with a real decision: the route preview contacts no provider and spends no token.
Routing rules
The activity × effort grid, cell by cell, with each rule’s ordered chain and the global limits.
Providers & models
Providers, keys, connection tests and, under each one, its models with capabilities, price and p50.
Local bridge
Pairing by code, the health of every machine and the inventory of local models.
API keys
One key per product or automation, with scopes, test mode and a single reveal.
Usage & cost
The paginated log of every request, the cost by provider and the history export.
// what holds it up
The part that never shows on screen
Everything belongs to the Space
Providers, models, rules, keys and budget live in the partition of the Space that created them. Any member reads; configuring and issuing keys is for admins.
s7k_* keys for machines
Shown once, stored only as a hash, and revoked with effect on the very next request. On the video routes the scope is genuinely enforced.
Metadata only, by default
The log keeps model, tokens, cost and latency — never the prompt nor the answer. Keeping the body is an explicit admin decision, and what is kept expires in 30 days.
A monthly cap per Space
A µUSD limit that joins the decision as one more filter: the cheap local model stays reachable once the expensive one no longer fits.
Secrets in a vault, not in a table
Every (Space, provider) pair owns its own secret. The screen shows the last four characters, and a blank field means “leave it alone”.
Four doors and an MCP server
One rule serves the console, integrations with a user token, machines with a key and the bridge. An agent arrives over MCP and asks for a capability, never a model.
Under construction
Said here before you go looking for it on screen.
- Audio and video have no playground yet: today those modalities are called through the API and the s7k_* keys.
- An API documentation page inside the console itself is still missing.
- Local inference beyond ~18 s in single-answer mode falls through to the next model in the chain — the long path is the playground’s streaming.