S7 AI Relay
Platform capability · AI routing

One endpoint. Every model. Under control.

One router for every AI at Seventy Sete — ask for the capability, the Relay picks the model, meters the cost in µUSD and shows you why.

routable activities

code, chat, summary, vision, reasoning, embeddings, transcription, speech and video

9

effort levels

from low to superior — a lookup key, never a guess

5

endpoint for every provider

no product stores an AI key

1

the unit behind every measurement

whole micro-dollars — money is never floating point

µUSD

// metered routing

Route by activity and effort — and see every µUSD.

Every request matches a rule by (activity, effort), picks the best model across providers and is metered in µUSD. The playground shows the route — and the cost — before you spend a token.

Changing an activity's model is editing one cell of the grid: no product changes code, no deploy happens.

A rule is an ordered chain — when the first one fails, the Relay tries the next within the same request.

Eleven filters judge every candidate, and the trace names which one stopped whom.

// how the Relay decides

From the request to the metered µUSD

The preview and the real call return the same trace — there is no brochure version and a production one.

  1. Requestcapability + effort
  2. Rulematches the activity
  3. Capabilitiesfilters eligible models
  4. Choicebest cost/quality
  5. Fallbackon failure, next in line
  6. MeteringµUSD, tokens and latency in the log

// modalities

Not text alone

Every modality is a routable activity, with the same trace, the same budget and the same metering.

Text, code and reasoning

Chat, long-context summaries, code with tool use and reasoning models — as a single answer or streamed through the local bridge.

Embeddings

From 1 to 96 texts per call. The required capability is enforced by the route, not by whoever wrote the rule remembering it.

Vision

Images arrive as signed references, never as bytes in the body: the Relay fetches within an allowlist and forces the vision capability.

Audio: transcription and speech

Speech becomes text and text becomes speech. Metering counts seconds of audio or characters, in the same µUSD arithmetic.

Avatar video

asynchronous · 202 + poll

A render takes minutes: the submit answers immediately and the poll is the engine. No webhook required, and a ceiling of concurrent renders per Space.

// providers

A closed catalog, switched on by you

The operator switches a provider on, imports the models it advertises in one click and reviews the inferred capabilities. Curated catalog: nobody plugs in an arbitrary adapter.

Anthropic
OpenAI
Google Gemini
OpenRouter
Groq
Qwen
Cloudflare
MiniMax
ElevenLabs
HeyGen
Ollama
// no adapter yetMistral AI · coming soonxAI Grok · coming soon

Each provider's key goes to the platform vault — never to the database, never to the browser. A product talking to the Relay knows no AI key at all.

Local bridge

Your own machine is a provider too

A light CLI pairs the machine, exposes Ollama through an outbound tunnel and starts taking jobs from the Relay — no open port, no public address.

A local model runs at zero cost, in the same catalog as the paid ones.

The answer can grow on screen as it streams, outside the gateway's time ceiling.

“Sensitive data” is a Space policy: switched on, it drops every candidate that is not provably local.

// the console

Seven screens to run all of it

The console is the operator surface: configure, test and account for it. Nothing here needs a deploy.

Overview

Today's requests, cost, p50 latency and fallbacks, with the traffic split by provider.

Playground

A pretend call with a real decision: the route preview contacts no provider and spends no token.

Routing rules

The activity × effort grid, cell by cell, with each rule’s ordered chain and the global limits.

Providers & models

Providers, keys, connection tests and, under each one, its models with capabilities, price and p50.

Local bridge

Pairing by code, the health of every machine and the inventory of local models.

API keys

One key per product or automation, with scopes, test mode and a single reveal.

Usage & cost

The paginated log of every request, the cost by provider and the history export.

// what holds it up

The part that never shows on screen

Everything belongs to the Space

Providers, models, rules, keys and budget live in the partition of the Space that created them. Any member reads; configuring and issuing keys is for admins.

s7k_* keys for machines

Shown once, stored only as a hash, and revoked with effect on the very next request. On the video routes the scope is genuinely enforced.

Metadata only, by default

The log keeps model, tokens, cost and latency — never the prompt nor the answer. Keeping the body is an explicit admin decision, and what is kept expires in 30 days.

A monthly cap per Space

A µUSD limit that joins the decision as one more filter: the cheap local model stays reachable once the expensive one no longer fits.

Secrets in a vault, not in a table

Every (Space, provider) pair owns its own secret. The screen shows the last four characters, and a blank field means “leave it alone”.

Four doors and an MCP server

One rule serves the console, integrations with a user token, machines with a key and the bridge. An agent arrives over MCP and asks for a capability, never a model.

Under construction

Said here before you go looking for it on screen.

  • Audio and video have no playground yet: today those modalities are called through the API and the s7k_* keys.
  • An API documentation page inside the console itself is still missing.
  • Local inference beyond ~18 s in single-answer mode falls through to the next model in the chain — the long path is the playground’s streaming.

Switch a provider on and watch the first route

A new Space is deliberately born empty — no placeholder model that looks configured. You switch a provider on, import the models, point one cell of the grid, and the playground shows the decision before a token is spent.