Skip to main content

Governed LLM Gateway

The Governed LLM Gateway exposes an OpenAI-compatible inbound API so any existing app can point its base_url at Viglet Turing ES with zero code change, and, in doing so, transparently gain the governance and intelligence Turing already provides: virtual keys, per-key budgets and rate limits, spend tracking, a semantic cache, guardrails, model-name routing into your agents and Semantic Navigation sites, and provider-native tools.

The gateway is opt-in and off by default. Enable it with:

turing.gateway.enabled=true

Endpoints

Base URL: https://<your-turing-host>/v1

Method & pathPurpose
POST /v1/chat/completionsChat completions: blocking, or streamed when "stream": true (OpenAI data: chunks, terminated by data: [DONE])
POST /v1/embeddingsEmbeddings
GET /v1/modelsLists the models the gateway advertises (every enabled LLM instance)

Because the wire format is OpenAI's, the official SDKs work unchanged:

from openai import OpenAI

client = OpenAI(
base_url="https://your-turing-host/v1",
api_key="sk-turing-...", # a Turing virtual key
)
resp = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
)
curl https://your-turing-host/v1/chat/completions \
-H "Authorization: Bearer sk-turing-..." \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o", "messages": [{"role": "user", "content": "Hello"}]}'

Virtual keys

Every caller authenticates with a virtual key (sk-turing-...), created and managed by an admin under /bento/gateway. A key is shown once at creation: copy it then; only a masked prefix is stored for display.

Each key can be scoped and governed:

SettingEffect
Allowed modelsComma-separated allow-list; blank = any model. A request for an out-of-scope model is rejected.
Monthly budget (USD)Soft cap: once month-to-date spend reaches it, the key auto-downgrades to a cheaper model (if one is configured) instead of the requested one.
Hard cap (USD)Hard kill-switch: once reached, requests are refused with HTTP 429 before any upstream LLM call.
Rate limit (req/min)Per-key request rate limit; excess requests get HTTP 429.
ExpiryA key past its expiry no longer authenticates.

Rotate a key to issue a fresh secret (the old one stops working); revoke to disable it immediately. All inbound spend is metered per key and shown in the /bento/gateway spend dashboard.

Model-name routing

The model string routes the request into the Turing stack:

model valueBehaviour
gpt-4o, claude-…, any instance id/namePassthrough to that provider (parity)
turing-agent:<id>Runs the full AI Agent: its system prompt, tools, MCP servers and grounding mode
turing-sn:<site>A RAG-grounded answer over a Semantic Navigation site (retrieval + citations + reranker)
turing-local:*Forces the embedded, local ($0) model
turing-router:<strategy>/<id1,id2,…>Load-balances across several deployments with a fallback chain. Strategy: rr (round-robin, default), latency (least observed latency), failover (declared order), or the catalog-driven cost (cheapest), quality (highest intelligence index) or ttft (lowest time-to-first-token). Example: turing-router:cost/inst-a,inst-b picks the cheapest first, the rest as fallback

The cost / quality / ttft strategies order the declared deployments by the public model catalog's pricing / intelligence index / time-to-first-token for each deployment's configured model. A deployment whose model has no catalog value for the chosen metric sorts last, so a missing figure never wins a "cheapest" or "fastest" pick.

Cross-cutting headers

Stack these on any model to add behaviour without changing your payload:

HeaderEffect
x-turing-rag-site: <site>RAG-as-a-header: retrieves the site's top passages and grounds the answer, on any base model
x-turing-cache: semanticServes an identical prompt from an in-memory response cache (no upstream call, no spend)
x-turing-guardrails: strictModerates the request input (and, on the blocking endpoint, the output); flagged content is refused
x-turing-tools: web_search,web_fetch,code_executionEnables the provider's native tools on a plain chat call, the capabilities must first be enabled on the model instance by an admin

Example, a legacy OpenAI app gains RAG grounding by changing its base URL and adding one header:

curl https://your-turing-host/v1/chat/completions \
-H "Authorization: Bearer sk-turing-..." \
-H "x-turing-rag-site: wknd" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o", "messages": [{"role": "user", "content": "What is WKND?"}]}'

Admin

The /bento/gateway console page (admin only) provides virtual-key CRUD (create / rotate / revoke with scope + budget), a per-key spend dashboard, and a copy-paste playground snippet. Inbound spend also flows into the standard Cost Governance view, keyed by virtual key.

Learning from traffic (optional)

With turing.gateway.capture-traffic=true, the gateway records inbound base-model and router turns as rows of a reusable evaluation dataset (prompts as the replay turns; the answer kept as reference material, asserting nothing until a human curates it). That dataset runs through the standard evaluation graders to measure quality on real traffic, and an admin can request a propose-only distilled/cheaper-model suggestion for an agent (POST /api/gateway/traffic/distill?agentId=…, gated by turing.distillation.enabled): the recommendation is never applied automatically. Capturing prompts/answers is off by default and admin-managed.

Notes & limits

  • The gateway targets the subset of the OpenAI wire contract that most clients use (chat, embeddings, streaming and tools) not full wire fidelity.
  • It is additive and default-off: nothing changes for non-gateway paths.
  • Native-tool turns (x-turing-tools) are not token-metered, since the provider-native services surface no usage.