Case study, 2026
Open Switch
A Rust LLM gateway with deterministic routing simulation, provider fallbacks, per-key budgets, and streamed OpenAI-compatible responses.
01
The routing decision is a moving target.
A deployment can look cheap and healthy until a run of failures or a latency spike changes the tradeoff. The routing core is deterministic and separate from HTTP, so the browser demo can exercise the same decision model without contacting a provider.
The axum gateway accepts OpenAI-compatible chat requests, lists configured models, emits SSE for streaming responses, and exposes plain-text request and latency metrics.
02
Failure has a state machine.
A deployment breaker starts closed, opens after consecutive failures, and later admits a limited half-open probe set. Successful probes return it to closed. The simulation exposes that state for every deployment after a run.
Latency is an exponentially weighted moving average. Unlike a fixed window, the alpha makes the reaction speed explicit. Cost-aware routing selects the cheapest healthy choice within a configured latency budget.
03
Protect the slow tail with a second attempt.
When the first request crosses the hedge delay, the gateway starts the next deployment and uses the first completed answer. Tokio cancels the losing attempt when the request future is dropped. The simulation reports hedge count and modeled tail latency saved.
Bearer keys have token ceilings per minute. Model groups and prices stay in deployment configuration, and mock, OpenAI-compatible, and Anthropic adapters share the request shape at the gateway boundary.
04, Measured
Small enough to inspect.
38
tests passing in the workspace
212 ns
per routing decision over 100,000 release iterations
578,428
in-process gateway requests per second at concurrency 64
0.11 / 0.23 ms
added gateway latency, p50 / p99, over 20 seconds
101,983 B
optimized browser WebAssembly binary