Switchyard routes each LLM call to the cheapest model that can still do the job. Without changing a line of your agent.
*Total cost based on average ISP token cost
Switchyard picks which model serves each LLM call.
Switchyard runs inside gateways you may already have.
- NeMo Relay — a native plugin. Load a
routes.tomlinto a Relay deployment you already run. Setup → - LiteLLM — a routing plugin for LiteLLM's
Routerand proxy. Setup → - More integrations coming soon.
flowchart LR
subgraph R["LiteLLM · NeMo Relay"]
P["Switchyard"]
end
P--> M["Efficient model"]
P--> N["Capable model"]
P--> O[etc.]
G[You] -->|"request"| P
style P fill:#76B900,stroke:#5A8F00,color:#000
Embed the routing algorithms in your own. Switchyard picks the model; your harness makes the call, so your transport, retries, and credentials stay untouched.
- Install:
pip install nemo-switchyard - Then follow Path 2 — Embed the Library: construct an algorithm, drive its step stream, make the answer call.
- Also available for Rust as
switchyard-libsy; Path 2 has theCargo.tomlblock.
flowchart LR
subgraph R["Your LLM gateway / harness"]
P["Switchyard"]
end
P--> M["Efficient model"]
P--> N["Capable model"]
P--> O[etc.]
G["Your users"] -->|"request"| P
style P fill:#76B900,stroke:#5A8F00,color:#000
A server in front of an agent, when you have no gateway to put Switchyard in. Point Claude Code, Codex CLI, or any OpenAI/Anthropic SDK client at it; Switchyard decides per turn which model serves it.
- Install:
cargo install --locked switchyard-server - Then follow Path 3 — Run the Standalone Proxy:
write
routes.toml, start the server, point your agent at it.
flowchart LR
P["Switchyard<br/>standalone proxy"]
P--> M["Efficient model"]
P--> N["Capable model"]
P--> O[etc.]
G[You] -->|"unchanged native API"| P
style P fill:#76B900,stroke:#5A8F00,color:#000
Pre-1.0 software. APIs, configuration, and routing behavior can change between releases — pin the version you integrate.
| Component | Stability | Use it for | Guidance |
|---|---|---|---|
switchyard-libsy |
Beta | Routing embedded in your own gateway or harness. You own model calls, credentials, and retries. | Trial integrations. API will change before v1.0. |
switchyard-llm-client |
Alpha | HTTP model calls and protocol translation alongside libsy. | Experiments and pilots. |
switchyard-runner |
Alpha | Running configured routes inside another runtime, such as NeMo Relay. | Integration work and supervised pilots. |
switchyard-server |
Demo | A standalone OpenAI- and Anthropic-compatible proxy. | Demos and evaluation only. Not for production. |
Three paths, in the same order as above. Using Claude Code or Codex? Point it at this README and ask it to set up the path you want.
You finish with an existing NeMo Relay deployment routing through Switchyard.
Requires NeMo Relay >=0.8.0, <1.0.0 and a Rust toolchain.
Follow the plugin README's
Install and
Configure Relay
sections. For the deployment file, use the routes.toml from
Path 3, step 2.
You finish with your own harness picking a model per request and still making every model call itself. Shown in Python; the Rust API has the same shape.
1. Install.
pip install git+https://github.com/NVIDIA-NeMo/Switchyard.gitThe API below is newer than nemo-switchyard 0.2.0 on PyPI, so install from
source until the next release. Rust: depend on switchyard-libsy and
switchyard-protocol from this repository instead. Pin both to the commit you
tested — @<sha> for pip, rev = "<sha>" for Cargo — before depending on them.
2. Construct an algorithm. It selects a category — efficient or
capable — and you map categories to model IDs when each request runs.
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router
algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)3. Drive it. run_stream yields steps. Serve each CallModel with your own
client — call.models is ordered by preference, and call.fail(error) reports
a failed call; Done carries the pick.
models = {"efficient": ["fast"], "capable": ["quality"], "any": ["quality", "fast"]}
async for step in algorithm.run_stream(request, models):
match step:
case Step.CallModel(call):
call.respond(LlmResponse.Agg(await my_client(call.request, call.models[0])))
case Step.Done(outcome):
model, request = outcome.selected_model_ids[0], outcome.request4. Make the answer call with model and request, using your own HTTP
client, retries, and credentials.
The complete runnable version — streaming and a working client — is
examples/libsy.py. Types:
switchyard-libsy,
switchyard-protocol.
You finish with a server on localhost:4000 that any OpenAI or Anthropic client
can call. Needs Rust with Cargo.
For v0.3.0, the standalone server is release-validated on Ubuntu 24.04,
Linux x86_64. Other platforms are outside the release-validation scope.
1. Install the server.
cargo install --locked switchyard-server2. Write routes.toml. A stage router over the same model pair as the
benchmark above: how to reach a provider, which models to use, how to choose
between them. --config takes any path; this writes it to the current directory.
cat > routes.toml <<'TOML'
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"
[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"
[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
TOMLEvery key is documented in the TOML schema reference.
3. Start it. --dry-run loads the config, prints server OK: and the model
IDs it exposes, then exits without starting the server.
export OPENROUTER_API_KEY="your-openrouter-key" # pragma: allowlist secret
switchyard-server --config routes.toml --dry-run
switchyard-server --config routes.toml --host 127.0.0.1 --port 40004. Send a request. The route's id is the model name clients ask for.
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"switchyard","messages":[{"role":"user","content":"hello"}]}'The same route also answers on /v1/messages (Anthropic Messages) and
/v1/responses (OpenAI Responses). /v1/stats reports which target served
what, and /metrics exposes Prometheus counters for requests, errors, latency,
tokens, and routing overhead.
5. Point a coding agent at it.
export ANTHROPIC_BASE_URL="http://localhost:4000"
export ANTHROPIC_MODEL="switchyard"
export ANTHROPIC_API_KEY="placeholder" # pragma: allowlist secret
claudeThe placeholder satisfies Claude Code's client-side auth check. In this local
setup, Switchyard uses the server's OPENROUTER_API_KEY for upstream requests.
Do not use the placeholder with forward_auth = true or a gateway that requires
a real client credential.
Codex CLI and other OpenAI clients use the OpenAI variables instead:
export OPENAI_BASE_URL="http://localhost:4000/v1"Start with Auto. Choose Task or Execution when you want more control over how requests move between an efficient model and a capable one.
| Choice | Use it when | Route type |
|---|---|---|
| Auto | You want Switchyard's recommended preset. | auto |
| Task | You want an LLM to judge which model can handle the task. | llm_classifier |
| Execution | You want tool results and agent progress to guide each request. | stage_router |
Auto currently uses Execution with fixed defaults and no LLM judge. Task uses the LLM classifier's capability mode. Execution uses the stage router. These names do not change the TOML configuration keys.
Auto requires a source build until v0.3.0 is published. The quickstart above uses
stage_routerdirectly.
See the full routing catalog for composite routing, escalation, custom policies, and other options. Performance results are listed under Benchmark Provenance.
- Core Concepts: LLM clients, targets, routes, model IDs, and routing algorithms
- Routing Overview: choose and configure a routing algorithm
- TOML Schema: every configuration key
- Architecture: how the proxy and library components fit together
- switchyard-server: server configuration, routing algorithms, and metrics
- switchyard-libsy: embed routing algorithms in a Rust application
- switchyard-protocol: provider-neutral request, response, and streaming types
- switchyard-translation: request, response, and stream translation
- switchyard-nemo-relay-plugin: install Switchyard as a native NeMo Relay plugin
| Configuration | Accuracy | Total cost | vs. Opus 4.8 baseline |
|---|---|---|---|
| Opus 4.8 baseline | 76.0% | $98.06 | — |
| Escalation | 75.7% | $85.00 | 99.6% of accuracy, 13.3% cheaper |
| Execution (Stage) | 72.7% | $68.19 | 95.7% of accuracy, 30.5% cheaper |
| Task (Capability) | 71.2% | $79.32 | 93.7% of accuracy, 19.1% cheaper |
| Kimi K2.6 alone | 55.8% | $76.28 | |
| GLM 5.2 alone | 52.4% | $16.47 | |
| DeepSeek V4 Pro alone | 48.7% | $96.92 | |
| Ultra 3 alone | 39.0% | $29.66 |
These are the v0.2.0 Terminal-Bench 2.1 results from Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard. Those runs used NVIDIA-internal inference endpoints, so absolute solve rates may shift on another serving stack; the routing parameters are the ones that ran.
The escalation deployment is checked in at
benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml,
with OpenRouter targets substituted so it is publicly runnable. To run the
harness, see benchmark/README.md; for latency and
routing overhead rather than task success, see
Soak Testing.
- Issues: GitHub Issues
- Code of Conduct: Code of Conduct
Apache 2.0 License. Copyright NVIDIA Corporation.
