Swap models without code changes
Your code only says vm/<slug>. Change the main model, add a fallback or reorder in the dashboard and it applies immediately.
If a step fails (timeout, 429, 5xx, source unavailable) the next one takes over. Only the step that succeeds is billed.
Illustrative animation. 429, 5xx, timeouts or an unavailable source move to the next step; other client errors stop; nothing is retried once output has started. Steps that use platform credits require your consent.
Your app: model: "vm/writer" → vm/writer Next step only on failure
openai/gpt-6.1-sol Your own OpenAI key firstqwen/qwen3.8-27b Then your own GPUopenai/gpt-6-luna BazaarLink as the safety netYour code only says vm/<slug>. Change the main model, add a fallback or reorder in the dashboard and it applies immediately.
BYOK keys and BYOC nodes go first; platform credits step in only when needed, so cost stays lowest.
Sol overloaded? Fall back to Luna. Cloud outage? Fall back to your own GPU. Every step can use a different model.
Usage and spend are grouped by vm name, so each team or product line sees its own cost.
/v1/chat/completions, /v1/responses and /v1/messages — keep your OpenAI and Anthropic SDKs.
Accounts and organizations each have their own namespace; org members share the same virtual models.
Temperature, reasoning effort and system prompts pass through as sent; the virtual model only decides who answers.
Client errors such as 4xx are returned, not retried, and a stream that has started is never switched mid-way.
Open Virtual Models in the dashboard and pick a name such as writer.
Choose a source (BYOK, BYOC or platform credits) and a model for each step, then reorder.
Nothing else in your code changes.
Your first virtual model takes about a minute.
Create a virtual model