Your GPU capacity
Model and inference stay on your hardware
Illustrative animation. The agent connects outbound to the gateway, so no inbound port is needed. First bind byoc/<slug> to your local model ID under Model bindings.
The model and inference environment stay on your equipment. The Agent opens the gateway connection, the gateway routes your account’s work to your node, and the app keeps one API entry point.
Model and inference stay on your hardware
Can take over when your node is full or offline, per your fallback settings (platform rates)
No BazaarLink model charge for traffic your own node handles
Same endpoint, no client changes
Choose the local backend that runs your model, then pass its connection settings to the Agent.
Default backend; the default URL is http://localhost:11434.
Use --backend openai and --url with the model ID your local server accepts.
Uses the same OpenAI-compatible streaming path without model-name translation.
You control the model files, inference process, hardware, and local runtime.
For traffic your own BYOC node handles, BazaarLink doesn't charge a model fee; if you also use a platform feature such as web search, that feature is billed at platform prices. Equipment, power and operations remain yours.
The local model connects through the gateway, so product code does not need a transport layer per backend. The node currently receives requests for /v1/chat/completions, /v1/responses, /v1/messages, and image generation (/v1/images).
Work goes only to nodes bound to that account or organization. The node is not placed in a shared pool or compute marketplace.
The Agent forwards image generation requests to your backend's /images/generations.
tools and tool_choice are forwarded to your backend, and tool_calls come back in OpenAI format (OpenAI-compatible and Ollama backends).
Reasoning models return their thinking in the reasoning_content field.
On the Fallback Settings page, choose which BazaarLink API keys this node applies to. When the node is full, offline or errors, the request tries the next source in your fallback order; requests handled by a platform model are billed at platform rates.
Requests your node handles can add :online too, so the model checks the web before answering: the search runs on the platform side and the results go to your node together with the question. The model fee is zero; only the search fee is charged, at platform prices, so your account needs enough balance to cover the search.
You manage the local logs for BYOC; before a request reaches your GPU, the platform entry still applies moderation, and sensitive-data filtering is available as an opt-in.
Platform moderation runs at the API entry (/v1/chat/completions, /v1/responses) before the request is processed; BYOC cannot disable it. Blocked requests do not reach your node.
This is opt-in: organization keys use the organization rule, while personal keys are enabled by the key owner at /content-filter. Once enabled, inbound prompts replace API keys (OpenAI/Anthropic/GitHub/Google/AWS), JWTs, private keys, credit-card numbers, and Taiwan national ID numbers with [REDACTED:TYPE] before the node; prompt injection is blocked (403, code: content_filter).
Enable at /content-filter@bazaarlink/byoc-agent v0.2.0 is MIT licensed. Install the CLI globally from the public GitHub repository; the npm registry version is coming later. Authenticated only confirms the gateway connection. Then open /keys/byok?tab=byoc, map a canonical API ID such as byoc/<lowercase-slug> to the exact model ID accepted by your local backend, and call that canonical ID through the API.
Install the MIT-licensed CLI directly from the public GitHub repository.
npm install -g github:Bazaarlinkorg/bazaarlink-byoc-agentIf you already have one, continue; otherwise use register with your account email.
bazaarlink-byoc register \
--email you@example.com \
--max-concurrent 4login calls /auth/byoc/login and stores the key plus a short-lived token locally.
bazaarlink-byoc login \
--key "byoc_..." \
--gateway "https://byoc-gateway.bazaarlink.ai" \
--input-price 0 \
--output-price 0Authenticated only confirms the gateway connection. Then open /keys/byok?tab=byoc, map a canonical API ID such as byoc/<lowercase-slug> to the exact model ID accepted by your local backend, and call that canonical ID through the API.
bazaarlink-byoc start
# Authenticated only confirms the gateway connection.
# In /keys/byok?tab=byoc, map byoc/<lowercase-slug> to exact backend ID.These fields mirror the current CLI and package metadata. Install from the public GitHub repository until the npm registry version arrives.
The CLI is installed globally from the public GitHub repository; the npm registry version is coming later.
--backend ollama is the default. Add --ollama-url when Ollama runs elsewhere.
LM Studio, vLLM, and llama.cpp use --backend openai; --url is required and After connecting, map a canonical API ID to the model ID accepted by your backend in the BYOC bindings page.
You have a GPU host and local model, and want the model files and inference environment to stay on equipment you control while your product keeps a familiar API entry point.
You do not want to rent out idle compute or place requests into a shared node pool. BYOC connects your equipment, your account, and one gateway path.
Bring one local model machine and a BYOC key, then validate how your compute, your node, and one API entry point work together.
Manage BYOC nodes