LangChain Free LLM API Setup: Python and TypeScript with BazaarLink
Configure ChatOpenAI with BazaarLink in Python or TypeScript, build a small LCEL chain, and check tool support, shared free quotas and streaming usage before deployment.
For a basic LangChain chat call, set three things explicitly: the BazaarLink base URL, your API key and model="auto:free". Check that small request before adding tools, memory or an agent loop. A working chat completion does not establish support for every feature in ChatOpenAI.
The Python and JavaScript setup below follows the official Python ChatOpenAI integration and JavaScript integration, checked on 17 September 2026. These are configuration examples, not a claim that a particular package version and live BazaarLink model were tested together in this revision.
Create a key and install only the packages you need
Sign in at /login, then create an API key under /keys. Registration does not require a payment method. Set BAZAARLINK_API_KEY privately in the shell that will run your application; do not commit its value.
For Python, the chat integration is a separate package:
python -m pip install -U langchain-openai
For a Node/TypeScript application:
npm install @langchain/openai @langchain/core
Once the initial request works, lock the dependency versions in your project. Save the Python package list or npm lockfile alongside your test results. Installing “latest” is convenient for an example, but it is not a reproducible production dependency policy.
Python: a minimal ChatOpenAI call
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="auto:free",
api_key=os.environ["BAZAARLINK_API_KEY"],
base_url="https://api.bazaarlink.ai/v1",
use_responses_api=False,
max_retries=0,
)
message = llm.invoke("Explain an API rate limit in one sentence.")
print(message.content)
use_responses_api=False makes the intended Chat Completions transport explicit. max_retries=0 keeps an initial failure visible and avoids repeated calls while you are debugging. Decide the production retry policy separately.
base_url and api_key are current documented parameters. Older aliases such as openai_api_base may still be supported, but mixing old and new recipes makes diagnosis harder. The legacy https://bazaarlink.ai/api/v1 endpoint also remains supported.
TypeScript: put baseURL inside configuration
import { ChatOpenAI } from "@langchain/openai";
const apiKey = process.env.BAZAARLINK_API_KEY;
if (!apiKey) throw new Error("Set BAZAARLINK_API_KEY before running this example.");
const llm = new ChatOpenAI({
model: "auto:free",
apiKey,
maxRetries: 0,
configuration: {
baseURL: "https://api.bazaarlink.ai/v1",
},
});
const message = await llm.invoke("Explain an API rate limit in one sentence.");
console.log(message.content);
Run this in your existing TypeScript environment with ESM/top-level-await support, or wrap the invocation in an async function. The connection setting belongs under configuration.baseURL; changing only a generic environment variable can leave an application using a different default endpoint.
Build a small LCEL chain after chat works
LangChain's prompt and output-parser components compose around the configured model. This Python sketch reuses llm from above:
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
prompt = ChatPromptTemplate.from_template(
"Summarize this text in one sentence: {text}"
)
chain = prompt | llm | StrOutputParser()
result = chain.invoke({"text": "Our test application sends one chat request at a time."})
print(result)
This checks prompt construction and one model invocation. It does not test agent tools, JSON Schema enforcement, image inputs or embeddings. Add each required capability as a separate test so an unsupported feature does not get mistaken for an authentication problem.
Tool calling needs a model-level check
ChatOpenAI exposes tool binding, but the target model must accept the schema and return usable tool calls. auto:free chooses from the current free pool, so it is not a guarantee of a fixed capability set.
For a tool-using application, my recommendation is to choose a specific ID from /models, then test one harmless tool with known input and output. Inspect the actual returned tool call and arguments, execute the tool, and send its result back. A natural-language answer saying “I used the tool” is not proof.
Neither choosing a paid model nor setting temperature to zero guarantees deterministic behavior. Keep tool permission checks and argument validation in application code. For current agent construction APIs, use the LangChain agent documentation that matches your locked version; this article does not reuse the old AgentExecutor recipe as if it were a current tested setup.
Free quota: count model calls, not chain runs
BazaarLink's live free-tier page currently shows a base 10 requests/minute and 50 requests/day, with ×1/×2 account-tier multipliers. The actual tier depends on current balance/subscription eligibility and configured thresholds. Free-model traffic shares that budget across the account. A chain with one invocation uses one model call; an agent task with several tools and retries can use many.
After the quota, requests on free-quota models can continue at normal paid rates when the account meets paid-fallback conditions; otherwise they are rate-limited. Check the free-model rules and usage dashboard before running CI batches. Adding credit does not turn the free pool into unlimited free inference.
| Failure | Useful next check |
|---|---|
| Authentication rejected | Key status and the configured base URL; never log the credential |
| Model not found | Exact ID in the live catalog, not a copied historical example |
| Rate limit | Shared quota and actual calls per chain, including retries |
| Tool/schema error | Model feature support and the exact request being generated |
| Stream text works but usage is missing | Provider support for streaming usage, not just text streaming |
LangChain documents separate streaming-usage settings: Python stream_usage and JavaScript streamUsage. Enable or disable them according to the endpoint contract rather than assuming every proxy supports the same option.
Before deployment, record package versions, model ID, transport, successful call count and the required features you actually exercised. That gives you a useful regression check when a package or model changes. For the broader free-inference tradeoffs, see the free API comparison.
FAQ
How do I use a free LLM API with LangChain?
Create a BazaarLink key and configure ChatOpenAI with your key, model auto:free and base_url https://api.bazaarlink.ai/v1 in Python. In TypeScript use configuration.baseURL. First test a basic chat request; free-model quotas and capabilities still apply.
Does auto:free guarantee LangChain agent tool calling?
No. LangChain can generate tool requests, but the selected model must support the required schema and tool-call behavior. Choose and test a specific model when the application depends on those features.
How many free LangChain runs can I make?
Count model calls, not application runs. The live base allowance is currently 10 requests/minute and 50/day with account multipliers, shared across free-model traffic. Tools, retries and multiple chain invocations can consume several requests per run.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API