BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-07-25 · · Author:BazaarLink · Relay detection · Claude Code relay · reverse proxy · AI API security

Does Claude Code Get Dumber on Relays? Hidden Prompt Test

Claude Code getting dumber on a relay is usually not the model. A hidden system prompt in reverse-proxy endpoints is to blame. Here is how to verify it.

"Claude Code has gotten dumber lately" is a frequent complaint in the community. If you are using a relay, before suspecting the model, rule out a more common cause first: the hidden system prompt built into reverse-proxy endpoints.

Where Cheap Claude Quota Comes From

Claude quota below official pricing generally comes from three kinds of sources:

SourceCharacteristicsImpact on Claude Code
Reverse-proxied from another AI product's internal endpointLowest price; comes with a hidden system promptObvious: instructions conflict with the hidden prompt
Reverse-proxied from a pool of subscription accountsMid-range price; has rate and concurrency limitsModerate: mainly intermittent failures and queuing
Legitimate API resaleClose to official price; high multiplierNone

The first type is the main source of "getting dumber" complaints, and its problem is not that the model has been swapped out, but that the model has been constrained.

Why a Hidden System Prompt Makes It Seem Dumb

The endpoints being reverse-proxied were originally internal endpoints of a particular product. The product uses a system prompt to confine the model to its own use case: answering only programming questions, outputting in a specific format, or playing a particular role. The reverse proxy simply exposes that endpoint for resale, and the prompt does not go away.

The result is that every one of your requests is tugging against an invisible set of instructions:

  • Your system prompt is down-weighted, because an earlier and stronger instruction sits in front of it
  • Context budget is consumed. We have measured cases where roughly 2,000 tokens were added to each request; in a long conversation, that is a real loss of usable context
  • Tool-call behavior changes. The original product may have restricted the available toolset, so Claude Code's tool use becomes unstable
  • Non-programming requests may be refused outright

From the user's side, all of this shows up as "getting dumber." But the model may really be Claude; it is just not free.

How to Tell an Official Problem from an Endpoint Problem

Run this A/B test, and do not reverse the order:

  1. Run a fixed set of questions (including long-context and tool-call tasks) against the official direct connection
  2. Run the same set of questions against your relay endpoint
  3. Compare the differences

If both sides get worse, it is an official-side change, and switching relays will not help. If only the relay side is worse, the problem is the endpoint.

Skipping step 1 and switching relays directly is the most common waste in the community. Switching to a different relay every month, with each one seeming "better, then worse again," is usually just chasing random fluctuation.

Some Things Self-Testing Cannot Detect

Asking the model "what model are you?" does not help, because the system prompt can dictate how it answers. Response speed is not very reliable either: we have seen an endpoint with a TTFT of only 182 ms whose body was completely empty, with headers sent but zero chunks. A TTFT value does not mean the endpoint is normal; it may instead be a signature of the reverse-proxy layer.

To measure the hidden system prompt, you have to compare the gap between the claimed input token count and the actual billed token count, and observe whether the model shows unexpected behavioral constraints under controlled prompts. This cannot be done by hand.

Run the Verification Once

Paste the endpoint's Base URL and key into the relay detection tool. Leak detection measures the size of any injected system prompt, and fingerprint classification tells you the true model family and version behind the endpoint. If it claims Opus but the fingerprint points to a smaller model, the report will flag that directly.

The decision logic is public, and the page shows the full decision tree, including conservative branches such as "abstain when evidence is insufficient, and identify only the family, not the exact version." A detection that tells you "uncertain" is more credible than one that always gives a definitive answer.

Further reading: Three Methods for Detecting Claude Degradation and Token Padding Detection.

FAQ

Why do reverse-proxied endpoints come with a system prompt built in?

Those endpoints were originally internal endpoints of a particular product, not general-purpose model endpoints. The product uses a system prompt to confine the model to its own use case, for example answering only programming questions or playing a specific role. A reverse proxy simply exposes that endpoint and resells it, and the prompt does not disappear. It becomes noise that rides along with every one of your requests.

Can the relay see the code I send through Claude Code?

Technically, yes. When Claude Code works, it reads large amounts of local file content and sends it with each request, and that content is visible to the relay along the transmission path. This is an inherent risk of using any relay service, regardless of whether the model is genuine. Projects involving sensitive code should factor this into their evaluation.

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
Claude Code · Claude Code update · changelog · mods · permission rules
Claude Code 2.1.289: 27 Changes, Permission-Rule Fixes and Mods Stability
Claude Code · Claude Code mods · plugins · TypeScript · AI coding tools
Claude Code Mods: Customize Behavior and UI with TypeScript
web search API · AI agent · LLM · Hermes Agent
Hermes Agent Web Search: Use a BazaarLink Model with :online
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.