BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-07-25 · · Author:BazaarLink · relay detection · Claude downgrading · full-strength Claude · AI API security

What Is 降智 (Model Downgrading)? 3 Detection Methods

Jiangzhi (降智) means a relay swaps your paid premium model for a cheaper one. Here are three detection methods and why gut feelings are unreliable.

If the Claude quota you bought suddenly "got dumber" (the same prompt yields lower code quality, skipped reasoning steps, or long summaries that start dropping sections), the community calls this phenomenon "downgrading" (降智). Downgrading is not necessarily your imagination, but it is not necessarily the relay's fault either. This article explains how to tell the difference, and three detection methods you can actually run.

Downgrading, mixing, full-strength: define the terms first

The community uses these three terms interchangeably, but they refer to different situations:

TermActual meaningDetection difficulty
Downgrading (降智)The premium model you bought is replaced with a cheaper one. You pay for Opus but get Haiku, GLM or DeepSeekMedium: a persistent problem; an A/B test shows it
Mixing (摻假)Most requests to the same endpoint go to the real model, but a few are secretly swappedHigh: single tests pass, so statistics are needed to catch it
Full-strength (滿血)The model really is the one claimed, with no interference from extra system prompts or tool restrictions—

"Mixing" is far more dangerous than "downgrading." An endpoint that fails every time gets dropped within three minutes; an endpoint that succeeds intermittently can run for months, and each time something goes wrong you first suspect your own prompt.

Why gut feelings are unreliable

Three reasons:

1. Self-declaration can be forged. A relay only needs to put one sentence in the system prompt, "You are Claude Opus, developed by Anthropic," and any open-source model will go along with it. The answer to "who are you" has no evidentiary value.

2. The official side also changes. Adjustments to default thinking effort and issues at the SDK layer can make everyone feel the model got dumber at the same time. Switching relays does not help in that case, because the problem is not the relay at all.

3. Style fluctuation causes false positives. The output style of the same model drifts with different temperatures and context lengths. Concluding that the model was swapped based only on "this answer was worse this time" has a very high false-positive rate, which is why rigorous testing must use distribution comparison rather than single samples.

Method 1: Knowledge cutoff and capability boundary comparison

Pick a set of questions that only the target model can answer correctly and where the answer is unique. Events near the knowledge cutoff date, long calculations in a specific format, and obscure syntax in a specific language all fall into this category.

The key is that the answers must be unique and verifiable, not a subjective comparison of "which answer is better." If the endpoint's accuracy on this set is clearly below that of the official direct connection, the model is very likely not the one claimed.

Limitation: this method can tell apart "different families" (Claude vs GLM), but it struggles to tell apart "different models within the same family" (Opus vs Haiku), because models in the same family share highly overlapping knowledge.

Method 2: Behavioral fingerprinting (model fingerprinting)

Compare the statistical distribution of responses rather than the content of a single response. When facing boundary questions, the same model has reproducible distribution features in response length, refusal tendency, and preference for certain tokens. These features are much harder to forge than the content itself.

This is the method that can distinguish sub-models within the same family, but it needs a trustworthy official baseline for comparison and enough samples to be statistically meaningful. Doing it by hand is not realistic, which is the reason detection tools exist.

Method 3: Hidden system prompt detection

Many cheap Claude quotas are reverse-proxied from the internal interfaces of other products. These interfaces come with a hidden system prompt that restricts the model to that product's purpose.

In our actual testing we measured such an endpoint: each request carried a hidden system prompt of about 2,000 tokens, the model claimed to be another product's name, and it refused all non-programming questions, even though it advertised itself as general-purpose Claude Opus. In Claude Code, this endpoint shows up as "getting dumber," because your instructions keep fighting that hidden prompt.

The detection method is to measure the gap between the "claimed input token count" and the "actual billed token count," and to observe whether the model shows unexpected behavioral constraints under controlled prompts.

By the way: downgrading often comes with extra token billing

Endpoints that alter the model usually also alter the usage field. The motives are the same (cut costs, raise margins), and the mechanism works at the same layer. See token padding detection for details.

Run a detection yourself

We have automated all three methods above as tests. Paste the endpoint's Base URL and key into relay detection, and it runs the full decision chain and outputs a 0 to 100 score, including family identification, same-family sub-model re-verification, hidden system prompt measurement, and token calculation checks. The decision process is public: the page shows the complete decision tree, so you can see which branch produced each conclusion.

One principle is worth remembering: any detection result that does not tell you its basis is not worth trusting. Including ours.

FAQ

What is the difference between Claude downgrading and official adjustments?

Both exist and should be looked at separately. Official changes (such as default thinking-effort adjustments or SDK harness issues) affect everyone at once, and the official side announces them or they can be traced through release notes. Downgrading caused by a relay affects only people using that endpoint: if the same prompts are fine when sent directly to the official API but worse through the relay, it is an endpoint problem. The way to judge is to run an A/B test with the same set of questions against both.

Can't I just ask the model what model it is?

No, that does not work. A relay only needs to write "You are Claude Opus" in the system prompt, and any model will answer accordingly. Self-declaration is the easiest signal to forge, so rigorous testing never trusts the seller's or the model's own claims and looks only at behavioral fingerprints.

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
claude plans · opencode go · subscription · pay-as-you-go API · AI API pricing
Subscription vs Pay-Per-Token API: Claude Pro, OpenCode Go
Usage rebate · Usage Rebate · AI API fees · OpenAI GPT · Gemini · DeepSeek · Enterprise AI API
Usage Rebate Rules: How BazaarLink Milestone Credits Work
AI gateway · AI API Gateway · AI Gateway · LLM Gateway · model router · relay · BYOK · upstream failover
AI API Gateway vs Router vs Relay: Differences and Choices
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.