What Is Token Padding? How to Detect It in AI APIs
Token padding means an AI API relay reports inflated usage, so you pay for more tokens than were consumed. Learn the signs and check with BazaarLink Probe.
Have you ever wondered why the same prompt sent to different services shows very different token counts on your bill? You may not be imagining it. Some AI API relays use "token padding" to report false numbers in the usage field, so you pay more. This article explains how the problem works, how to recognize it, and how to detect it automatically with tools.
What Is Token Padding?
Token padding (also called token inflation) is when an AI API relay, after calling the upstream model, modifies the usage object it returns, inflating the actual consumed token counts before passing them back to you:
// Actual response from the upstream model
{
"usage": {
"prompt_tokens": 120,
"completion_tokens": 80,
"total_tokens": 200
}
}
// Response after the relay tampered with it and returned it to you
{
"usage": {
"prompt_tokens": 156, // inflated by 30%
"completion_tokens": 104, // inflated by 30%
"total_tokens": 260
}
}
Because most developers do not check usage line by line, minor padding (5–15%) is rarely noticed. But at a monthly volume of several million tokens, the extra cost adds up considerably.
Token Padding vs Model Substitution: Which Is Harder to Spot?
Model substitution is commonly called "jiangzhi" (降智, capability downgrading) in Chinese-speaking communities. For how to tell, see What Does "Jiangzhi" (Capability Downgrade) Mean? Three Detection Methods.
| Issue type | Method | Impact | Difficulty of detection |
|---|---|---|---|
| Model substitution | Replaces the high-priced model you specified (e.g., GPT-4o) with a cheaper one (e.g., GPT-3.5) | Response quality drops, but you still pay the price of the higher-priced model | Medium (detectable through capability tests) |
| Token padding | Uses the correct model, but inflates token counts in the usage field | Response quality is normal, but you pay 5–30% more | High (requires comparing against raw token counts) |
| System prompt injection | The relay inserts hidden instructions before and/or after your prompt | Model behavior is manipulated, and your prompt_tokens rise because of the injection | High (requires comparing token counts and response behavior) |
How Large Is Typical Padding?
Based on public research and community reports, the size of token padding varies widely:
- Minor (5–10%): the most common; rarely noticed, but significant over time
- Moderate (10–20%): requires careful bill checking to detect
- Severe (30% or more): the difference is obvious even in a single request
Note: Illustrative calculation
If your monthly usage is 5 million prompt_tokens (GPT-4o billed at $2.50 per 1M tokens), 15% padding means you overpay by 750K tokens × $2.50 = about $1.875 per month. At scale (50 million tokens per month), that is a hidden loss of about $18.75.
How to Check for Token Padding Manually
Method 1: Send a prompt of known length
Use tiktoken or another tokenizer to count the exact tokens in your prompt, then compare with the usage.prompt_tokens the API returns:
# Python example (requires tiktoken)
import tiktoken, openai
enc = tiktoken.encoding_for_model("gpt-4o")
prompt = "What is 2+2?"
expected_tokens = len(enc.encode(prompt))
resp = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}]
)
actual = resp.usage.prompt_tokens
diff_pct = (actual - expected_tokens) / expected_tokens * 100
print(f"expected: {expected_tokens}, reported: {actual}, gap: {diff_pct:.1f}%")
Method 2: Detect automatically with BazaarLink Probe
Manual counting is slow and requires knowing each model's tokenizer differences. A faster option is BazaarLink Probe, which automatically runs a set of standardized tests, including token count comparison, and outputs a 0–100 score with a detailed report.
- Enter your API endpoint and key
- Probe automatically sends prompts with precisely calculated token counts
- Compares the returned usage against expected values
- Gap exceeds the threshold → flagged as "Token padding risk"
- Also checks for model substitution, system prompt injection, and other risks
Choosing an AI API Provider: Billing Transparency Compared
The most fundamental way to avoid token padding is to choose a provider that connects directly to the official vendor, or a platform with clear, transparent billing. Here is a comparison of several major options:
| Platform | Billing transparency | Taiwan unified invoice | Relay risk | Model count |
|---|---|---|---|---|
| BazaarLink | Provider and usage data available for cross-checking | Depends on checkout channel or enterprise contract | Verify with Probe and your actual records | Mainstream AI models |
| SiliconFlow (硅基流動) | Billed in RMB (mainland China) | Per provider billing documents | Medium | 100+ |
| Groq | USD billing (US), speed-focused | Per provider billing documents | Low (own inference infrastructure) | 20+ |
| Together AI | USD billing (US) | Per provider billing documents | Low (own inference) | 50+ |
| Unidentified relay | Unknown, no published rates | ✗ | High | Varies |
Technical Background on Token Padding
LLM supply chain attacks have drawn ongoing attention from academia and security researchers in recent years. The paper Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain (ACM CCS 2026) systematically classifies relay attack techniques, including:
- AC-1 Payload Injection: model substitution
- AC-1.b Conditional Injection: conditional system prompt injection
- Token Inflation: token padding (tampering with the usage field)
- AC-2 Secret Exfiltration: key theft risk
BazaarLink Probe's design draws on this research framework and covers automated detection for each of the attack categories above.
Conclusion
Token padding is a hidden way for AI API relays to profit. It is hard to see with the naked eye, but automated tools can detect it systematically. Whether you are an individual developer or an enterprise user, billing transparency and verifiability are key considerations when choosing an AI API provider.
Go to BazaarLink Probe now to test the endpoint you are currently using and check for token padding or other supply chain risks.
FAQ
What is token padding?
Token padding (token inflation) is when an AI API relay, when returning the usage field, reports more prompt_tokens or completion_tokens than were actually consumed, so the user pays more. It is one of the common ways relays make money.
How is token padding different from model substitution?
Model substitution is when a relay uses a cheaper model (such as GPT-3.5) in place of the model you specified (such as GPT-4o), while token counts may still be accurate. Token padding uses the correct model but inflates token counts in the usage field to charge more. Both can be detected automatically with BazaarLink Probe.
How can I tell whether an API is padding tokens?
The most reliable method is to send a fixed prompt of known length and compare the returned prompt_tokens against the expected value. BazaarLink Probe runs this test automatically and presents results as a 0–100 score; a gap above the threshold is flagged as a risk.
How large is the typical token padding gap?
Minor padding is usually within 5–15% and hard to notice; severe padding can exceed 30%. With high monthly token usage, the cumulative overpayment can be substantial.
How does BazaarLink Probe detect token padding?
Probe sends a prompt whose token count has been precisely calculated, then compares it against the values in the usage field returned. If the gap exceeds a reasonable threshold, it is flagged as a "token padding risk" and shown in the detection report.
Which situations are most prone to token padding?
Relays that resell API access on a pay-per-use basis are the highest risk, especially those that do not clearly publish their billing rules. Connecting directly to official providers through BazaarLink lets you check your bill against official usage, with no intermediary layer that could inflate figures.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API