# Model Substitution in the Black-Box LLM API Resale Market

## A 14-Day, 171-Endpoint, 625-Probe Empirical Measurement Brief

- Paper ID: `OR-2026-04-26-resale-substitution`
- Publisher: BazaarLink Research
- Published: 2026-04-26
- Last updated: 2026-07-29
- Measurement window: 2026-04-12 to 2026-04-25
- Canonical: https://bazaarlink.ai/research/llm-resale-substitution-2026

## Abstract

BazaarLink Research measured 171 distinct black-box LLM API reseller endpoints over 14 calendar days and collected 625 probe runs. The study looked for observable disagreement between an endpoint's advertised model and independent behavioral or protocol signals. 10 of 171 endpoints showed at least one detected substitution signal. Under a stricter repeated-observation rule, 2 of 149 evaluable endpoints showed persistent signals.

These measurements identify signals, not legal conclusions about an operator. Endpoint identities are anonymized in this public brief.

## Methodology

The measurement system combines four independent black-box channels:

1. Protocol and surface fingerprints, including response shape and metadata.
2. Behavioral fingerprints designed to distinguish model families.
3. Deterministic sub-model checks for version-level consistency.
4. Verdict fusion that requires agreement across independent evidence channels.

Synthetic, cost-constrained checks reported 94.4% true-positive detection for the tested intra-family substitutions and 0% false positives for the tested genuine-model controls. Sample sizes and the exact evaluation boundary are part of the limitations below.

## Principal measurements

| Measure | Result |
|---|---:|
| Measurement window | 14 days |
| Distinct endpoints | 171 |
| Probe runs | 625 |
| Endpoints with at least one signal | 10/171 (approximately 5.8%) |
| Persistent-signal endpoints | 2/149 (1.3%) |
| Synthetic intra-family true-positive rate | 94.4% |
| Synthetic genuine-model false-positive rate | 0% |
| Synthetic substitution samples | n=18 |
| Synthetic genuine-model controls | n=6 |

## Limitations

- Black-box inference is probabilistic; it cannot directly inspect an upstream routing decision.
- Low-sample endpoints are not used for the principal persistent-substitution estimate.
- An endpoint may route differently during detected test traffic than during ordinary customer traffic.
- Synthetic validation used cost-constrained sample sizes and does not cover every model, language, or deployment configuration.
- The measurement window ended on 2026-04-25; current endpoint behavior may differ.
- This public version omits the internal full-domain mapping and does not make legal claims about any operator.

## Research ethics

Measurements used legitimately purchased access and ordinary API requests. The study did not attempt denial of service, billing bypass, unauthorized access, or access beyond the public service contract. Publication is intended for consumer protection and market transparency.

## Provenance and reproducibility

- Live public Probe explorer: https://bazaarlink.ai/probe
- Model-level public datasets: https://bazaarlink.ai/probe/stats/
- Relay-level public datasets: https://bazaarlink.ai/probe/relay
- Open-source detection engine: https://github.com/Bazaarlinkorg/LLMprobe-engine
- Machine-readable citation metadata: embedded as Dataset JSON-LD on the canonical HTML page

## Citation

BazaarLink Research. "Model Substitution in the Black-Box LLM API Resale Market: A 14-Day, 171-Endpoint, 625-Probe Empirical Measurement Brief." OR-2026-04-26-resale-substitution, 2026-04-26. https://bazaarlink.ai/research/llm-resale-substitution-2026
