AI Search API: How It Works, How to Evaluate (30 Tests)
AI search APIs let models check latest information. A 30-question test covers how they work, depth vs. news-mode tradeoffs, and how to judge speed and cost.
A large language model's knowledge stops at its training cutoff. For questions like "news from the past week" or "regulations announced only this month," it can only guess. AI search APIs fill that gap: they first look things up on the web for the model, then hand the organized results to the model to answer.
This article covers two things: how a search API actually works within a single conversation, and what we learned from running a fixed question bank about how to evaluate a search API. All numbers in this article come from the same controlled test; the test conditions are listed at the end so you can reproduce it yourself.
How AI search APIs work
A conversation with search typically goes through four steps:
- Decide whether to search: The model (or your program) looks at the question and decides whether it needs real-time data.
- Send the query: Rewrite the question as a search query, call the search API, and usually specify how many results you want.
- Retrieve results: The search API returns a set of URLs, titles, and content summaries, some with publication dates.
- Feed the model: Put this content into the prompt so the model answers based on the sources and cites them.
Unlike a regular Google search, the recipient is a program, not a person. So the evaluation focus is not how nice the page looks, but three things: whether the content found is relevant, whether it is fresh, and whether it is good enough for the model to quote directly.
If you just want the model's answers to include web data, you don't need to wire up a search API yourself: BazaarLink supports web search (:online). See the web search feature guide.
How to evaluate a search API: fixed question bank, fixed scoring, rules set before testing
The most common mistake when comparing search APIs is "asking a few questions at random and deciding this one seems more accurate." That kind of comparison can't be reproduced, and it's easily skewed by the one or two questions you remember best. Our approach was:
- The question bank was fixed and publicly archived before testing: 30 questions in total, 25 in Traditional Chinese and 5 in English, covering Taiwan news, regulations, government announcements, medical health education, academic topics, local businesses, and place names. 8 were marked as "requires recent data." No questions were swapped or changed during testing.
- Each question was run once per configuration: no retries, no cherry-picking results, a concurrency limit of 2, and a 30-second timeout per call. Every call requested 5 results.
- Scoring rules were written as a rubric first: each question had a 0–5 scoring standard, and the same fixed judge model with the same prompt scored every item, so the standard never changed based on the question or source.
- Raw data was kept for every number: search responses, scoring responses, and costs were all archived item by item so they can be checked later.
The key point of this process is reproducibility: swap in a different search API or different settings, and as long as you rerun with the same question bank and rubric, you get numbers that can be compared side by side.
Depth setting: twice the cost, what do you get?
Most search APIs offer a "search depth" option: shallow is faster and cheaper, while deep performs an extra round of content extraction. We compared two depths (called "Standard" and "Advanced" below) on the same search API. Results for the 30 general questions:
| Standard | Advanced | |
|---|---|---|
| Successful calls | 30/30 | 30/30 |
| Average relevance (0–5) | 3.260 | 3.153 |
| Latency p50 | 1,998 ms | 5,061 ms |
| Latency p95 | 4,033 ms | 7,014 ms |
| Cost per search (BazaarLink customer price) | About $0.0096 | About $0.0192 |
In this question bank, the advanced depth did not bring higher average relevance, while costing about twice as much time and money. So "deeper is better" is not the default answer: run your own question bank with the cheaper setting first, and only move up if it proves insufficient.
Language also made a difference. Standard depth scored 3.392 on Traditional Chinese questions and 2.600 on English; advanced depth scored 3.200 and 2.920. In other words, test with questions in the language your users mainly ask in, and don't infer Traditional Chinese performance from someone else's English test results. Also, 96% of results for Traditional Chinese questions came from Taiwan or Traditional Chinese sources, versus only 24% for English questions. If you need local sources, check this separately.
News mode: getting a date isn't the same as answering correctly
When asking "what happened in the past week," the hardest part for a search API is timeliness. We split the 8 questions that need recent data into two test conditions:
General mode: none of the 300 results had a parseable publication date. This doesn't mean the content is outdated; it means you cannot tell from the data itself whether it is new or old, so your program has no way to filter it.
News mode + time range (limited to the past week or past month): the date field became usable, and all results with dates fell within 30 days. But relevance differed greatly:
| Standard | Advanced | |
|---|---|---|
| Results with a publication date | 36/40 | 34/40 |
| Average relevance (0–5, 8 questions) | 0.475 | 1.900 |
| Share of Taiwan / Traditional Chinese sources | 37.5% | 82.5% |
Two key points:
- Having a date doesn't mean it's useful. At standard depth, 90% of results had dates, but the first few Taiwan news questions almost all got irrelevant international news, with relevance of only 0.475.
- Evaluate news questions separately from general questions. Advanced depth clearly did better on news questions (1.900), but still below the 3+ level for general questions. Timeliness questions are harder than they look; don't assume the problem is solved just because the date field has values.
In practice: the program first decides whether the question "needs recent data." General questions use the cheaper setting, news questions switch to the deeper setting with a time range added, and finally the date and source are shown to the user alongside the answer.
Speed and cost
- Look at p95 as well as p50. The median may look fast, but that doesn't mean it's never slow: the p95 for both depths is about 2 seconds slower than p50. Search typically completes before the model answers, so this waiting time is added directly to the reply time the user experiences.
- Count cost together with model tokens. Search fees are only part of the picture. Stuffing 5 results into the prompt increases input tokens, and that must be counted in the cost of each answer.
- Estimate usage before choosing a plan. Many teams' actual search volume is far below expectations, and the free quota is enough. Measure your real monthly search count first, then decide whether to pay to upgrade.
Your evaluation checklist
- First write 20–30 of your own questions covering your language and domain, and mark which ones need recent data.
- Write the rubric first (what counts as relevant, what counts as wrong) before you start testing.
- Run each configuration once against the same question bank, and keep the raw responses.
- Evaluate general and news questions separately, and different languages separately.
- Record p50, p95, and the cost per search at the same time.
- Results only represent this question bank: whenever the bank changes or the search API is updated, retest.
FAQ
How is an AI search API different from a regular search engine?
A search engine's results are pages meant for people; a search API returns structured data meant for programs (URL, title, content summary, date), which can be placed directly into the model's prompt.
Why not just let the model "go online" itself?
The model itself cannot access the internet. For it to look up information, your program or platform has to call the search API on its behalf and then hand the results to the model. That's why how accurate the lookup is depends on the search API, not just the model.
Can I use this test result directly as a basis for purchasing decisions?
Not recommended. This is a controlled test of 30 questions, suitable for learning evaluation methods and observing tradeoffs, not a permanent ranking across all languages, regions, or question types. Please retest with your own question bank.
Is there an easier way to have the model's answers include web data?
Yes. BazaarLink supports web search (:online), so you don't need to wire up a search API yourself; see the web search feature guide.
Test conditions (for reproduction)
- 30-question fixed bank (25 Traditional Chinese, 5 English; 8 requiring recent data); the question bank and rubric were archived before testing.
- Each question was run once per configuration, requesting 5 results per call, with a concurrency limit of 2, a 30-second timeout, and no retries.
- Scoring: fixed judge model, fixed prompt, 0–5 points.
- Costs are BazaarLink web search customer prices (based on pricing as of September 29, 2026). Latency is measured at the time of testing, for reference only, and does not represent any service level commitment.
- News mode supplementary test: for the 8 recent-data questions, added news topic and time range (past week or past month).
- Test date: September 28, 2026.
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API