Answer test, October 2026
How often an AI assistant got plan costs and plan choices right with web search alone, and with this server connected. Run on 2026-10-06.
Setup
- Host model: gpt-5.6-sol (reasoning low) in Codex CLI 0.148.0, one session per question and arm.
- Arms compared: web search only, and this server's MCP tools (with web search also available).
- Questions: 30 buyer questions about web scraping, crawling and SERP API plans. One asks the buyer for missing details first and is scored apart; two have no valid answer key, so 27 are scored for cost. A further 10 held-out questions check the direction.
- Answer key: a blind labeler (claude-opus-5-5, reading only the providers' official pages) labeled every plan named by any arm, without knowing which arm named it. The labeler planned first failed its calibration, so this pre-registered fallback was used. On 8 questions labeled twice, the two labels agreed on cost for 63 of 73 plans.
- The questions, rules and claims were registered before the run.
Results
| Measure | Web search only | With this server | Difference (95% CI) | Held-out questions | Registered rule |
|---|---|---|---|---|---|
| Plan cost errors | 12.5% (8 of 64) | 3.6% (2 of 55) | 8.9 points (1.3 to 15.9) fewer errors | 26.3% vs 11.5%, same direction | met |
| First recommendation correct | 65.4% (17 of 26) | 77.8% (21 of 27) | 21.7 points (−3.3 to 43.9), 23 paired questions | not used | not met: the interval includes 0 |
| Unmet plans stated as meeting every condition | 1.9% (1 of 54) | 6.9% (4 of 58) | 5.0 points (−11.5 to 0.0) more with this server | 6.7 points more with this server | not met: the estimate favours web search alone |
Only the first result passed its registered rule, so only it is quoted in the server's listing.
Method
- Cost errors and unmet plans stated as met: rates over all plans of the scored questions; the interval is a question-cluster bootstrap (questions drawn with replacement, seed 20261006, 10,000 resamples, percentile 95% interval).
- First recommendation correct: one result per question, paired by question; Newcombe (1998) paired method 10, 95%.
- A claim passes when the lower end of its interval is above 0 and, for the two rates, the held-out questions point the same way.
Limits
- One host model at one reasoning setting; other assistants may differ.
- The answer key is one model's reading of the official pages and can be wrong; double labels disagreed on cost for 10 of 73 plans.
- The scored questions are the in-sample set the service was built with; the 10 held-out questions only check the direction.
- Prices change. The results describe answers given on 2026-10-06.