> Method
Measured, not guessed.
An MCP server is only useful if an agent picks the right tool and fills in the right arguments. I measure that instead of assuming it.
01 The test set
Each server gets 30 test prompts. Twenty are used while I tune. Ten are held back and not looked at until the end. Four of the 30 are cases where the right answer is to refuse or ask a question, and three are deliberately ambiguous. For every prompt I write down the expected tool and arguments before testing.
02 What I score
- Right tool chosen
- Right arguments filled in
- Correct refusal or clarifying question when no tool fits
- Safety: any write or delete tool called on a read-only request
- Efficiency: how many calls a task needs
03 Before and after
I run the set on the untuned server, improve the tool list and descriptions, then run it again. Agents vary from run to run, so each run is repeated three times. You get both sets of scores and a list of the cases that changed.
04 The risk report
For every operation in your OpenAPI file I record whether it is read-only, a create, an update, a delete, an action disguised as a GET, or something that costs money or contacts someone outside. I also report how complete the spec is (operation names, descriptions, response schemas, examples). A GET request that deletes data after you read it is flagged as destructive, because an agent would otherwise treat it as safe.
05 One tool at a time
I build and check each tool separately instead of generating everything at once. A tool is finished only when it has a clear description, a read-only, write or destructive label, and automated test calls that pass without a chat window, including a case that should fail. Then it gets two reviews, the second one aimed at finding a prompt that would make an agent choose the wrong tool. You approve the tool list before the next tool starts.
The point is that mistakes are found early, one tool at a time, and every result is written down so you can see it.
What this does not show
Thirty prompts is a small sample, and ten held-back cases are noisy. Results apply to the model and prompt wording I tested, not to every agent. The test does not measure adoption, traffic or revenue, and it is not a security audit or certification.