Agents need fast feedback loops: why an MCP server should be testable without a chat client

By

An MCP server that can only be tried by opening a chat client is slow to improve. Every change means restarting a client, writing a prompt, waiting for a model, and guessing whether a failure came from the tool, its description or the model. When a server can also be run from a script, I find it quicker to improve, because I can see which part failed.

What “headless” means here

Run the server, list its tools, call each tool with fixed inputs and compare the output to what you expect, with no model in the loop. Because no model is involved, the same inputs should give the same outputs, and a failed check points at the server rather than at the model.

Three checks worth running on every change

  1. The tool list. Names, descriptions and input schemas load, and each description says what the tool does and when not to use it.
  2. Known-good calls. For each tool, one input with a stored expected output (a fixture), including a failure case such as a missing required field or a bad ID.
  3. Safety labels. Read-only tools are marked read-only, and anything that writes or deletes is clearly labelled, ideally behind an explicit confirmation step.

Then add the model, last

Once the plain checks pass, test with a model using a small fixed set of prompts and record which tool it picks and with what arguments. On my own tests I use 30 prompts, 20 for tuning and 10 held back (see the Method page). When a model picks the wrong tool, the cause is often in the tool descriptions, so that is where I look first.

Where to start

If you want a first look at your own API, the free spec check on this site is the starting point: send a link to your public OpenAPI file and I will read it offline, without calling your API, and send back a short written snapshot.